No description
  • Dockerfile 100%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Brian Clemens 198f50320b
All checks were successful
Build and publish / build (vulkan) (push) Successful in 26s
Build and publish / build (rocm-gfx1151) (push) Successful in 3m26s
Merge pull request 'Update Unsloth llama.cpp prebuilt to v11541' (#19) from renovate/unslothai-llama.cpp-11541.x into main
Reviewed-on: #19
2026-10-11 15:31:20 +00:00
.forgejo/workflows Serialize matrix builds to avoid Forgejo registry blob race 2026-09-13 14:03:01 +09:00
Containerfile Merge pull request 'Update Unsloth llama.cpp prebuilt to v11541' (#19) from renovate/unslothai-llama.cpp-11541.x into main 2026-10-11 15:31:20 +00:00
README.md Add Forgejo Actions workflow publishing to the built-in registry 2026-09-13 13:43:14 +09:00
renovate.json Add Renovate: track unsloth releases + base image digest 2026-09-13 13:03:33 +09:00

llama-unsloth-container

Container image for Unsloth's llama.cpp builds (Vulkan backend), overlaid on the upstream ghcr.io/ggml-org/llama.cpp:full-vulkan runtime image.

Deployed on shikigami (Strix Halo / Radeon 8060S, gfx1151) as the rootful podman quadlet llama.service (/etc/containers/systemd/llama.container).

Why not stock upstream?

Unsloth's MTP draft heads (e.g. unsloth/Qwen3.8-Flash-Next-GGUF/MTP/*) use the fork's tensor naming (blk.48.nextn.hc_head_*). Mainline's merged MTP implementation expects top-level output_hc_* names and fails to load them:

llama_model_load: error loading model: check_tensor_dims: tensor 'output_hc_norm.weight' not found

Same tensors, same bytes, different names. Running the fork's binaries makes Unsloth's heads (and their day-0 model support / extra quant types) work as published.

Why not stock unsloth?

They don't publish container images, only per-backend tarballs. The tarball is self-contained (RUNPATH=$ORIGIN, glibc ≤ 2.34) but needs a Vulkan loader + Mesa RADV from the OS, which the upstream image already provides.

Images

CI (.forgejo/workflows/build.yaml) builds and publishes to the Forgejo registry on every push to main that touches the Containerfile; pull requests get build-only validation (so Renovate bumps are proven before merge):

tag backend
git.tiuxo.com/brian/llama-unsloth:latest Vulkan (deployed)
git.tiuxo.com/brian/llama-unsloth:<UNSLOTH_TAG> Vulkan, pinned
git.tiuxo.com/brian/llama-unsloth:latest-rocm-gfx1151 ROCm gfx1151
git.tiuxo.com/brian/llama-unsloth:<UNSLOTH_TAG>-rocm-gfx1151 ROCm gfx1151, pinned

Manual build

sudo podman build -t localhost/llama-unsloth:latest .

Backend variants

UNSLOTH_BACKEND selects the release asset (default vulkan):

sudo podman build --build-arg UNSLOTH_BACKEND=rocm-gfx1151 \
    -t localhost/llama-unsloth:b10909-mix-bea84f7-rocm-gfx1151 .

The ROCm tarballs bundle the full ROCm userspace (HIP/HSA/rocBLAS + librocm_sysdeps_*), so they run on the same Vulkan base image — only /dev/kfd + /dev/dri come from the host. HIP sees the full GTT as device memory on Strix Halo.

Vulkan is the deployed backend. A/B on shikigami (b10909, Qwen3.8-Flash-Next UD-Q4_K_XL + MTP n-max 2, temp 1.0, 2 reps each):

workload Vulkan (tok/s) ROCm gfx1151 (tok/s)
short-ctx decode 41.5–42.5 36.1–38.0
~6.7k-ctx decode 38.3–39.7 32.2–34.1
~6.7k-ctx prefill 280 280

Draft acceptance was identical (~0.78) on both. Worth re-testing on future releases; the gap is backend kernels, not speculation.

Updates (Renovate)

renovate.json manages both moving parts:

  • UNSLOTH_TAG — a custom regex manager reads the # renovate: annotation above the ARG and tracks unslothai/llama.cpp releases (versioning compares the upstream build number in b<NNNNN>-mix-<sha>)
  • FROM digest — the stock dockerfile manager updates the pinned digest for the :full-vulkan tag

Integrity: the build downloads the release's own llama-prebuilt-sha256.json manifest and verifies the tarball against it, so a Renovate bump is a single-line diff with no manual checksum step.

The full pipeline: Renovate opens a PR (CI validates the build) → merge → CI publishes to the registry → podman-auto-update.timer on shikigami pulls the new :latest and restarts llama.service.

Notes

  • The quadlet uses AutoUpdate=registry; podman auto-update rolls back if the updated container fails to start. A restart drops the loaded model until the next request triggers a reload (~80 s).
  • RUN /app/llama-server --version in the build doubles as a linker smoke test; a glibc/ABI mismatch fails the build, not the deployment.