- Dockerfile 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
|
||
| .forgejo/workflows | ||
| Containerfile | ||
| README.md | ||
| renovate.json | ||
llama-unsloth-container
Container image for Unsloth's llama.cpp builds
(Vulkan backend), overlaid on the upstream ghcr.io/ggml-org/llama.cpp:full-vulkan
runtime image.
Deployed on shikigami (Strix Halo / Radeon 8060S, gfx1151) as the rootful
podman quadlet llama.service (/etc/containers/systemd/llama.container).
Why not stock upstream?
Unsloth's MTP draft heads (e.g. unsloth/Qwen3.8-Flash-Next-GGUF/MTP/*) use the
fork's tensor naming (blk.48.nextn.hc_head_*). Mainline's merged MTP
implementation expects top-level output_hc_* names and fails to load them:
llama_model_load: error loading model: check_tensor_dims: tensor 'output_hc_norm.weight' not found
Same tensors, same bytes, different names. Running the fork's binaries makes Unsloth's heads (and their day-0 model support / extra quant types) work as published.
Why not stock unsloth?
They don't publish container images, only per-backend tarballs. The tarball is
self-contained (RUNPATH=$ORIGIN, glibc ≤ 2.34) but needs a Vulkan
loader + Mesa RADV from the OS, which the upstream image already provides.
Images
CI (.forgejo/workflows/build.yaml) builds and publishes to the Forgejo
registry on every push to main that touches the Containerfile; pull
requests get build-only validation (so Renovate bumps are proven before
merge):
| tag | backend |
|---|---|
git.tiuxo.com/brian/llama-unsloth:latest |
Vulkan (deployed) |
git.tiuxo.com/brian/llama-unsloth:<UNSLOTH_TAG> |
Vulkan, pinned |
git.tiuxo.com/brian/llama-unsloth:latest-rocm-gfx1151 |
ROCm gfx1151 |
git.tiuxo.com/brian/llama-unsloth:<UNSLOTH_TAG>-rocm-gfx1151 |
ROCm gfx1151, pinned |
Manual build
sudo podman build -t localhost/llama-unsloth:latest .
Backend variants
UNSLOTH_BACKEND selects the release asset (default vulkan):
sudo podman build --build-arg UNSLOTH_BACKEND=rocm-gfx1151 \
-t localhost/llama-unsloth:b10909-mix-bea84f7-rocm-gfx1151 .
The ROCm tarballs bundle the full ROCm userspace (HIP/HSA/rocBLAS +
librocm_sysdeps_*), so they run on the same Vulkan base image — only
/dev/kfd + /dev/dri come from the host. HIP sees the full GTT as device
memory on Strix Halo.
Vulkan is the deployed backend. A/B on shikigami (b10909, Qwen3.8-Flash-Next UD-Q4_K_XL + MTP n-max 2, temp 1.0, 2 reps each):
| workload | Vulkan (tok/s) | ROCm gfx1151 (tok/s) |
|---|---|---|
| short-ctx decode | 41.5–42.5 | 36.1–38.0 |
| ~6.7k-ctx decode | 38.3–39.7 | 32.2–34.1 |
| ~6.7k-ctx prefill | 280 | 280 |
Draft acceptance was identical (~0.78) on both. Worth re-testing on future releases; the gap is backend kernels, not speculation.
Updates (Renovate)
renovate.json manages both moving parts:
UNSLOTH_TAG— a custom regex manager reads the# renovate:annotation above theARGand tracks unslothai/llama.cpp releases (versioning compares the upstream build number inb<NNNNN>-mix-<sha>)FROMdigest — the stock dockerfile manager updates the pinned digest for the:full-vulkantag
Integrity: the build downloads the release's own llama-prebuilt-sha256.json
manifest and verifies the tarball against it, so a Renovate bump is a
single-line diff with no manual checksum step.
The full pipeline: Renovate opens a PR (CI validates the build) → merge →
CI publishes to the registry → podman-auto-update.timer on shikigami
pulls the new :latest and restarts llama.service.
Notes
- The quadlet uses
AutoUpdate=registry;podman auto-updaterolls back if the updated container fails to start. A restart drops the loaded model until the next request triggers a reload (~80 s). RUN /app/llama-server --versionin the build doubles as a linker smoke test; a glibc/ABI mismatch fails the build, not the deployment.