Update Unsloth llama.cpp prebuilt to v11443 - autoclosed #16

Closed
brian wants to merge 1 commit from renovate/unslothai-llama.cpp-11443.x into main
Owner

This PR contains the following updates:

Package Update Change
unslothai/llama.cpp major b11160-mix-a6922cc → b11443-mix-d65395f

Release Notes

unslothai/llama.cpp

vb11443-mix-d65395f: llama.cpp prebuilt b11443-mix-d65395f

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b11443, merged with:

vb11438-mix-d65395f: llama.cpp prebuilt b11438-mix-d65395f

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b11438, merged with:

vb11408-mix-1e24fc5: llama.cpp prebuilt b11408-mix-1e24fc5

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b11408, merged with:


Configuration

📅 Schedule: (UTC)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

♻ Rebasing: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about this update again.


  • If you want to rebase/retry this PR, check this box

This PR has been generated by Mend Renovate CLI.

This PR contains the following updates: | Package | Update | Change | |---|---|---| | [unslothai/llama.cpp](https://github.com/unslothai/llama.cpp) | major | `b11160-mix-a6922cc` → `b11443-mix-d65395f` | --- ### Release Notes <details> <summary>unslothai/llama.cpp</summary> ### [`vb11443-mix-d65395f`](https://github.com/unslothai/llama.cpp/releases/tag/b11443-mix-d65395f): llama.cpp prebuilt b11443-mix-d65395f [Compare Source](https://github.com/unslothai/llama.cpp/compare/b11438-mix-d65395f...b11443-mix-d65395f) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b11443](https://github.com/ggml-org/llama.cpp/releases/tag/b11443), merged with: - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [6431891](https://github.com/ggml-org/llama.cpp/pull/24423/commits/643189144c98fba431e8d2da169210b7b6d5d930)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [efd2b13](https://github.com/ggml-org/llama.cpp/pull/25731/commits/efd2b13f80c68df745f0134a2a0f56a434a2d33d)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1_XS, IQ1_XXS, IQ1_XXXS: three quant types below IQ1_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [ef45f21](https://github.com/unslothai/llama.cpp/pull/61/commits/ef45f212c90d5769e0dd8d482f99fd66760312e7)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - ggml-cuda: avoid direct ROCm_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml_cuda_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML_CUDA_ENABLE_UNIFIED_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [e2519eb](https://github.com/unslothai/llama.cpp/pull/176/commits/e2519ebc09468b477847a7529834a8da998ed2da)) - Load GLM-5-Next GGUFs converted with the earlier glm5next arch name ([unslothai/llama.cpp#239](https://github.com/unslothai/llama.cpp/pull/239), commit [bec2164](https://github.com/unslothai/llama.cpp/pull/239/commits/bec216444552a20bc39bf2156fa5396154970608)) - qwen4exp: run MTP heads that share the target's token_embd and output ([unslothai/llama.cpp#240](https://github.com/unslothai/llama.cpp/pull/240), commit [36175a8](https://github.com/unslothai/llama.cpp/pull/240/commits/36175a8325e24303704cfbab1b7d8a0bd30500a5)) - Carry what the dropped GLM-5-Next, Qwen MTP and readahead pins had beyond upstream ([unslothai/llama.cpp#241](https://github.com/unslothai/llama.cpp/pull/241), commit [b96a713](https://github.com/unslothai/llama.cpp/pull/241/commits/b96a713a48313a32c92fd6f77b367ea9d803417b)) - glm5-next : keep the indexer compute buffer bounded at long contexts ([unslothai/llama.cpp#243](https://github.com/unslothai/llama.cpp/pull/243), commit [aba4a4c](https://github.com/unslothai/llama.cpp/pull/243/commits/aba4a4c1db96d56846e86d1169dd35241bafd55d)) - EmbeddingGemma-2 support ([unslothai/llama.cpp#247](https://github.com/unslothai/llama.cpp/pull/247), commit [73c2f73](https://github.com/unslothai/llama.cpp/pull/247/commits/73c2f733e8beef7dd104de747442eb9e21f2bb96)) ### [`vb11438-mix-d65395f`](https://github.com/unslothai/llama.cpp/releases/tag/b11438-mix-d65395f): llama.cpp prebuilt b11438-mix-d65395f [Compare Source](https://github.com/unslothai/llama.cpp/compare/b11408-mix-1e24fc5...b11438-mix-d65395f) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b11438](https://github.com/ggml-org/llama.cpp/releases/tag/b11438), merged with: - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [6431891](https://github.com/ggml-org/llama.cpp/pull/24423/commits/643189144c98fba431e8d2da169210b7b6d5d930)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [efd2b13](https://github.com/ggml-org/llama.cpp/pull/25731/commits/efd2b13f80c68df745f0134a2a0f56a434a2d33d)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1_XS, IQ1_XXS, IQ1_XXXS: three quant types below IQ1_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [ef45f21](https://github.com/unslothai/llama.cpp/pull/61/commits/ef45f212c90d5769e0dd8d482f99fd66760312e7)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - ggml-cuda: avoid direct ROCm_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml_cuda_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML_CUDA_ENABLE_UNIFIED_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [e2519eb](https://github.com/unslothai/llama.cpp/pull/176/commits/e2519ebc09468b477847a7529834a8da998ed2da)) - Load GLM-5-Next GGUFs converted with the earlier glm5next arch name ([unslothai/llama.cpp#239](https://github.com/unslothai/llama.cpp/pull/239), commit [bec2164](https://github.com/unslothai/llama.cpp/pull/239/commits/bec216444552a20bc39bf2156fa5396154970608)) - qwen4exp: run MTP heads that share the target's token_embd and output ([unslothai/llama.cpp#240](https://github.com/unslothai/llama.cpp/pull/240), commit [36175a8](https://github.com/unslothai/llama.cpp/pull/240/commits/36175a8325e24303704cfbab1b7d8a0bd30500a5)) - Carry what the dropped GLM-5-Next, Qwen MTP and readahead pins had beyond upstream ([unslothai/llama.cpp#241](https://github.com/unslothai/llama.cpp/pull/241), commit [b96a713](https://github.com/unslothai/llama.cpp/pull/241/commits/b96a713a48313a32c92fd6f77b367ea9d803417b)) - glm5-next : keep the indexer compute buffer bounded at long contexts ([unslothai/llama.cpp#243](https://github.com/unslothai/llama.cpp/pull/243), commit [aba4a4c](https://github.com/unslothai/llama.cpp/pull/243/commits/aba4a4c1db96d56846e86d1169dd35241bafd55d)) - EmbeddingGemma-2 support ([unslothai/llama.cpp#247](https://github.com/unslothai/llama.cpp/pull/247), commit [73c2f73](https://github.com/unslothai/llama.cpp/pull/247/commits/73c2f733e8beef7dd104de747442eb9e21f2bb96)) ### [`vb11408-mix-1e24fc5`](https://github.com/unslothai/llama.cpp/releases/tag/b11408-mix-1e24fc5): llama.cpp prebuilt b11408-mix-1e24fc5 [Compare Source](https://github.com/unslothai/llama.cpp/compare/b11160-mix-a6922cc...b11408-mix-1e24fc5) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b11408](https://github.com/ggml-org/llama.cpp/releases/tag/b11408), merged with: - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [6431891](https://github.com/ggml-org/llama.cpp/pull/24423/commits/643189144c98fba431e8d2da169210b7b6d5d930)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [efd2b13](https://github.com/ggml-org/llama.cpp/pull/25731/commits/efd2b13f80c68df745f0134a2a0f56a434a2d33d)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1_XS, IQ1_XXS, IQ1_XXXS: three quant types below IQ1_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [ef45f21](https://github.com/unslothai/llama.cpp/pull/61/commits/ef45f212c90d5769e0dd8d482f99fd66760312e7)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - ggml-cuda: avoid direct ROCm_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml_cuda_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML_CUDA_ENABLE_UNIFIED_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [e2519eb](https://github.com/unslothai/llama.cpp/pull/176/commits/e2519ebc09468b477847a7529834a8da998ed2da)) - Load GLM-5-Next GGUFs converted with the earlier glm5next arch name ([unslothai/llama.cpp#239](https://github.com/unslothai/llama.cpp/pull/239), commit [bec2164](https://github.com/unslothai/llama.cpp/pull/239/commits/bec216444552a20bc39bf2156fa5396154970608)) - qwen4exp: run MTP heads that share the target's token_embd and output ([unslothai/llama.cpp#240](https://github.com/unslothai/llama.cpp/pull/240), commit [36175a8](https://github.com/unslothai/llama.cpp/pull/240/commits/36175a8325e24303704cfbab1b7d8a0bd30500a5)) - Carry what the dropped GLM-5-Next, Qwen MTP and readahead pins had beyond upstream ([unslothai/llama.cpp#241](https://github.com/unslothai/llama.cpp/pull/241), commit [b96a713](https://github.com/unslothai/llama.cpp/pull/241/commits/b96a713a48313a32c92fd6f77b367ea9d803417b)) </details> --- ### Configuration 📅 **Schedule**: (UTC) - Branch creation - At any time (no schedule defined) - Automerge - At any time (no schedule defined) 🚦 **Automerge**: Disabled by config. Please merge this manually once you are satisfied. ♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox. 🔕 **Ignore**: Close this PR and you won't be reminded about this update again. --- - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this box --- This PR has been generated by [Mend Renovate CLI](https://github.com/renovatebot/renovate). <!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0NC4xMzguMCIsInVwZGF0ZWRJblZlciI6IjQ0LjE0Mi4xIiwidGFyZ2V0QnJhbmNoIjoibWFpbiIsImxhYmVscyI6WyJyZW5vdmF0ZSJdfQ==-->
Update Unsloth llama.cpp prebuilt to v11443
All checks were successful
Build and publish / build (vulkan) (pull_request) Successful in 23s
Build and publish / build (rocm-gfx1151) (pull_request) Successful in 1m43s
37625ff643
brian changed title from Update Unsloth llama.cpp prebuilt to v11443 to Update Unsloth llama.cpp prebuilt to v11443 - autoclosed 2026-10-09 00:06:34 +00:00
brian closed this pull request 2026-10-09 00:06:34 +00:00
All checks were successful
Build and publish / build (vulkan) (pull_request) Successful in 23s
Build and publish / build (rocm-gfx1151) (pull_request) Successful in 1m43s

Pull request closed

Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brian/llama-unsloth-container!16
No description provided.