Update Unsloth llama.cpp prebuilt to v11160 #12

Merged
brian merged 1 commit from renovate/unslothai-llama.cpp-11160.x into main 2026-09-30 07:30:26 +00:00
Owner

This PR contains the following updates:

Package Update Change
unslothai/llama.cpp major b10909-mix-bea84f7 → b11160-mix-a6922cc

Release Notes

unslothai/llama.cpp (unslothai/llama.cpp)

vb11160-mix-a6922cc: llama.cpp prebuilt b11160-mix-a6922cc

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b11160, merged with:

vb11139-mix-a6922cc: llama.cpp prebuilt b11139-mix-a6922cc

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b11139, merged with:

vb11115-mix-a6922cc: llama.cpp prebuilt b11115-mix-a6922cc

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b11115, merged with:

vb11030-mix-5ff778e: llama.cpp prebuilt b11030-mix-5ff778e

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b11030, merged with:

vb11027-mix-3e83366: llama.cpp prebuilt b11027-mix-3e83366

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b11027, merged with:

vb11007-mix-3e83366: llama.cpp prebuilt b11007-mix-3e83366

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b11007, merged with:

vb10995-mix-3e83366: llama.cpp prebuilt b10995-mix-3e83366

Compare Source

Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b10995, merged with:


Configuration

📅 Schedule: (UTC)

  • Branch creation
    • At any time (no schedule defined)
  • Automerge
    • At any time (no schedule defined)

🚦 Automerge: Disabled by config. Please merge this manually once you are satisfied.

♻ Rebasing: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox.

🔕 Ignore: Close this PR and you won't be reminded about this update again.


  • If you want to rebase/retry this PR, check this box

This PR has been generated by Mend Renovate CLI.

This PR contains the following updates: | Package | Update | Change | |---|---|---| | [unslothai/llama.cpp](https://github.com/unslothai/llama.cpp) | major | `b10909-mix-bea84f7` → `b11160-mix-a6922cc` | --- ### Release Notes <details> <summary>unslothai/llama.cpp (unslothai/llama.cpp)</summary> ### [`vb11160-mix-a6922cc`](https://github.com/unslothai/llama.cpp/releases/tag/b11160-mix-a6922cc): llama.cpp prebuilt b11160-mix-a6922cc [Compare Source](https://github.com/unslothai/llama.cpp/compare/b11139-mix-a6922cc...b11160-mix-a6922cc) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b11160](https://github.com/ggml-org/llama.cpp/releases/tag/b11160), merged with: - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [3dac51d](https://github.com/ggml-org/llama.cpp/pull/24423/commits/3dac51dc80b1400929266453c6ced6748ccf6862)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [3870aa4](https://github.com/ggml-org/llama.cpp/pull/25731/commits/3870aa48bbed3d5b1ee7cded0de688baa42f2470)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1\_XS, IQ1\_XXS, IQ1\_XXXS: three quant types below IQ1\_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [46cbf0e](https://github.com/unslothai/llama.cpp/pull/61/commits/46cbf0e95786fe8f5b7c0e86d57aaf8f8eceea7f)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - model: add GLM-5-Next (GLM-5.3-Flash) ([#&#8203;27754](https://github.com/ggml-org/llama.cpp/pull/27754), commit [86ebfef](https://github.com/ggml-org/llama.cpp/pull/27754/commits/86ebfef2c6a0f3359a2a07d2c215d61b0fa885c9)) - llama: batched readahead for lazily read gather tables ([unslothai/llama.cpp#137](https://github.com/unslothai/llama.cpp/pull/137), commit [4e1865e](https://github.com/unslothai/llama.cpp/pull/137/commits/4e1865e34ec5f6ca39403215c89129c13731be70)) - ggml-cuda: avoid direct ROCm\_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml\_cuda\_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML\_CUDA\_ENABLE\_UNIFIED\_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - MTP for Qwen3.8-Flash-Next ([unslothai/llama.cpp#144](https://github.com/unslothai/llama.cpp/pull/144), commit [ca14269](https://github.com/unslothai/llama.cpp/pull/144/commits/ca1426903fabe9af26cd10c42034cb4bbd2e0e11)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [09ce1a4](https://github.com/unslothai/llama.cpp/pull/176/commits/09ce1a4d2939844e211f7b4d30a296f4c1aed9a8)) ### [`vb11139-mix-a6922cc`](https://github.com/unslothai/llama.cpp/releases/tag/b11139-mix-a6922cc): llama.cpp prebuilt b11139-mix-a6922cc [Compare Source](https://github.com/unslothai/llama.cpp/compare/b11115-mix-a6922cc...b11139-mix-a6922cc) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b11139](https://github.com/ggml-org/llama.cpp/releases/tag/b11139), merged with: - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [3dac51d](https://github.com/ggml-org/llama.cpp/pull/24423/commits/3dac51dc80b1400929266453c6ced6748ccf6862)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [3870aa4](https://github.com/ggml-org/llama.cpp/pull/25731/commits/3870aa48bbed3d5b1ee7cded0de688baa42f2470)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1\_XS, IQ1\_XXS, IQ1\_XXXS: three quant types below IQ1\_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [46cbf0e](https://github.com/unslothai/llama.cpp/pull/61/commits/46cbf0e95786fe8f5b7c0e86d57aaf8f8eceea7f)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - model: add GLM-5-Next (GLM-5.3-Flash) ([#&#8203;27754](https://github.com/ggml-org/llama.cpp/pull/27754), commit [86ebfef](https://github.com/ggml-org/llama.cpp/pull/27754/commits/86ebfef2c6a0f3359a2a07d2c215d61b0fa885c9)) - llama: batched readahead for lazily read gather tables ([unslothai/llama.cpp#137](https://github.com/unslothai/llama.cpp/pull/137), commit [4e1865e](https://github.com/unslothai/llama.cpp/pull/137/commits/4e1865e34ec5f6ca39403215c89129c13731be70)) - ggml-cuda: avoid direct ROCm\_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml\_cuda\_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML\_CUDA\_ENABLE\_UNIFIED\_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - MTP for Qwen3.8-Flash-Next ([unslothai/llama.cpp#144](https://github.com/unslothai/llama.cpp/pull/144), commit [ca14269](https://github.com/unslothai/llama.cpp/pull/144/commits/ca1426903fabe9af26cd10c42034cb4bbd2e0e11)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [09ce1a4](https://github.com/unslothai/llama.cpp/pull/176/commits/09ce1a4d2939844e211f7b4d30a296f4c1aed9a8)) ### [`vb11115-mix-a6922cc`](https://github.com/unslothai/llama.cpp/releases/tag/b11115-mix-a6922cc): llama.cpp prebuilt b11115-mix-a6922cc [Compare Source](https://github.com/unslothai/llama.cpp/compare/b11030-mix-5ff778e...b11115-mix-a6922cc) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b11115](https://github.com/ggml-org/llama.cpp/releases/tag/b11115), merged with: - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [3dac51d](https://github.com/ggml-org/llama.cpp/pull/24423/commits/3dac51dc80b1400929266453c6ced6748ccf6862)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [3870aa4](https://github.com/ggml-org/llama.cpp/pull/25731/commits/3870aa48bbed3d5b1ee7cded0de688baa42f2470)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1\_XS, IQ1\_XXS, IQ1\_XXXS: three quant types below IQ1\_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [46cbf0e](https://github.com/unslothai/llama.cpp/pull/61/commits/46cbf0e95786fe8f5b7c0e86d57aaf8f8eceea7f)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - model: add GLM-5-Next (GLM-5.3-Flash) ([#&#8203;27754](https://github.com/ggml-org/llama.cpp/pull/27754), commit [86ebfef](https://github.com/ggml-org/llama.cpp/pull/27754/commits/86ebfef2c6a0f3359a2a07d2c215d61b0fa885c9)) - llama: batched readahead for lazily read gather tables ([unslothai/llama.cpp#137](https://github.com/unslothai/llama.cpp/pull/137), commit [4e1865e](https://github.com/unslothai/llama.cpp/pull/137/commits/4e1865e34ec5f6ca39403215c89129c13731be70)) - ggml-cuda: avoid direct ROCm\_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml\_cuda\_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML\_CUDA\_ENABLE\_UNIFIED\_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - MTP for Qwen3.8-Flash-Next ([unslothai/llama.cpp#144](https://github.com/unslothai/llama.cpp/pull/144), commit [ca14269](https://github.com/unslothai/llama.cpp/pull/144/commits/ca1426903fabe9af26cd10c42034cb4bbd2e0e11)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [09ce1a4](https://github.com/unslothai/llama.cpp/pull/176/commits/09ce1a4d2939844e211f7b4d30a296f4c1aed9a8)) ### [`vb11030-mix-5ff778e`](https://github.com/unslothai/llama.cpp/releases/tag/b11030-mix-5ff778e): llama.cpp prebuilt b11030-mix-5ff778e [Compare Source](https://github.com/unslothai/llama.cpp/compare/b11027-mix-3e83366...b11030-mix-5ff778e) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b11030](https://github.com/ggml-org/llama.cpp/releases/tag/b11030), merged with: - Qwen3.8-Flash-Next MTP Fix (2x faster now) - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [12e0a96](https://github.com/ggml-org/llama.cpp/pull/24423/commits/12e0a9627d02c6395fd4bbf2aadff93d0d46a0e4)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [946fc11](https://github.com/ggml-org/llama.cpp/pull/25731/commits/946fc11d1afb1e6cd316e853f1b23487754a74b9)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1\_XS, IQ1\_XXS, IQ1\_XXXS: three quant types below IQ1\_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [46cbf0e](https://github.com/unslothai/llama.cpp/pull/61/commits/46cbf0e95786fe8f5b7c0e86d57aaf8f8eceea7f)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - model: add GLM-5-Next (GLM-5.3-Flash) ([#&#8203;27754](https://github.com/ggml-org/llama.cpp/pull/27754), commit [86ebfef](https://github.com/ggml-org/llama.cpp/pull/27754/commits/86ebfef2c6a0f3359a2a07d2c215d61b0fa885c9)) - llama: batched readahead for lazily read gather tables ([unslothai/llama.cpp#137](https://github.com/unslothai/llama.cpp/pull/137), commit [4e1865e](https://github.com/unslothai/llama.cpp/pull/137/commits/4e1865e34ec5f6ca39403215c89129c13731be70)) - ggml-cuda: avoid direct ROCm\_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml\_cuda\_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML\_CUDA\_ENABLE\_UNIFIED\_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - MTP for Qwen3.8-Flash-Next ([unslothai/llama.cpp#144](https://github.com/unslothai/llama.cpp/pull/144), commit [ca14269](https://github.com/unslothai/llama.cpp/pull/144/commits/ca1426903fabe9af26cd10c42034cb4bbd2e0e11)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [09ce1a4](https://github.com/unslothai/llama.cpp/pull/176/commits/09ce1a4d2939844e211f7b4d30a296f4c1aed9a8)) ### [`vb11027-mix-3e83366`](https://github.com/unslothai/llama.cpp/releases/tag/b11027-mix-3e83366): llama.cpp prebuilt b11027-mix-3e83366 [Compare Source](https://github.com/unslothai/llama.cpp/compare/b11007-mix-3e83366...b11027-mix-3e83366) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b11027](https://github.com/ggml-org/llama.cpp/releases/tag/b11027), merged with: - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [12e0a96](https://github.com/ggml-org/llama.cpp/pull/24423/commits/12e0a9627d02c6395fd4bbf2aadff93d0d46a0e4)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [946fc11](https://github.com/ggml-org/llama.cpp/pull/25731/commits/946fc11d1afb1e6cd316e853f1b23487754a74b9)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1\_XS, IQ1\_XXS, IQ1\_XXXS: three quant types below IQ1\_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [46cbf0e](https://github.com/unslothai/llama.cpp/pull/61/commits/46cbf0e95786fe8f5b7c0e86d57aaf8f8eceea7f)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - model: add GLM-5-Next (GLM-5.3-Flash) ([#&#8203;27754](https://github.com/ggml-org/llama.cpp/pull/27754), commit [86ebfef](https://github.com/ggml-org/llama.cpp/pull/27754/commits/86ebfef2c6a0f3359a2a07d2c215d61b0fa885c9)) - llama: batched readahead for lazily read gather tables ([unslothai/llama.cpp#137](https://github.com/unslothai/llama.cpp/pull/137), commit [4e1865e](https://github.com/unslothai/llama.cpp/pull/137/commits/4e1865e34ec5f6ca39403215c89129c13731be70)) - ggml-cuda: avoid direct ROCm\_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml\_cuda\_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML\_CUDA\_ENABLE\_UNIFIED\_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - MTP for Qwen3.8-Flash-Next ([unslothai/llama.cpp#144](https://github.com/unslothai/llama.cpp/pull/144), commit [f40f64a](https://github.com/unslothai/llama.cpp/pull/144/commits/f40f64a81cf4d9f8c4539920828ad9d20f4ba7c5)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [09ce1a4](https://github.com/unslothai/llama.cpp/pull/176/commits/09ce1a4d2939844e211f7b4d30a296f4c1aed9a8)) ### [`vb11007-mix-3e83366`](https://github.com/unslothai/llama.cpp/releases/tag/b11007-mix-3e83366): llama.cpp prebuilt b11007-mix-3e83366 [Compare Source](https://github.com/unslothai/llama.cpp/compare/b10995-mix-3e83366...b11007-mix-3e83366) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b11007](https://github.com/ggml-org/llama.cpp/releases/tag/b11007), merged with: - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [12e0a96](https://github.com/ggml-org/llama.cpp/pull/24423/commits/12e0a9627d02c6395fd4bbf2aadff93d0d46a0e4)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [946fc11](https://github.com/ggml-org/llama.cpp/pull/25731/commits/946fc11d1afb1e6cd316e853f1b23487754a74b9)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1\_XS, IQ1\_XXS, IQ1\_XXXS: three quant types below IQ1\_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [46cbf0e](https://github.com/unslothai/llama.cpp/pull/61/commits/46cbf0e95786fe8f5b7c0e86d57aaf8f8eceea7f)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - model: add GLM-5-Next (GLM-5.3-Flash) ([#&#8203;27754](https://github.com/ggml-org/llama.cpp/pull/27754), commit [86ebfef](https://github.com/ggml-org/llama.cpp/pull/27754/commits/86ebfef2c6a0f3359a2a07d2c215d61b0fa885c9)) - llama: batched readahead for lazily read gather tables ([unslothai/llama.cpp#137](https://github.com/unslothai/llama.cpp/pull/137), commit [4e1865e](https://github.com/unslothai/llama.cpp/pull/137/commits/4e1865e34ec5f6ca39403215c89129c13731be70)) - ggml-cuda: avoid direct ROCm\_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml\_cuda\_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML\_CUDA\_ENABLE\_UNIFIED\_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - MTP for Qwen3.8-Flash-Next ([unslothai/llama.cpp#144](https://github.com/unslothai/llama.cpp/pull/144), commit [f40f64a](https://github.com/unslothai/llama.cpp/pull/144/commits/f40f64a81cf4d9f8c4539920828ad9d20f4ba7c5)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [09ce1a4](https://github.com/unslothai/llama.cpp/pull/176/commits/09ce1a4d2939844e211f7b4d30a296f4c1aed9a8)) ### [`vb10995-mix-3e83366`](https://github.com/unslothai/llama.cpp/releases/tag/b10995-mix-3e83366): llama.cpp prebuilt b10995-mix-3e83366 [Compare Source](https://github.com/unslothai/llama.cpp/compare/b10909-mix-bea84f7...b10995-mix-3e83366) Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream [b10995](https://github.com/ggml-org/llama.cpp/releases/tag/b10995), merged with: - DiffusionGemma ([#&#8203;24423](https://github.com/ggml-org/llama.cpp/pull/24423), commit [12e0a96](https://github.com/ggml-org/llama.cpp/pull/24423/commits/12e0a9627d02c6395fd4bbf2aadff93d0d46a0e4)) - Add TML Inkling architecture ([#&#8203;25731](https://github.com/ggml-org/llama.cpp/pull/25731), commit [946fc11](https://github.com/ggml-org/llama.cpp/pull/25731/commits/946fc11d1afb1e6cd316e853f1b23487754a74b9)) - kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes ([unslothai/llama.cpp#70](https://github.com/unslothai/llama.cpp/pull/70), commit [883f2c9](https://github.com/unslothai/llama.cpp/pull/70/commits/883f2c9ba78f3847148454adf025da29385fff3e)) - IQ1\_XS, IQ1\_XXS, IQ1\_XXXS: three quant types below IQ1\_S ([unslothai/llama.cpp#61](https://github.com/unslothai/llama.cpp/pull/61), commit [46cbf0e](https://github.com/unslothai/llama.cpp/pull/61/commits/46cbf0e95786fe8f5b7c0e86d57aaf8f8eceea7f)) - sampling: index penalties by token id instead of scanning every candidate ([unslothai/llama.cpp#95](https://github.com/unslothai/llama.cpp/pull/95), commit [3db8cb5](https://github.com/unslothai/llama.cpp/pull/95/commits/3db8cb5b2e9bf291057b9f19960e8601a162da81)) - model: add GLM-5-Next (GLM-5.3-Flash) ([#&#8203;27754](https://github.com/ggml-org/llama.cpp/pull/27754), commit [86ebfef](https://github.com/ggml-org/llama.cpp/pull/27754/commits/86ebfef2c6a0f3359a2a07d2c215d61b0fa885c9)) - llama: batched readahead for lazily read gather tables ([unslothai/llama.cpp#137](https://github.com/unslothai/llama.cpp/pull/137), commit [4e1865e](https://github.com/unslothai/llama.cpp/pull/137/commits/4e1865e34ec5f6ca39403215c89129c13731be70)) - ggml-cuda: avoid direct ROCm\_Host compute on HIP integrated GPUs (port of [ggml-org#25863](https://github.com/ggml-org/llama.cpp/issues/25863)) ([unslothai/llama.cpp#158](https://github.com/unslothai/llama.cpp/pull/158), commit [abfc45b](https://github.com/unslothai/llama.cpp/pull/158/commits/abfc45b9cb21eae4848cb82196e659f42c9a8341)) - ggml-cuda: use cudaMemcpyDefault in the ggml\_cuda\_cpy 2D fast path ([unslothai/llama.cpp#157](https://github.com/unslothai/llama.cpp/pull/157), commit [6c6da89](https://github.com/unslothai/llama.cpp/pull/157/commits/6c6da89266ba7839d825c9997782af4f4d26b81b)) - ggml-cuda: make GGML\_CUDA\_ENABLE\_UNIFIED\_MEMORY=0 actually disable it, and say so on HIP ([unslothai/llama.cpp#149](https://github.com/unslothai/llama.cpp/pull/149), commit [b65a2dc](https://github.com/unslothai/llama.cpp/pull/149/commits/b65a2dce12c14a489e19a059cb6ee59112f1b733)) - MTP for Qwen3.8-Flash-Next ([unslothai/llama.cpp#144](https://github.com/unslothai/llama.cpp/pull/144), commit [f40f64a](https://github.com/unslothai/llama.cpp/pull/144/commits/f40f64a81cf4d9f8c4539920828ad9d20f4ba7c5)) - llama: map each contiguous run of a context's tensors, not one span over all of them ([unslothai/llama.cpp#152](https://github.com/unslothai/llama.cpp/pull/152), commit [b2b5ed9](https://github.com/unslothai/llama.cpp/pull/152/commits/b2b5ed9ff86427a530b762a45d3fdbd453bcd4e8)) - mtmd: test that every projector is registered and uniquely named ([unslothai/llama.cpp#176](https://github.com/unslothai/llama.cpp/pull/176), commit [09ce1a4](https://github.com/unslothai/llama.cpp/pull/176/commits/09ce1a4d2939844e211f7b4d30a296f4c1aed9a8)) </details> --- ### Configuration 📅 **Schedule**: (UTC) - Branch creation - At any time (no schedule defined) - Automerge - At any time (no schedule defined) 🚦 **Automerge**: Disabled by config. Please merge this manually once you are satisfied. ♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the rebase/retry checkbox. 🔕 **Ignore**: Close this PR and you won't be reminded about this update again. --- - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this box --- This PR has been generated by [Mend Renovate CLI](https://github.com/renovatebot/renovate). <!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0NC4xMDYuMCIsInVwZGF0ZWRJblZlciI6IjQ0LjEwNi4wIiwidGFyZ2V0QnJhbmNoIjoibWFpbiIsImxhYmVscyI6WyJyZW5vdmF0ZSJdfQ==-->
Update Unsloth llama.cpp prebuilt to v11160
All checks were successful
Build and publish / build (vulkan) (pull_request) Successful in 12s
Build and publish / build (rocm-gfx1151) (pull_request) Successful in 1m39s
d8623d5a05
brian merged commit c8ac6865c8 into main 2026-09-30 07:30:26 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brian/llama-unsloth-container!12
No description provided.