github
GitHub APIExtracts: repo, release, stars delta · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public API; token recommended for production rate limits.
Last ingested: 2026-08-26 · run status: success with data
- 2026-07-31ggerganov/llama.cpp b10215: b10215
<details open> vulkan: Introduce driver version check for Windows Intel GPU to mitigate crashing (#25192) * Removed crash guard for Intel Crash fixed from driver 32.0.101.8860 * Added driver version check for windows * Change to convert from driverVersion rather than string
github:ggerganov/llama.cpp - 2026-07-31ggerganov/llama.cpp b10214: b10214
<details open> mtmd: add n_embd_head (#26342) Co-authored-by: Daniel Han <unslothai@gmail.com> </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10214/llama-b10214-bin-macos-a
github:ggerganov/llama.cpp - 2026-07-31ggerganov/llama.cpp b10213: b10213
<details open> Support rotated kv cache quant (#26180) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10213/llama-b10213-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64
github:ggerganov/llama.cpp - 2026-07-31ggerganov/llama.cpp b10212: b10212
<details open> llama : load MTP tensors only if they are really used (#26296) * llama : load MTP tensors only if they are really used * llama : skip loading MTP (if not used) in remaining models that support MTP --------- Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.co
github:ggerganov/llama.cpp - 2026-07-31ggerganov/llama.cpp b10211: b10211
<details open> vulkan: update vulkan sdk to 1.4.357.0 (#26303) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10211/llama-b10211-bin-macos-arm64.tar.gz) - macOS Apple Silico
github:ggerganov/llama.cpp - 2026-07-31ggerganov/llama.cpp b10210: b10210
<details open> server: correct accepted tokens when need draft token replay (#26320) * spec: correct accepted tokens when need draft token replay * cont : naming --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> </details> **Website:** - <https://llama.app>
github:ggerganov/llama.cpp - 2026-07-31ggerganov/llama.cpp b10209: b10209
<details open> cuda: extract Q2_0 elements via __byte_perm (#25603) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10209/llama-b10209-bin-macos-arm64.tar.gz) - macOS Apple S
github:ggerganov/llama.cpp - 2026-07-31ggerganov/llama.cpp b10208: b10208
<details open> SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt proc… (#25025) * SYCL: add oneMKL GEMM flash attention for XMX-accelerated prompt processing * fattn-mkl: fix interleaved dst layout in normalize kernel - Fix mkl_fa_normalize_head: use interleave
github:ggerganov/llama.cpp - 2026-07-30ggerganov/llama.cpp b10186: b10186
<details open> ggml : Fix issue with kleidiai ci and stringop overflow warning (#26277) Signed-off-by: Jonathan Clohessy <Jonathan.Clohessy@arm.com> </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama
github:ggerganov/llama.cpp - 2026-07-30ggerganov/llama.cpp b10184: b10184
<details open> mimo2: address MTP review feedback (#26228) Co-authored-by: tnhnyc <115956684+tnhnyc@users.noreply.github.com> </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/
github:ggerganov/llama.cpp - 2026-07-29ggerganov/llama.cpp b10182: b10182
<details open> llama: move suppress_tokens handling to common/sampling (#26276) * llama: move suppress_tokens handling to common/sampling * address security issues * rm has_logit_bias </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm6
github:ggerganov/llama.cpp - 2026-07-29ggerganov/llama.cpp b10181: b10181
<details open> ggml-cuda : disable MMQ on devices with less than 48 KiB shared memory (#26141) ggml_cuda_should_use_mmq() selects MMQ purely from the quantization type. The current MMQ configurations are designed and maintained against a minimum of 48 KiB per-block shared memor
github:ggerganov/llama.cpp - 2026-07-29ggerganov/llama.cpp b10180: b10180
<details open> sycl: contiguous fast path + 32-bit index math for unary elementwise ops (#25946) * sycl: contiguous fast path + 32-bit index math for unary elementwise ops * sycl: use fastdiv for elementwise index math </details> **Website:** - <https://llama.app> **macOS/i
github:ggerganov/llama.cpp - 2026-07-29ggerganov/llama.cpp b10179: b10179
<details open> vendor: update BoringSSL to 0.20260728.0 (#26241) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10179/llama-b10179-bin-macos-arm64.tar.gz) - macOS Apple Sili
github:ggerganov/llama.cpp - 2026-07-29ggerganov/llama.cpp b10178: b10178
<details open> server : add trace logging for slot similarity checking (#26271) Adds trace logging in server-context.cpp for slot similarity checking during prompt cache slot selection, including skip reasons and similarity calculation details. Assisted-by: llama.cpp:Qwen3.6-2
github:ggerganov/llama.cpp - 2026-07-29ggerganov/llama.cpp b10176: b10176
<details open> RPC: add tensor_memset (#25912) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10176/llama-b10176-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, Kleidi
github:ggerganov/llama.cpp - 2026-07-29ggerganov/llama.cpp b10175: b10175
<details open> add rdna3.5, and 3 to mmq configs so they can be tuned independently. (#26199) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10175/llama-b10175-bin-macos-arm
github:ggerganov/llama.cpp - 2026-07-29ggerganov/llama.cpp b10174: b10174
<details open> model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2) (#25980) * model: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2) Adds GLM-5.2 NextN/MTP as a --spec-type draft-mtp target: nextn tensor loading via the qwen35moe/step35-sty
github:ggerganov/llama.cpp - 2026-07-28ggerganov/llama.cpp b10173: b10173
<details open> model: Add Laguna-S-2.1 LLM_TYPE (#26233) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10173/llama-b10173-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm
github:ggerganov/llama.cpp - 2026-07-28ggerganov/llama.cpp b10172: b10172
<details open> ggml-webgpu: Fix some binding alias issues to support all archs, fix recurrent-state-rollback test (#25931) * Add overlap glu variant to support all archs, fix recurrent-state-rollback test * format * Fix all arch overlapped ranges * format * diagnose bus err
github:ggerganov/llama.cpp - 2026-07-28ggerganov/llama.cpp b10171: b10171
<details open> opencl: skip the Adreno KQ/KQV image kernels for multi-stream batches (#26189) The Adreno KQ/KQV image1d kernels (ggml_cl_mul_mat_kq_kqv_adreno) ignore dim 3 entirely: the sub-buffer covers only nb02*ne02 bytes and the kernel receives no ne03/ne13/nb03/nb13 argum
github:ggerganov/llama.cpp - 2026-07-28ggerganov/llama.cpp b10167: b10167
<details open> server: abstract llama_memory calls to common_memory (#26221) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10167/llama-b10167-bin-macos-arm64.tar.gz) - macO
github:ggerganov/llama.cpp - 2026-07-28ggerganov/llama.cpp b10166: b10166
<details open> ggml : set output of view src (#25729) * llama-graph: set_outputs to t->view_src * change set_output to GGML_ASSERT about views not being outputs * sampler : avoid views in outputs * cont : fix dist sampler * cont : consistent logits handling * ggml : set ou
github:ggerganov/llama.cpp - 2026-07-28ggerganov/llama.cpp b10165: b10165
<details open> vulkan: add iq4_nl support back to FA (#24585) * vulkan: add iq4_nl support back to FA I was originally concerned about wasting shared memory on the LUT, but it's small and unlikely to matter in practice. Also support q1_0 for non-coopmat2. Fixes #23681 * rem
github:ggerganov/llama.cpp - 2026-07-28ggerganov/llama.cpp b10164: b10164
<details open> ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration (#22675) * ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration * cuda: added SSD CICD fixes for CUDA / HIP / MUSA / MSVC. * ggml-cuda: review comments fixed. * ggml-cuda: Fuse M
github:ggerganov/llama.cpp - 2026-07-28ggerganov/llama.cpp b10159: b10159
<details open> ggml-metal: FWHT kernel for metal backend (#25924) * metal fwht wip * shape guard and formatting * formatting * Formatting and typos Co-authored-by: YiChen Lv <63285796+forforever73@users.noreply.github.com> * fix narrowing issue Co-authored-by: YiChen Lv <
github:ggerganov/llama.cpp - 2026-07-28ggerganov/llama.cpp b10156: b10156
<details open> Disable -ffast-math on HIP (#25495) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10156/llama-b10156-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, Kl
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10155: b10155
<details open> mtmd: support MiMo-V2.5 audio input (RVQ-based model) (#26190) * gguf converter for mimo audio * fix conv * cpp impl * nits * nits 2 </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/ll
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10154: b10154
<details open> common : add common_print_available_devices() (#26170) Signed-off-by: Adrien Gallouët <angt@huggingface.co> </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10153: b10153
<details open> model: Add support for Nanbeige4.2 (#25994) * support nanbeige4.2 model * fix * fix flake8 Lint check * fix loop bound check and drop redundant head_dim --------- Co-authored-by: root <lizongqiang@kanzhun.com> </details> **Website:** - <https://llama.app>
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10152: b10152
<details open> fit : count nextn (MTP) blocks in n_gpu_layers so front layers stay on GPU (#26177) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10152/llama-b10152-bin-maco
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10151: b10151
<details open> sycl(build): parallelize ocloc invocations (#25903) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10151/llama-b10151-bin-macos-arm64.tar.gz) - macOS Apple Si
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10150: b10150
<details open> ggml : adjust logic for offloading ops to weight's backend (#25832) * ggml : adjust logic for offloading ops to weight's backend * llama : dsv4 graph fixes </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://gi
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10149: b10149
<details open> tests : remove unnecessary sync in test-save-load-state (#26166) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10149/llama-b10149-bin-macos-arm64.tar.gz) - m
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10148: b10148
<details open> common: fix explicit -md precedence over draft sidecar resolution (#26165) * common: fix explicit -md precedence over draft sidecar resolution Follow-up of #25955, an explicit --model-draft file given with -hfd was silently overridden by the sidecar resolution o
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10146: b10146
<details open> ggml-cpu: Enable BF16 tiled gemm optimization on PowerPC (#26068) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10146/llama-b10146-bin-macos-arm64.tar.gz) -
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10144: b10144
<details open> server + ui: fix stream routes for model names containing a slash (#26137) * server + ui: refactor resumable stream routes to query string conv_id The conversation id can embed a model name containing slashes (ggml-org/...) in router mode, which the decoded path
github:ggerganov/llama.cpp - 2026-07-27ggerganov/llama.cpp b10142: b10142
<details open> mtmd: Add Vision Support for Minimax-M3 (#25113) * Add preliminary MiniMax-M3 support Text-only port that re-uses existing components: MiniMax-M2 style GQA with per-head QK-norm and partial rotary, DeepSeek-V3 style leading-dense and routed/shared experts, and s
github:ggerganov/llama.cpp - 2026-07-26ggerganov/llama.cpp b10141: b10141
<details open> mtmd: fix android build (#26150) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10141/llama-b10141-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, Kleid
github:ggerganov/llama.cpp - 2026-07-25vllm-project/vllm v0.26.0: v0.26.0
# vLLM v0.26.0 Release Notes ## Highlights This release features 411 commits from 212 contributors (61 new)! * **New Inkling model family** with a full support stack: base modeling (#48799), piecewise CUDA graph support (#48822), Hopper FA4 relative attention (#48858), M
github:vllm-project/vllm - 2026-07-24ggerganov/llama.cpp b10107: b10107
<details open> hexagon: fix Windows crash when op_poll is enabled (#26029) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10107/llama-b10107-bin-macos-arm64.tar.gz) - macOS
github:ggerganov/llama.cpp - 2026-07-24ggerganov/llama.cpp b10106: b10106
<details open> CUDA: fix external compilation of q1_0 MMQ (#25778) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10106/llama-b10106-bin-macos-arm64.tar.gz) - macOS Apple Si
github:ggerganov/llama.cpp - 2026-07-24ggerganov/llama.cpp b10105: b10105
<details open> args: refactor mlock/mmap/directio into load-mode (#20834) * args: overhaul mmap/mlock/dio into single arg Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * docs: update docs with llama-gen-docs Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * chore: satisfy cod
github:ggerganov/llama.cpp - 2026-07-24ggerganov/llama.cpp b10103: b10103
<details open> metal : add f16 type support to leaky relu (#25981) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10103/llama-b10103-bin-macos-arm64.tar.gz) - macOS Apple Si
github:ggerganov/llama.cpp - 2026-07-23ggerganov/llama.cpp b10099: b10099
<details open> CUDA: Improve NVFP4 W4A4 activation quantization (#25730) * Squash history before conflict-resolution during rebase on master WIP commit Add 32-byte loads, restore per-block amax Use nvfp4x4 intrinsic when available Fuse per-channel amax and quantization kern
github:ggerganov/llama.cpp - 2026-07-23ggerganov/llama.cpp b10098: b10098
<details open> hexagon: activation ops update (#25974) * hex-geglu: optimized all-in-one geglu microkernel * hex-geglu: enable non-contiguous src and strided DMA * hex-act: enable non-contiguous srs and strided DMA for rest of ACT ops * hex-act: generalize GLU per-thread fun
github:ggerganov/llama.cpp - 2026-07-23ggerganov/llama.cpp b10094: b10094
<details open> common: infer the speculative type from the draft repo sidecars (#25989) With -hfd pointing to a repo that ships mtp-/dflash-/eagle3- sidecars and no --spec-type given, the draft resolved to a full model while the sidecar was the intended draft. When the specula
github:ggerganov/llama.cpp - 2026-07-23ggerganov/llama.cpp b10093: b10093
<details open> Fix DeepSeek4 crafted template (#25414) * chat: fix DS4 template to explicitly follow reference behavior * Support DeepSeekv4 flag (`drop_reasoning`). * fix: hook DS3.2 parser for DS4 as well * fix: add tool result reordering * fix: post-merge </details> **
github:ggerganov/llama.cpp - 2026-07-23ggerganov/llama.cpp b10092: b10092
<details open> ggml: enable PowerPC backend variants on AIX (#25983) * ggml: enable PowerPC backend variants on AIX Allow the PowerPC CPU backend variants to be built on AIX by extending the platform check in the CMake configuration. This reuses the existing PowerPC backend im
github:ggerganov/llama.cpp - 2026-07-22ggerganov/llama.cpp b10091: b10091
<details open> ci : fix SYCL package shared library lookup (#25987) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10091/llama-b10091-bin-macos-arm64.tar.gz) - macOS Apple S
github:ggerganov/llama.cpp