github
GitHub APIExtracts: repo, release, stars delta · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public API; token recommended for production rate limits.
Last ingested: 2026-08-25 · run status: success with data
- 2026-08-13ggerganov/llama.cpp b10415: b10415
<details open> spec : auto-detect mtp draft model type (#27005) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10415/llama-b10415-bin-macos-arm64.tar.gz) - macOS Apple Silic
github:ggerganov/llama.cpp - 2026-08-13ggerganov/llama.cpp b10414: b10414
<details open> metal : add TQ2_0 support (#26980) * metal: add TQ2_0 support Add support for the GGML_TYPE_TQ2_0 (ternary, 2 bits per element) type in the Metal backend. Assisted-by: llama.cpp:DeepSeek-v4-Flash-0731 * cont : optimize mul_mv kernel - float ops over integer o
github:ggerganov/llama.cpp - 2026-08-13ggerganov/llama.cpp b10413: b10413
<details open> common : auto-detect spec type from draft GGUF metadata (#26814) * common : auto-detect spec type from draft GGUF metadata When -md loads a local draft model without --spec-type, the sidecar inference in common_models_handler_apply only checks HF repo sidecars a
github:ggerganov/llama.cpp - 2026-08-13NVIDIA/cutlass v4.7.0: CUTLASS 4.7.0
### CuTe DSL * New features: - Introduced the Primitives API which provides a lower-level abstraction beneath CuTe enabling Tensor Core programming through SIMT. This provides a stable, thin wrapper over NVVM operations to use where CuTe abstractions reduce development veloc
NVDAgithub:NVIDIA/cutlass - 2026-08-13ggerganov/llama.cpp b10400: b10400
<details open> ggml : fix arm builds, unused var (#26991) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10400/llama-b10400-bin-macos-arm64.tar.gz) - macOS Apple Silicon (ar
github:ggerganov/llama.cpp - 2026-08-12ggerganov/llama.cpp b10375: b10375
<details open> chat : tighten bare function parsing for Qwen models (#26793) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10375/llama-b10375-bin-macos-arm64.tar.gz) - macO
github:ggerganov/llama.cpp - 2026-08-12NVIDIA/cutlass 4.7.0: CUTLASS 4.7.0
### CuTe DSL * New features: - Introduced the Primitives API which provides a lower-level abstraction beneath CuTe enabling Tensor Core programming through SIMT. This provides a stable, thin wrapper over NVVM operations to use where CuTe abstractions reduce development veloc
NVDAgithub:NVIDIA/cutlass - 2026-08-12ggerganov/llama.cpp b10373: b10373
<details open> imatrix.cpp: Move finite check and only check touched experts (#26861) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10373/llama-b10373-bin-macos-arm64.tar.g
github:ggerganov/llama.cpp - 2026-08-12ggerganov/llama.cpp b10369: b10369
<details open> mtmd: support pocket-tts (#26871) * adapt the api * text model ok * working impl, need verify and clean up * mtmd: build the pocket-tts transposed convolutions as GEMM + col2im ggml_conv_transpose_1d has no grouped mode, so the depthwise upsample was built as
github:ggerganov/llama.cpp - 2026-08-11ggerganov/llama.cpp b10362: b10362
<details open> tests : disable backend sampler hip multi output (#26878) * test-backend-sampler: skip multi_output_sampling_chain on HIP The new multi_output_sampling_chain test uses top_k, whose backend probs path needs CUB (unavailable on HIP), so sampled_probs is null and t
github:ggerganov/llama.cpp - 2026-08-11ggerganov/llama.cpp b10361: b10361
<details open> model : fix SWA not being enabled for EXAONE 4.5 (#26848) * model : fix SWA not being enabled for EXAONE 4.5 load_arch_hparams tests `hparams.n_layer() == 64` before LLM_KV_NEXTN_PREDICT_LAYERS has been read. n_layer() returns n_layer_all - n_layer_nextn and n_l
github:ggerganov/llama.cpp - 2026-08-11ggerganov/llama.cpp b10360: b10360
<details open> common/peg : suppress incomplete escape sequences (#26780) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10360/llama-b10360-bin-macos-arm64.tar.gz) - macOS A
github:ggerganov/llama.cpp - 2026-08-11ggerganov/llama.cpp b10359: b10359
<details open> ggml-webgpu: fix CI errors from #25025 and #25262 (#26566) * test new flash_attn test * rebase and fix to disable subgrou matrices when max_kv_tile == 0 * delete log output * Add i32 support to cpy and enables the all ops test * restore the non target ci test
github:ggerganov/llama.cpp - 2026-08-11vllm-project/vllm v0.27.1: v0.27.1
This is a patch release on top of v0.27.0. - Support quantized DSpark Markov heads (#50424)
github:vllm-project/vllm - 2026-08-11ggerganov/llama.cpp b10358: b10358
<details open> Address review comment of PR 25532 (#26852) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10358/llama-b10358-bin-macos-arm64.tar.gz) - macOS Apple Silicon (a
github:ggerganov/llama.cpp - 2026-08-11ggerganov/llama.cpp b10357: b10357
<details open> opencl: transpose the K tile in local memory for FA prefill kernels (#26428) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10357/llama-b10357-bin-macos-arm64
github:ggerganov/llama.cpp - 2026-08-11ggerganov/llama.cpp b10356: b10356
<details open> ci : target ROCm 7.14 for build and release (#25775) * Switch ROCm from 7.2.1 to 7.14 ROCm 7.14 is the first production release using TheRock build system. It can be installed using multi-arch deliverables from wheels, debs, rpms, tarballs or runfiles. Adjust R
github:ggerganov/llama.cpp - 2026-08-10ggerganov/llama.cpp b10355: b10355
<details open> llama : support multi-output backend sampling (#25532) * Enable backend sampling with token speculation * Clamp the mask sum before converting it into the sampled index * Add a numeric context parameter declaring the maximum outputs one sequence * More fixes
github:ggerganov/llama.cpp - 2026-08-10ggerganov/llama.cpp b10354: b10354
<details open> ggml-cpu : fix CPU affinity mask being ignored on Android (#26838) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10354/llama-b10354-bin-macos-arm64.tar.gz) -
github:ggerganov/llama.cpp - 2026-08-10ggerganov/llama.cpp b10353: b10353
<details open> ggml : require contiguous src for ROLL on CUDA and Metal (#25928) ggml_roll only asserts nb[0] == ggml_type_size, so a permuted src is a valid input, but the CUDA and Metal roll kernels index by ne alone and never read the nb strides. A non-contiguous src therefo
github:ggerganov/llama.cpp - 2026-08-10vllm-project/vllm v0.27.0: v0.27.0
# vLLM v0.27.0 Release Notes ## Highlights This release features 561 commits from 242 contributors (64 new)! * **Kimi K3 support** with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRe
github:vllm-project/vllm - 2026-08-10ggerganov/llama.cpp b10344: b10344
<details open> model: add MTP support for Nemotron model (#26725) * model: add MTP support for Nemotron Nano model * model: add mtp_flags for nemotron model * address review comments </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64
github:ggerganov/llama.cpp - 2026-08-10ggerganov/llama.cpp b10343: b10343
<details open> vendor : update cpp-httplib to 0.53.0 (#26821) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10343/llama-b10343-bin-macos-arm64.tar.gz) - macOS Apple Silicon
github:ggerganov/llama.cpp - 2026-08-10ggerganov/llama.cpp b10342: b10342
<details open> model : Granite-Switch Architecture (#25107) * granite-switch: add llama.cpp backend (POC, CPU) New "granite-switch" architecture: a dense, all-attention Granite-4.1 model with N embedded LoRA adapters selected per-token by control tokens. - gguf-py schema (arc
github:ggerganov/llama.cpp - 2026-08-10ggerganov/llama.cpp b10338: b10338
<details open> model-saver : fix expert shared/chunk FFN length key clobber (#26693) The saver called add_kv with LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH twice, the second time passing n_ff_chexp. gguf_set_val_u32 removes-then-appends, so the second call clobbers the first: th
github:ggerganov/llama.cpp # Release v5.15.0 ## New Model additions ### Meta Muse Glimmer Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it can be deployed to
github:huggingface/transformers- 2026-08-10ggerganov/llama.cpp b10336: b10336
<details open> ggml-webgpu : refactor several wgsl files and simplify flash_attn wgsl. (#26134) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10336/llama-b10336-bin-macos-a
github:ggerganov/llama.cpp - 2026-08-09ggerganov/llama.cpp b10333: b10333
<details open> ggml-cpu : fix missing Q5_0 dispatch in SpaceMiT backend (#26792) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10333/llama-b10333-bin-macos-arm64.tar.gz) -
github:ggerganov/llama.cpp - 2026-08-09ggerganov/llama.cpp b10332: b10332
<details open> ci: rm `GGML_HIP_ROCWMMA_FATTN` (#26760) Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10332/llama-b10332-bin-
github:ggerganov/llama.cpp - 2026-08-08ggerganov/llama.cpp b10331: b10331
<details open> server: report the isolate working directory from get_info (#26773) * server: report the isolate working directory from get_info Without an explicit cwd, get_info fell back to the server process working directory even when a tools runtime was configured. That na
github:ggerganov/llama.cpp - 2026-08-08ggerganov/llama.cpp b10330: b10330
<details open> CUDA: fuse rms_norm + mul + rope (+ view + set_rows) (#26767) * CUDA: fuse rms_norm + mul + rope (+ view + set_rows) * tests: add broadcast weight case to rms_norm_mul_rope * CUDA: check memory ranges before rms_norm rope fusion * CUDA: check memory ranges in
github:ggerganov/llama.cpp - 2026-08-08ggerganov/llama.cpp b10329: b10329
<details open> server, ui: only offer a working directory when a tool reads it (#26762) The working directory chip showed up as soon as the server exposed any builtin tool, so a server started with just get_datetime, or a user who turned every filesystem tool off in the setting
github:ggerganov/llama.cpp - 2026-08-08ggerganov/llama.cpp b10328: b10328
<details open> server: add initial tool isolation support (via docker) (#26507) * server: add initial tool isolation support (via docker) * add docs * adapt get_info * py: fix type check * cont * separate tools_io_sandbox / tools_io_docker * rename sandbox --> isolate *
github:ggerganov/llama.cpp - 2026-08-08NVIDIA/cutlass v4.6.2: CUTLASS 4.6.2
### CuTe DSL * Bug fixing and improvements - Reverted the TMA bulk copy elect_one change from 4.6.0 to reset the behavior to align with 4.5.x releases - Fixed a vectorized fp32->f8 conversion issue ([!3382](https://github.com/NVIDIA/cutlass/issues/3382)) - Fixed a CuTe
NVDAgithub:NVIDIA/cutlass - 2026-08-08ggerganov/llama.cpp b10327: b10327
<details open> CUDA: fix thread/block count in quantized cpy kernel launches (#26731) * CUDA: fix thread/block count in quantized cpy kernel launches * tests: add uneven block count cpy case </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon
github:ggerganov/llama.cpp - 2026-08-07ggerganov/llama.cpp b10326: b10326
<details open> tts: account for the vocoder pass in the timings line (#26733) get_output runs the waveform work the pipeline defers to it, from a single trailing window to a full pass depending on the model. Measuring it keeps the reported total and the audio to process ratio h
github:ggerganov/llama.cpp - 2026-08-07ggerganov/llama.cpp b10322: b10322
<details open> sycl: coalesce the ssm_conv window loads (#26612) test-backend-ops perf -o SSM_CONV on an Arc Pro B70, interleaved A/B against master, 6 reps, us/run: ne_a=[515,3328,1,1] ne_b=[4,3328,1,1] n_t=512 97.68 -> 52.95 1.85x ne_a=[937,8192,1,1] ne_b=[4,8192
github:ggerganov/llama.cpp - 2026-08-07ggerganov/llama.cpp b10321: b10321
<details open> metal : fix NORM/RMS_NORM for row lengths that leave a partial simdgroup (#26708) ggml_metal_op_norm sized the threadgroup with `nth = std::min(nth, args.ne00_t)`, which can leave nth not a multiple of the simdgroup size. The kernels finish their row reduction wi
github:ggerganov/llama.cpp - 2026-08-07ggerganov/llama.cpp b10319: b10319
<details open> mtmd: fix longest_edge ignoring min/max pixels (#26638) * mtmd: fix longest_edge ignoring min/max pixels * nits </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/downloa
github:ggerganov/llama.cpp - 2026-08-07ggerganov/llama.cpp b10318: b10318
<details open> sync : ggml </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10318/llama-b10318-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLE
github:ggerganov/llama.cpp - 2026-08-07ggerganov/llama.cpp b10313: b10313
<details open> server: (router) add LRU scheduler (#26572) * add lru_sched * handle coalescing (req leaves waiting queue) * add tests * fix stream case * address review comments </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](
github:ggerganov/llama.cpp - 2026-08-07ggerganov/llama.cpp b10312: b10312
<details open> server: (router) do not evict busy models (#26567) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10312/llama-b10312-bin-macos-arm64.tar.gz) - macOS Apple Sil
github:ggerganov/llama.cpp - 2026-08-07ggerganov/llama.cpp b10311: b10311
<details open> mtmd: stop feeding the text stream again during Qwen3-TTS generation (#26706) The reference implementation has two mutually exclusive prompt layouts. In non streaming mode the prefill carries the whole utterance text plus tts_eos summed with codec_pad, and the tr
github:ggerganov/llama.cpp - 2026-08-07ggerganov/llama.cpp b10310: b10310
<details open> ggml : add aarch64 HWCAP fallbacks and fix fp16 variant detection (#25554) * ggml : add fallback definitions for missing aarch64 HWCAP bits * ggml : require HWCAP_ASIMDHP for the aarch64 fp16 cpu variants Also rename has_fp16_va to has_fp16, the field gates the
github:ggerganov/llama.cpp - 2026-08-06ggerganov/llama.cpp b10298: b10298
<details open> mtmd: add chunk save/load function (#26645) * mtmd: add chunk save/load function * nits * add tests * rn _MAX --> _COUNT </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/relea
github:ggerganov/llama.cpp - 2026-08-06ggerganov/llama.cpp b10297: b10297
<details open> server: fix empty response for /cors-proxy (#26656) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10297/llama-b10297-bin-macos-arm64.tar.gz) - macOS Apple Si
github:ggerganov/llama.cpp - 2026-08-06ggerganov/llama.cpp b10295: b10295
<details open> model-loader : fix quantized reshaped tensor strides (#26672) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10295/llama-b10295-bin-macos-arm64.tar.gz) - macO
github:ggerganov/llama.cpp - 2026-08-06ggerganov/llama.cpp b10293: b10293
<details open> ci : onboard AMD ROCm CI with gfx1151 fixes (#26544) * ci: prepare for amd rocm ci Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * ci: fix editorconfig-checker Signed-off-by: Aaron Teo <aaron.teo1@ibm.com> * ci: fix device not recognised Signed-off-by: Aaron
github:ggerganov/llama.cpp - 2026-08-06ggerganov/llama.cpp b10291: b10291
<details open> vulkan: fix submission batching size, add debug tools for diagnosing causes of DeviceLost drivers errors (#26371) * vulkan: add debug tooling to get more information about a DeviceLost error * fix submission threshold applied too late * use logging macros, thro
github:ggerganov/llama.cpp - 2026-08-06ggerganov/llama.cpp b10290: b10290
<details open> mtmd/ggml: add ggml_build_forward_order (#26649) * ggml: add ggml_build_forward_order ggml_build_forward_expand marks the tensor and all its ancestors for compute, so using it as a pure ordering hint (keeping q, k and v together) defeats ggml_build_forward_selec
github:ggerganov/llama.cpp