github
GitHub APIExtracts: repo, release, stars delta · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public API; token recommended for production rate limits.
Last ingested: 2026-08-25 · run status: success with data
- 2026-07-16ggerganov/llama.cpp b10052: b10052
<details open> hexagon: L2 cache handling rework (dirty bit tracking with lazy flushing) and more MUL_MAT updates (#25762) * hex-mm: fix artificial limit in the solver that restricted number of act-prep threads * hex-mm: fix warning * hex-prof: do not apply --top to the timel
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10051: b10051
<details open> kleidiai: Add SME vs SME2 distinction in kernel dispatch (#25478) The current integration treats SME as a single capability (CPU_FEATURE_SME) with no distinction between SME(v1) and SME2. The kernels dispatched under CPU_FEATURE_SME use SME2-specific instructions
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10050: b10050
<details open> vulkan: when using transfer queue for async copies, sync on event_wait to avoid race (#25229) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10050/llama-b1005
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10048: b10048
<details open> TP: fix Phi3, Bert, Plamo2/3, ChatGLM (#25536) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10048/llama-b10048-bin-macos-arm64.tar.gz) - macOS Apple Silicon
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10047: b10047
<details open> vendor: update BoringSSL to 0.20260713.0 (#25624) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10047/llama-b10047-bin-macos-arm64.tar.gz) - macOS Apple Sili
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10046: b10046
<details open> tests: actually exercise `test-recurrent-state-rollback` (#25758) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10046/llama-b10046-bin-macos-arm64.tar.gz) -
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10045: b10045
<details open> server : allow text-only slot save/restore with mtmd (#25076) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10045/llama-b10045-bin-macos-arm64.tar.gz) - macO
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10043: b10043
<details open> CUDA: Support CUDA Virtual Devices (#25228) * support cuda virtual devices * disable NCCL path when virtual devices are used * label virtual devices in description; add GPUx2 server CI jobs * code refactor </details> **Website:** - <https://llama.app> **mac
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10042: b10042
<details open> Enable CUDA graphs on volta+turing (#25749) </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10042/llama-b10042-bin-macos-arm64.tar.gz) - macOS Apple Silicon (a
github:ggerganov/llama.cpp # Patch release v5.14.1 This patch solves a few issues which appeared when integrating Inkling model, most notably an issue affecting models using EncoderDecoderCache during assisted generation. It also fixes an issue that could appear during prefill with StaticCache and sdpa
github:huggingface/transformers- 2026-07-16ggerganov/llama.cpp b10038: b10038
<details open> ci : add official website link to release notes (#25728) Assisted-by: pi:llama.cpp/Qwen3.6-27B </details> **Website:** - <https://llama.app> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10038/llama-b10
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10037: b10037
<details open> quant : allow using manual tensor types with --pure (#25716) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10037/llama-b10037-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enab
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10036: b10036
<details open> opencl: disable FA and MoE weights repack to work around compiler issues for Adreno 850 GPU (#25745) * opencl: workaround for A850 compiler compat * opencl: fix DX compiler version parsing and cleanup --------- Co-authored-by: Li He <lih@qti.qualcomm.com> </d
github:ggerganov/llama.cpp - 2026-07-16ggerganov/llama.cpp b10035: b10035
<details open> cuda: extract Q1_0 elements via __byte_perm (#25628) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10035/llama-b10035-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DI
github:ggerganov/llama.cpp # ROCm Core SDK 7.14.0 release notes ROCm Core SDK 7.14.0 transitions ROCm to [TheRock](https://github.com/ROCm/TheRock), a build and release system that introduces a modular architecture to improve flexibility, maintainability, and alignment with community use cases: * **L
AMDgithub:ROCm/ROCm