github
GitHub APIExtracts: repo, release, stars delta · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public API; token recommended for production rate limits.
Last ingested: 2026-08-25 · run status: success with data
- 2026-07-14ggerganov/llama.cpp b10012: b10012
<details open> hexagon: fix hmx-queue signal enum-narrowing problem (#25677) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10012/llama-b10012-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI ena
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b10011: b10011
<details open> server : refactor prompt cache state ownership (#25649) * server : clear checkpoints upon prompt clear * server : move the prompt state data to the server_prompt_cache Assisted-by: pi:llama.cpp/Qwen3.6-27B * server : handle batched slot being cleared </detail
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b10010: b10010
<details open> server: add --cors-* options (#25655) * server: add --cors-* options * add special "localhost" value * add tests * fix test * add link to PR </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b1
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b10007: b10007
<details open> opencl: fix a dp4a bug for devices where cl_khr_integer_dot_product is unavailable (#25639) * opencl: do not fail backend init on devices without cl_khr_integer_dot_product * opencl: do not call dp4 kernels when dp is unavailable --------- Co-authored-by: Li H
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b10005: b10005
<details open> DeepseekV4: fix seq_rm (#25588) * DeepseekV4: fix seq_rm * implement proper seq_cp * create actual update context </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10005/llama-b10005-bin-macos-a
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b9999: b9999
<details open> kleidiai : add SME2 f32 kernel (#24414) * kleidiai : add SME2 f32 kernel * enable dynamic scheduling for SME2 f32 kernel </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9999/llama-b9999-bin-mac
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b9996: b9996
<details open> arg: Flush log before exiting after usage() (#25504) Under certain conditions, it's possible for messages emitted via LOG() to get lost before exit, apparently because they are emitted by another thread. common_params_print_usage() uses printf directly, and is no
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b9995: b9995
<details open> sycl: set fattn_vec_nthreads to 256 for Battlemage (#25205) Currently detects lunarlake + battlemage / xe2 and sets the value to 256. Keeps default at 128, Intel's ARC Alchemist's prefered value. </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https:
github:ggerganov/llama.cpp - 2026-07-14vllm-project/vllm v0.25.1: v0.25.1
# vLLM v0.25.1 ## Highlights This release features 2 commits from 2 contributors (1 new)! v0.25.1 is a patch release containing two targeted bug fixes on top of v0.25.0. ### Bug Fixes * **Avoid blocking model launching when no system FFmpeg is available for TorchCode
github:vllm-project/vllm - 2026-07-14ggerganov/llama.cpp b9994: b9994
<details open> metal : add Q2_0 support (#25419) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9994/llama-b9994-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED](https://githu
github:ggerganov/llama.cpp