github
GitHub APIExtracts: repo, release, stars delta · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public API; token recommended for production rate limits.
Last ingested: 2026-08-25 · run status: success with data
- 2026-07-15ggerganov/llama.cpp b10034: b10034
<details open> opencl: exclude some moe kernels on Adreno a7x (#25698) * opencl: exclude Adreno A7x from using Adreno MoE kernels Some compilers for A7x devices miscompile the repack kernels, corrupting the weights and causing MoE models to generate garbage output * opencl: e
github:ggerganov/llama.cpp - 2026-07-15ggerganov/llama.cpp b10032: b10032
<details open> cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) (#25545) * cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel) * chore : remove indentation of #pragma unroll * cuda : remove unnec
github:ggerganov/llama.cpp - 2026-07-15ggerganov/llama.cpp b10031: b10031
<details open> tokenize : drop --stdin mutual-exclusion check (#25672) match cli and completion, which don't enforce it </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10031/llama-b10031-bin-macos-arm64.tar.gz
github:ggerganov/llama.cpp # Release v5.14.0 ## New Model additions ### Inkling (fresh from Thinking Machines): 975B total, 41B active * Add Inkling model #47347 by @molbap @Cyrilvallez @eustlb and @zucchini-nlp <img width="3840" height="2160" alt="image" src="https://github.com/user-attachmen
github:huggingface/transformers- 2026-07-15ggerganov/llama.cpp b10025: b10025
<details open> cuda : relax tensor contiguity requirements for quantized concat (#25678) * cuda : relax tensor contiguity requirements for quantized concat * tests : add test cases for non-contiguous quantized concat * ggml : relax contiguity requirements for quantized concat
github:ggerganov/llama.cpp - 2026-07-15ggerganov/llama.cpp b10021: b10021
<details open> DeepseekV4: reduce graph splits (#25702) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10021/llama-b10021-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED](http
github:ggerganov/llama.cpp - 2026-07-15ggerganov/llama.cpp b10020: b10020
<details open> sycl : fix get_rows Q2_K, Q4_K, Q5_K (#25656) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10020/llama-b10020-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED]
github:ggerganov/llama.cpp - 2026-07-15NVIDIA/cutlass v4.6.1: CUTLASS 4.6.1
### CuTe DSL * Bug fixing and improvements - Fixed following issues: - https://github.com/NVIDIA/cutlass/issues/3243 - https://github.com/NVIDIA/cutlass/issues/3359 - https://github.com/NVIDIA/cutlass/issues/3365 - https://github.com/NVIDIA/cutlass/issue
NVDAgithub:NVIDIA/cutlass - 2026-07-15ggerganov/llama.cpp b10015: b10015
<details open> opencl: do not use `clCreateBufferWithProperties` when targeting CL 2.x (#25673) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10015/llama-b10015-bin-macos-arm64.tar.gz) - macOS Apple Silicon (
github:ggerganov/llama.cpp