github
GitHub APIExtracts: repo, release, stars delta · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public API; token recommended for production rate limits.
Last ingested: 2026-08-26 · run status: success with data
- 2026-07-15NVIDIA/cutlass v4.6.1: CUTLASS 4.6.1
### CuTe DSL * Bug fixing and improvements - Fixed following issues: - https://github.com/NVIDIA/cutlass/issues/3243 - https://github.com/NVIDIA/cutlass/issues/3359 - https://github.com/NVIDIA/cutlass/issues/3365 - https://github.com/NVIDIA/cutlass/issue
NVDAgithub:NVIDIA/cutlass - 2026-07-15ggerganov/llama.cpp b10015: b10015
<details open> opencl: do not use `clCreateBufferWithProperties` when targeting CL 2.x (#25673) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10015/llama-b10015-bin-macos-arm64.tar.gz) - macOS Apple Silicon (
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b10012: b10012
<details open> hexagon: fix hmx-queue signal enum-narrowing problem (#25677) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10012/llama-b10012-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI ena
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b10011: b10011
<details open> server : refactor prompt cache state ownership (#25649) * server : clear checkpoints upon prompt clear * server : move the prompt state data to the server_prompt_cache Assisted-by: pi:llama.cpp/Qwen3.6-27B * server : handle batched slot being cleared </detail
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b10010: b10010
<details open> server: add --cors-* options (#25655) * server: add --cors-* options * add special "localhost" value * add tests * fix test * add link to PR </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b1
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b10007: b10007
<details open> opencl: fix a dp4a bug for devices where cl_khr_integer_dot_product is unavailable (#25639) * opencl: do not fail backend init on devices without cl_khr_integer_dot_product * opencl: do not call dp4 kernels when dp is unavailable --------- Co-authored-by: Li H
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b10005: b10005
<details open> DeepseekV4: fix seq_rm (#25588) * DeepseekV4: fix seq_rm * implement proper seq_cp * create actual update context </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b10005/llama-b10005-bin-macos-a
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b9999: b9999
<details open> kleidiai : add SME2 f32 kernel (#24414) * kleidiai : add SME2 f32 kernel * enable dynamic scheduling for SME2 f32 kernel </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9999/llama-b9999-bin-mac
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b9996: b9996
<details open> arg: Flush log before exiting after usage() (#25504) Under certain conditions, it's possible for messages emitted via LOG() to get lost before exit, apparently because they are emitted by another thread. common_params_print_usage() uses printf directly, and is no
github:ggerganov/llama.cpp - 2026-07-14ggerganov/llama.cpp b9995: b9995
<details open> sycl: set fattn_vec_nthreads to 256 for Battlemage (#25205) Currently detects lunarlake + battlemage / xe2 and sets the value to 256. Keeps default at 128, Intel's ARC Alchemist's prefered value. </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https:
github:ggerganov/llama.cpp - 2026-07-14vllm-project/vllm v0.25.1: v0.25.1
# vLLM v0.25.1 ## Highlights This release features 2 commits from 2 contributors (1 new)! v0.25.1 is a patch release containing two targeted bug fixes on top of v0.25.0. ### Bug Fixes * **Avoid blocking model launching when no system FFmpeg is available for TorchCode
github:vllm-project/vllm - 2026-07-14ggerganov/llama.cpp b9994: b9994
<details open> metal : add Q2_0 support (#25419) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9994/llama-b9994-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED](https://githu
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9993: b9993
<details open> model: add Hy3 (hy_v3) support with MTP speculative decoding (#25395) * model: add Hy3 (hy_v3) architecture support Adds Tencent Hunyuan 3 (HF architecture HYV3ForCausalLM, GGUF arch hy_v3): a MoE decoder stack with per-head Q/K RMSNorm, a sigmoid router with ex
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9992: b9992
<details open> CUDA: refactor MMQ kernel configuration (#24127) * CUDA: refactor MMQ kernel configuration * fix Blackwell config * remove legacy code </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9992/llam
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9990: b9990
<details open> spec: add Minimax2 eagle3 support * Fix nullptr in minimax2 EAGLE3 * minor : add newline --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/
github:ggerganov/llama.cpp - 2026-07-13NVIDIA/cutlass v4.6.0: CUTLASS 4.6.0
* Release [documentation](https://docs.nvidia.com/cutlass/latest/media/docs/cpp/gemm_performance_measurement_methodology_guidelines.md) that explains how to accurately profiling GEMM performance. ### CuTe DSL * New features - New fine-grained compilation API: cute.compile_
NVDAgithub:NVIDIA/cutlass - 2026-07-13ggerganov/llama.cpp b9988: b9988
<details open> tests: Harmonize header use (#25616) * tests: Harmonize the use of private ggml includes * tests: In test-backend-ops, use quoted includes As with all other tests. This is to ensure that the build uses shipped headers over possibly system-installed ones. </det
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9987: b9987
<details open> gguf : add tensor shape accessor (#24405) * gguf : add tensor shape accessors * gguf : return tensor shape as const int64_t * * gguf : remove n_dims accessor, keep only gguf_get_tensor_ne </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://githu
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9986: b9986
<details open> chat : fix reasoning leak with force-opened bare <think> templates (#24674) * chat : fix reasoning leak with force-opened bare <think> templates The reasoning start tag inferred from prior turns can carry trailing whitespace (e.g. <think>\n) while a force-open t
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9985: b9985
<details open> sycl: add fused top-k MoE (#25217) * sycl: add fused top-k MoE * sycl: address review: GGML_SYCL_ENABLE_FUSION env, move fusion dispatch to topk-moe * sycl: print GGML_SYCL_ENABLE_FUSION at startup like other env vars Co-Authored-By: Claude Fable 5 <noreply@an
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9984: b9984
<details open> sycl: add Q2_K to DMMV reorder path (#25064) Signed-off-by: Todd Malsbary <todd.malsbary@intel.com> </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9984/llama-b9984-bin-macos-arm64.tar.gz) - mac
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9982: b9982
<details open> server: honour per-request reasoning_budget_tokens in chat completions (#23116) * server: honour per-request reasoning_budget_tokens in chat completions The reasoning-budget block in oaicompat_chat_params_parse read only the server-level default (opt.reasoning_b
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9981: b9981
<details open> vendor : update cpp-httplib to 0.50.1 (#25576) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9981/llama-b9981-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED](
github:ggerganov/llama.cpp - 2026-07-13ggerganov/llama.cpp b9980: b9980
<details open> server: Don't consider models with --no-mmproj-auto as multimodal (#25590) If mmproj is explicitly disabled via the model preset or command-line parameters then the model won't be able to handle image/audio inputs and this shouldn't be declared as supported input
github:ggerganov/llama.cpp - 2026-07-12ggerganov/llama.cpp b9979: b9979
<details open> mtmd: fix silent prompt truncation on embedded NUL (#25548) * mtmd: fix silent prompt truncation on embedded NUL mtmd_input_text carried the prompt as a bare const char* with no length, so a NUL byte in message content cut the prompt at the tokenizer boundary an
github:ggerganov/llama.cpp - 2026-07-12ggerganov/llama.cpp b9978: b9978
<details open> server : evict checkpoints within min-step of each other (#25472) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9978/llama-b9978-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI e
github:ggerganov/llama.cpp - 2026-07-12ggerganov/llama.cpp b9977: b9977
<details open> server : fix image blocks in tool_result being dropped during Anthropic OpenAI conversion (#22536) * server : fix image blocks in tool_result being dropped during Anthropic→OpenAI conversion server_chat_convert_anthropic_to_oai() silently discarded image blocks
github:ggerganov/llama.cpp - 2026-07-12ggerganov/llama.cpp b9976: b9976
<details open> Fix conditional to display 'LLAMA_SPLIT_MODE_TENSOR not implemented for architecture' message (#24926) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9976/llama-b9976-bin-macos-arm64.tar.gz) - m
github:ggerganov/llama.cpp - 2026-07-12ggerganov/llama.cpp b9975: b9975
<details open> gguf : reject empty metadata keys (#24917) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9975/llama-b9975-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED](http
github:ggerganov/llama.cpp - 2026-07-12ggerganov/llama.cpp b9969: b9969
<details open> Vulkan: route large matmuls to medium tile on Adreno (#24877) * [Vulkan] Fixes llama-cli breaking over longer promts sizes The llama-cli was breaking for longer promts sizes for q4_0 quantized networks. Causing due to insufficient shared memory. * Removed the u
github:ggerganov/llama.cpp - 2026-07-12ggerganov/llama.cpp b9968: b9968
<details open> opencl: add int8 dp4 dense and MoE prefill optimization for Adreno GPUs (#25537) * opencl: add int8 dp4 dense and moe GEMM * opencl: refactor --------- Co-authored-by: Li He <lih@qti.qualcomm.com> </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](htt
github:ggerganov/llama.cpp - 2026-07-12ggerganov/llama.cpp b9967: b9967
<details open> server: accept null sampling params (#25538) * server: accept null sampling params Extend the schema validation to treat a null value as absent, so clients can send null on nullable params (temperature, top_p, ...) to request the server default. This matches the
github:ggerganov/llama.cpp - 2026-07-11vllm-project/vllm v0.25.0: v0.25.0
# vLLM v0.25.0 Release Notes ## Highlights This release features 558 commits from 232 contributors (64 new)! * **Model Runner V2 is now the default for all dense models** (#44443). Building on quantized-model support from the previous release, MRv2 is now the standard ex
github:vllm-project/vllm - 2026-07-11ggerganov/llama.cpp b9966: b9966
<details open> llama : make tensor-split regex patterns static (#24710) llama_meta_device_get_split_state() recompiled 29 std::regex on every call. In -sm tensor mode the callback runs once per tensor per token, so this dominated the decode thread in profiling. Mark them static
github:ggerganov/llama.cpp - 2026-07-11ggerganov/llama.cpp b9965: b9965
<details open> hexagon: improve ARGSORT performance for small tensors (#25512) * hex-sort: add efficient bitomic sort in hvx regs up to 1024 elements * hex-sort: fix inverted vrors * hex-sort: specialize sort functions for the common cases * hex-sort: add tracing and local c
github:ggerganov/llama.cpp - 2026-07-11ggerganov/llama.cpp b9964: b9964
<details open> arg: prevent duplicate spec model downloads (#25527) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9964/llama-b9964-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISA
github:ggerganov/llama.cpp # Patch release v5.13.1 This patch is focused on enabling `transformers` for the latest release of vllm! - Be more defensive with remap_legacy_layer_types for custom models (#47245) from @hmellor - Fix custom code which doesn't know about the new linear layer type names
github:huggingface/transformers- 2026-07-11ggerganov/llama.cpp b9963: b9963
<details open> mtmd: deepseek-ocr v1 multi-tile (#24717) * mtmd: deepseek-ocr v1 multi-tile dynamic resolution + unified image-preprocessors for both versions (ds-ocr v1 and v2) * remove hacky API * fuse row into a long image * almost working * adapt to new preprocessor api
github:ggerganov/llama.cpp - 2026-07-11ggerganov/llama.cpp b9960: b9960
<details open> server: remove loading.html (#25500) * server: remove loading.html * apply ui changes </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9960/llama-b9960-bin-macos-arm64.tar.gz) - macOS Apple Sili
github:ggerganov/llama.cpp - 2026-07-11ggerganov/llama.cpp b9959: b9959
<details open> sync : ggml </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9959/llama-b9959-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED](https://github.com/ggml-org/llama.c
github:ggerganov/llama.cpp - 2026-07-10ggerganov/llama.cpp b9957: b9957
<details open> server: improve tools, remove apply_diff (#25498) * server: improve tools, remove apply_diff * improve edit tool * add tools_io abstraction * add tools_io_basic * fix build * move utils to class member * add const </details> **macOS/iOS:** - [macOS Apple
github:ggerganov/llama.cpp - 2026-07-10ggerganov/llama.cpp b9956: b9956
<details open> cli: fix crash on wrong server base url (#25497) * llama-cli: fix crash on wrong server base url by catching exceptions and graceful exit * review: leaner catch group: json error and standard exception </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](
github:ggerganov/llama.cpp - 2026-07-10ggerganov/llama.cpp b9952: b9952
<details open> llama : make all KQ masks f16 if FA is used, remove zero attention bias, remove raw_k repeats in DeepSeek V4 (#25370) * llama : make all KQ masks (except the lightning indexer one) f16 if FA is used and remove zero attention bias in DeepSeek V4 * llama : remove
github:ggerganov/llama.cpp - 2026-07-10ggerganov/llama.cpp b9951: b9951
<details open> ggml-et: Initial ET backend (#24179) * ggml-et: Add performance logging * ggml-et: Quants helpers * ggml-et: Add MUL_MAT kernel * ggml-et: Add ROPE kernel * ggml-et: Add RMS_NORM kernel * ggml-et: Add GLU kernel * ggml-et: Add SOFT_MAX kernel * ggml-et: A
github:ggerganov/llama.cpp - 2026-07-10ggerganov/llama.cpp b9950: b9950
<details open> llama-batch: add unit test (#25471) * llama-batch: add unit test * fix win32 builds * add not implemented assertion in unused methods * remove unreachable code </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/re
github:ggerganov/llama.cpp - 2026-07-10ggerganov/llama.cpp b9949: b9949
<details open> opencl: cluster-parallel decode FA for Adreno (#25473) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9949/llama-b9949-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DI
github:ggerganov/llama.cpp - 2026-07-10ggerganov/llama.cpp b9948: b9948
<details open> ggml : process data in smaller chunks in CUDA ggml_top_k() and ggml_argsort() to reduce temporary buffers memory usage (#24776) * ggml : process data in smaller chunks in CUDA ggml_top_k() implementation to reduce temporary buffers memory usage * ggml : allocate
github:ggerganov/llama.cpp - 2026-07-10ggerganov/llama.cpp b9947: b9947
<details open> cli: add --output option (#25484) </details> **macOS/iOS:** - [macOS Apple Silicon (arm64)](https://github.com/ggml-org/llama.cpp/releases/download/b9947/llama-b9947-bin-macos-arm64.tar.gz) - macOS Apple Silicon (arm64, KleidiAI enabled) [DISABLED](https://githu
github:ggerganov/llama.cpp - 2026-07-09ggerganov/llama.cpp b9946: b9946
<details open> hexagon: tiling, tracing and optimizations for unary ops (#25474) * hexagon: tile wide rows in pointwise unary ops to avoid VTCM overflow * unary: reject permuted tensors for now (not used by models) * hex-unary: replace divs with fastdiv * hex-unary: add vtcm
github:ggerganov/llama.cpp - 2026-07-09ggerganov/llama.cpp b9945: b9945
<details open> server : move chat-template thinking probe inside the init try/catch (#24093) A model whose chat template parses at init but fails parser generation at apply time (e.g. uses {% call %}) throws std::invalid_argument from common_chat_templates_support_enable_thinki
github:ggerganov/llama.cpp