Extracts: post title, subreddit, score · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public endpoints; reuse and sustained commercial access are terms-restricted.
Last ingested: 2026-08-26 · run status: success with data
- 2026-07-04Appreciation post!
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unn4j3/appreciation_post/"> <img alt="Appreciation post!" src="https://preview.redd.it/xokoffoerabh1.png?width=640&crop=smart&auto=webp&s=f2b1fafaf2adf5fb4935393cf2497f090bca41da" title="Appreciat
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>We moved our agent fleet's working memory off Markdown and onto TOON (Token-Oriented Object Notation) in December 2025 and just wrote up what 14 harnesses taught us.</p> <p>The honest numbers (tiktoken o200k, 100 uniform CRM records):</p> <p>- TO
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unif51/possible_evidence_of_literal_prompt_injection_by/"> <img alt="possible evidence of literal prompt injection by anthropic" src="https://external-preview.redd.it/72by0vodw29h1.png?width=640&crop=smar
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>I'm using the model with opencode and the issue is it's looping hard when reasoning. It's not a deranged babbling though, the reasoning is legit large spans of text, it just can't get outside of the loop and make a decision. So I babysit it, stop
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Is this a doable build? I know that it is possible to overflow models from the memory, so I'm wondering if a q4 or q6 quantization could work with this set up. Anyone running anything similar, with memory offload of models?</p> </div><!-- SC_ON -
reddit:LocalLLaMA- 2026-07-04Gemma 4 12B - MLX Kernel
<!-- SC_OFF --><div class="md"><p>I've mentioned this kernel project I was working on in a few posts and figured I would just open the project code for anyone curious: <a href="https://github.com/jscott3201/Helios">MLX Gemma 12B</a></p> <p>The main constraints for this on my end
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Check it out: <a href="https://github.com/fairydreaming/llama.cpp/tree/dsv4">https://github.com/fairydreaming/llama.cpp/tree/dsv4</a></p> <p>They are PRs <a href="https://github.com/ggml-org/llama.cpp/pull/25247">#25247</a>, <a href="https://gith
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Found cheap one RTX5000 16GB is it worth to add to RTX3060? And is it difficult to make them work together in LM Studio? Do I need to use studio drivers and is RTX5000 still supported with driver updates? Or should I just go for 5060Ti instead? T
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unbn4y/comparing_local_inference_speeds_across_a_few/"> <img alt="Comparing local inference speeds across a few real setups people are running (3090 vs 5090 vs dual 6000)" src="https://preview.redd.it/8aep4zi
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unbm45/ran_a_classicmedival_europe_fantasy_rpagentic/"> <img alt="Ran a classic(medival europe) fantasy RP/agentic benchmark across 8 local models Qwen3.6-27B held up better than its size suggests" src="https
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unbjr2/using_local_models_with_hermes_vs_claude_code/"> <img alt="Using local models with Hermes vs Claude code" src="https://preview.redd.it/ii7l8k4wa8bh1.png?width=640&crop=smart&auto=webp&s=350
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unbi4a/qwen36_27b_on_a_5090_64k_sample_toks_distribution/"> <img alt="Qwen3.6 27B on a 5090, 6.4k sample tok/s distribution after tuning MTP/cache settings" src="https://preview.redd.it/l2k7gu5cb8bh1.jpeg?wid
reddit:LocalLLaMA- 2026-07-04DGX Spark and Overtemps
<!-- SC_OFF --><div class="md"><p>For anyone who has a DGX-Spark and is having problems during these very hot summer months, you can underclock with:</p> <p>sudo nvidia-smi -lgc 0,900</p> <p>My temps dropped from 85C to 60C and this fixed my problem of overtemp lockups.</p> <p>Ed
reddit:LocalLLaMA - 2026-07-04Feedback needed
<!-- SC_OFF --><div class="md"><p>Dear LocalLLaMA,</p> <p>I’m looking for feedback on my project because I noticed it is mainly catered towards technical people.</p> <p><a href="https://github.com/heterodoxin/apostate">Apostate</a> (Apostate is an abliteration engine built from s
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>A lot of people seem to be confused or mystified about this so figured I'd spell it out.</p> <p>I played around with RYS and realized that it broke Gemma 4 models. Turns out there's a `layer_scalar` value that is applied at each layer. If you don
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Here's the setup I decided on for embedding gemma4-12b into a Tauri2 desktop app:</p> <ul> <li>Native Rust FFI into llama.cpp via <code>llama-cpp-2</code> (Metal enabled)</li> <li>Model:<code>gemma-4-12b-it-Q5_K_S</code> quantized by Unsloth, Q5_
reddit:LocalLLaMA- 2026-07-04First time I have seen this: my model seemed aware of its context usage ask me for compaction!
<!-- SC_OFF --><div class="md"><p>I was in a middle of a Claude Code session with GLM 5.2. Context usage 537k/1M. After finishing a task, GLM asked me this:</p> <blockquote> <p><strong>Context</strong> <strong>note:</strong> this session has run long and context is getting heavy.
reddit:LocalLLaMA <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1un9955/paper_gear_guided_endtoend_autoregression_for/"> <img alt="[Paper] GEAR: Guided End-to-End AutoRegression for Image Synthesis" src="https://preview.redd.it/vo0a7q0ut7bh1.png?width=640&crop=smart&am
reddit:LocalLLaMA- 2026-07-04Using structured insurance/captives to solve public pushback & zoning delays? (Idea discussion)
<!-- SC_OFF --><div class="md"><p>I’m looking for some blunt feedback on a strategy to handle the increasing community/municipal pushback on new builds—specifically regarding power grid strain, water use, and the "what's in it for the community" argument.<br /> Right no
reddit:datacenter - 2026-07-04Questions about AWS data center roles.
<!-- SC_OFF --><div class="md"><p>Hello everyone, I'm currently working at a data center on the operations side of things and am looking to move over to the compute side, something I can't do with my current employer. </p> <p>I'm wondering what are the differences between an Infr
reddit:datacenter