Extracts: post title, subreddit, score · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public endpoints; reuse and sustained commercial access are terms-restricted.
Last ingested: 2026-08-26 · run status: success with data
<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1uogpo4/us_navy_is_flighttesting_3d_printed_fighter_jet/"> <img alt="US Navy is flight-testing 3D printed fighter jet parts that cut repair times in half — forward-deployed 3D printers generate composite parts,
reddit:hardware- 2026-07-05Floating gates are the only analog compute element you can fab today with zero extra mask steps
<!-- SC_OFF --><div class="md"><p>Some might say I'm obsessed with this whole analog computing thing. But its only because I find it so interesting! For some quick background I wrote a previous introductory explainer about how computing some things in the analog domain can be far
reddit:semiconductors - 2026-07-05Graviton G5 STREAM memory bandwidth
<!-- SC_OFF --><div class="md"><p>STREAM memory bandwidth results on an c9g.48xlarge instance.</p> <p>For reference my own HPC code (memory bandwidth limited with most time spent doing parallel sparse linear algebra) ran about 40% faster compared to the last generation Graviton (
reddit:hardware <!-- SC_OFF --><div class="md"><p>I have an old Xeon rig with 512Gb of 4-channel DDR4 2133 memory and E5-2699v4 processor. For GPU I have GTX 1060 with 6Gb of VRAM, so I use CPU only mode. I can run GLM 5.2 with 40B active parameters in Q4_K_XL at 1.8 t/s, but as you can understa
reddit:LocalLLaMA- 2026-07-05A planetary test for local models
<!-- SC_OFF --><div class="md"><p>Here is a fun test prompt:</p> <p>Imagine a date in the next 1000 years where the Sun, along with its gravity, suddenly disappeared. When that happens, all planets in our solar system would stop orbiting and carry in a straight line. Is there a d
reddit:LocalLLaMA - 2026-07-055060 worth it?
<!-- SC_OFF --><div class="md"><p>I had built a pc in 2023 for pretty cheap, it has a RTX 4090 and intel i9 13900K. I’m thinking of adding a 5060Ti for additional LLM workloads. I primarily use my GPU for training SLMs and llama.cpp inference.<br /> Since the memory bandwidth of
reddit:LocalLLaMA - 2026-07-05I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads
<!-- SC_OFF --><div class="md"><h1>I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads — prefill dominates everything, and KV head count beats parameter count</h1> <p>I've been running local LLMs for agentic workflows (tool use, cod
reddit:LocalLLaMA - 2026-07-05Concurrency plus nvfp4 on Blackwell
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unqkjy/concurrency_plus_nvfp4_on_blackwell/"> <img alt="Concurrency plus nvfp4 on Blackwell" src="https://preview.redd.it/phylna3ajbbh1.png?width=140&height=127&auto=webp&s=9dabf768fd1b49bfd934f3c
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Not really a tutorial, but more of sharing my attempts at getting higher contexts on Q8 of Qwen3.6-27 with 32GB VRAM.</p> <p><strong>Disclaimer</strong>: Not in-depth research. Crowd wisdom suggests that Qwen is more tolerant of model quantizatio
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unobl4/using_applications_to_make_a_smaller_model_more/"> <img alt="Using "applications" to make a smaller model more effective at bigger tasks." src="https://external-preview.redd.it/a2trYXYycGQxYm
reddit:LocalLLaMA