Extracts: post title, subreddit, score · metadata or bounded excerpt
Retention: D1 event history with canonical source link and deduplication metadata.
Access basis: Public endpoints; reuse and sustained commercial access are terms-restricted.
Last ingested: 2026-08-26 · run status: success with data
<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1uogpo4/us_navy_is_flighttesting_3d_printed_fighter_jet/"> <img alt="US Navy is flight-testing 3D printed fighter jet parts that cut repair times in half — forward-deployed 3D printers generate composite parts,
reddit:hardware- 2026-07-05Floating gates are the only analog compute element you can fab today with zero extra mask steps
<!-- SC_OFF --><div class="md"><p>Some might say I'm obsessed with this whole analog computing thing. But its only because I find it so interesting! For some quick background I wrote a previous introductory explainer about how computing some things in the analog domain can be far
reddit:semiconductors - 2026-07-05Graviton G5 STREAM memory bandwidth
<!-- SC_OFF --><div class="md"><p>STREAM memory bandwidth results on an c9g.48xlarge instance.</p> <p>For reference my own HPC code (memory bandwidth limited with most time spent doing parallel sparse linear algebra) ran about 40% faster compared to the last generation Graviton (
reddit:hardware <!-- SC_OFF --><div class="md"><p>I have an old Xeon rig with 512Gb of 4-channel DDR4 2133 memory and E5-2699v4 processor. For GPU I have GTX 1060 with 6Gb of VRAM, so I use CPU only mode. I can run GLM 5.2 with 40B active parameters in Q4_K_XL at 1.8 t/s, but as you can understa
reddit:LocalLLaMA- 2026-07-05A planetary test for local models
<!-- SC_OFF --><div class="md"><p>Here is a fun test prompt:</p> <p>Imagine a date in the next 1000 years where the Sun, along with its gravity, suddenly disappeared. When that happens, all planets in our solar system would stop orbiting and carry in a straight line. Is there a d
reddit:LocalLLaMA - 2026-07-055060 worth it?
<!-- SC_OFF --><div class="md"><p>I had built a pc in 2023 for pretty cheap, it has a RTX 4090 and intel i9 13900K. I’m thinking of adding a 5060Ti for additional LLM workloads. I primarily use my GPU for training SLMs and llama.cpp inference.<br /> Since the memory bandwidth of
reddit:LocalLLaMA - 2026-07-05I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads
<!-- SC_OFF --><div class="md"><h1>I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads — prefill dominates everything, and KV head count beats parameter count</h1> <p>I've been running local LLMs for agentic workflows (tool use, cod
reddit:LocalLLaMA - 2026-07-05Concurrency plus nvfp4 on Blackwell
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unqkjy/concurrency_plus_nvfp4_on_blackwell/"> <img alt="Concurrency plus nvfp4 on Blackwell" src="https://preview.redd.it/phylna3ajbbh1.png?width=140&height=127&auto=webp&s=9dabf768fd1b49bfd934f3c
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Not really a tutorial, but more of sharing my attempts at getting higher contexts on Q8 of Qwen3.6-27 with 32GB VRAM.</p> <p><strong>Disclaimer</strong>: Not in-depth research. Crowd wisdom suggests that Qwen is more tolerant of model quantizatio
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unobl4/using_applications_to_make_a_smaller_model_more/"> <img alt="Using "applications" to make a smaller model more effective at bigger tasks." src="https://external-preview.redd.it/a2trYXYycGQxYm
reddit:LocalLLaMA- 2026-07-04Appreciation post!
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unn4j3/appreciation_post/"> <img alt="Appreciation post!" src="https://preview.redd.it/xokoffoerabh1.png?width=640&crop=smart&auto=webp&s=f2b1fafaf2adf5fb4935393cf2497f090bca41da" title="Appreciat
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>We moved our agent fleet's working memory off Markdown and onto TOON (Token-Oriented Object Notation) in December 2025 and just wrote up what 14 harnesses taught us.</p> <p>The honest numbers (tiktoken o200k, 100 uniform CRM records):</p> <p>- TO
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unif51/possible_evidence_of_literal_prompt_injection_by/"> <img alt="possible evidence of literal prompt injection by anthropic" src="https://external-preview.redd.it/72by0vodw29h1.png?width=640&crop=smar
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>I'm using the model with opencode and the issue is it's looping hard when reasoning. It's not a deranged babbling though, the reasoning is legit large spans of text, it just can't get outside of the loop and make a decision. So I babysit it, stop
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Is this a doable build? I know that it is possible to overflow models from the memory, so I'm wondering if a q4 or q6 quantization could work with this set up. Anyone running anything similar, with memory offload of models?</p> </div><!-- SC_ON -
reddit:LocalLLaMA- 2026-07-04Gemma 4 12B - MLX Kernel
<!-- SC_OFF --><div class="md"><p>I've mentioned this kernel project I was working on in a few posts and figured I would just open the project code for anyone curious: <a href="https://github.com/jscott3201/Helios">MLX Gemma 12B</a></p> <p>The main constraints for this on my end
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Check it out: <a href="https://github.com/fairydreaming/llama.cpp/tree/dsv4">https://github.com/fairydreaming/llama.cpp/tree/dsv4</a></p> <p>They are PRs <a href="https://github.com/ggml-org/llama.cpp/pull/25247">#25247</a>, <a href="https://gith
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Found cheap one RTX5000 16GB is it worth to add to RTX3060? And is it difficult to make them work together in LM Studio? Do I need to use studio drivers and is RTX5000 still supported with driver updates? Or should I just go for 5060Ti instead? T
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unbn4y/comparing_local_inference_speeds_across_a_few/"> <img alt="Comparing local inference speeds across a few real setups people are running (3090 vs 5090 vs dual 6000)" src="https://preview.redd.it/8aep4zi
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unbm45/ran_a_classicmedival_europe_fantasy_rpagentic/"> <img alt="Ran a classic(medival europe) fantasy RP/agentic benchmark across 8 local models Qwen3.6-27B held up better than its size suggests" src="https
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unbjr2/using_local_models_with_hermes_vs_claude_code/"> <img alt="Using local models with Hermes vs Claude code" src="https://preview.redd.it/ii7l8k4wa8bh1.png?width=640&crop=smart&auto=webp&s=350
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1unbi4a/qwen36_27b_on_a_5090_64k_sample_toks_distribution/"> <img alt="Qwen3.6 27B on a 5090, 6.4k sample tok/s distribution after tuning MTP/cache settings" src="https://preview.redd.it/l2k7gu5cb8bh1.jpeg?wid
reddit:LocalLLaMA- 2026-07-04DGX Spark and Overtemps
<!-- SC_OFF --><div class="md"><p>For anyone who has a DGX-Spark and is having problems during these very hot summer months, you can underclock with:</p> <p>sudo nvidia-smi -lgc 0,900</p> <p>My temps dropped from 85C to 60C and this fixed my problem of overtemp lockups.</p> <p>Ed
reddit:LocalLLaMA - 2026-07-04Feedback needed
<!-- SC_OFF --><div class="md"><p>Dear LocalLLaMA,</p> <p>I’m looking for feedback on my project because I noticed it is mainly catered towards technical people.</p> <p><a href="https://github.com/heterodoxin/apostate">Apostate</a> (Apostate is an abliteration engine built from s
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>A lot of people seem to be confused or mystified about this so figured I'd spell it out.</p> <p>I played around with RYS and realized that it broke Gemma 4 models. Turns out there's a `layer_scalar` value that is applied at each layer. If you don
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Here's the setup I decided on for embedding gemma4-12b into a Tauri2 desktop app:</p> <ul> <li>Native Rust FFI into llama.cpp via <code>llama-cpp-2</code> (Metal enabled)</li> <li>Model:<code>gemma-4-12b-it-Q5_K_S</code> quantized by Unsloth, Q5_
reddit:LocalLLaMA- 2026-07-04First time I have seen this: my model seemed aware of its context usage ask me for compaction!
<!-- SC_OFF --><div class="md"><p>I was in a middle of a Claude Code session with GLM 5.2. Context usage 537k/1M. After finishing a task, GLM asked me this:</p> <blockquote> <p><strong>Context</strong> <strong>note:</strong> this session has run long and context is getting heavy.
reddit:LocalLLaMA <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1un9955/paper_gear_guided_endtoend_autoregression_for/"> <img alt="[Paper] GEAR: Guided End-to-End AutoRegression for Image Synthesis" src="https://preview.redd.it/vo0a7q0ut7bh1.png?width=640&crop=smart&am
reddit:LocalLLaMA- 2026-07-04Using structured insurance/captives to solve public pushback & zoning delays? (Idea discussion)
<!-- SC_OFF --><div class="md"><p>I’m looking for some blunt feedback on a strategy to handle the increasing community/municipal pushback on new builds—specifically regarding power grid strain, water use, and the "what's in it for the community" argument.<br /> Right no
reddit:datacenter - 2026-07-04Questions about AWS data center roles.
<!-- SC_OFF --><div class="md"><p>Hello everyone, I'm currently working at a data center on the operations side of things and am looking to move over to the compute side, something I can't do with my current employer. </p> <p>I'm wondering what are the differences between an Infr
reddit:datacenter <!-- SC_OFF --><div class="md"><p>Hi everyone,<br /> I had three interviews scheduled with Google for a Data Center Technician role, but I only completed two so far. The last one was rescheduled for next week and will focus on behavioral questions.<br /> The first two interviews
reddit:datacenter<!-- SC_OFF --><div class="md"><p>I would like to work at a local wastewater treatment plant consulting firm in Electrical and then I was wondering can’t I go from that directly to a faang as an EE?</p> <p>Does a lot of the mission critical stuff transfer eg the ATS,Motors and Ge
reddit:datacenter- 2026-07-03Introducing myself
<!-- SC_OFF --><div class="md"><p>Hello,</p> <p>I am a newcomer to the <strong>DCT</strong> and am eager to begin my journey with <strong>AWS infrastructure</strong> in about a month. I possess a modest background in technology. I am seeking mentorship and guidance as I embark on
reddit:datacenter - 2026-07-03Automotive technician to data center
<!-- SC_OFF --><div class="md"><p>No doubt this has been asked in the past, but have any of you moved over to a data center role from being an automotive tech?</p> <p>A buddy of mine is a data center tech, and what he was telling me sounds like I could do it based on my electrica
reddit:datacenter - 2026-07-03Daily Discussion Friday 2026-07-03
  submitted by   <a href="https://www.reddit.com/user/AutoModerator"> /u/AutoModerator </a> <br /> <span><a href="https://www.reddit.com/r/AMD_Stock/comments/1um3smu/daily_discussion_friday_20260703/">[link]</a></span>   <span><a href="https://www.reddit.com/r/AMD_Sto
reddit:AMD_Stock <table> <tr><td> <a href="https://www.reddit.com/r/AMD_Stock/comments/1uliahe/technical_analysis_for_amd_72premarket/"> <img alt="Technical Analysis for AMD 7/2----------Pre-Market" src="https://preview.redd.it/0lowepuhitah1.png?width=140&height=87&auto=webp&s=f7b2d46
reddit:AMD_Stock<table> <tr><td> <a href="https://www.reddit.com/r/AMD_Stock/comments/1uleqft/photonic_ai_network_heads_for_first_commercial/"> <img alt="Photonic AI network heads for first commercial deployment with AMD" src="https://external-preview.redd.it/Ic_WChybAirlQH_96lvbcVmJr_KxHAOy0bB7
reddit:AMD_Stock- 2026-07-02Daily Discussion Thursday 2026-07-02
  submitted by   <a href="https://www.reddit.com/user/AutoModerator"> /u/AutoModerator </a> <br /> <span><a href="https://www.reddit.com/r/AMD_Stock/comments/1ul7jj0/daily_discussion_thursday_20260702/">[link]</a></span>   <span><a href="https://www.reddit.com/r/AMD_S
reddit:AMD_Stock <table> <tr><td> <a href="https://www.reddit.com/r/AMD_Stock/comments/1ukwedq/dont_worry_about_it_amd_is_still_going_all_the/"> <img alt="Don't worry about it. AMD is still going all the way to the edge." src="https://preview.redd.it/satgwymocoah1.jpeg?width=640&crop=smart&am
reddit:AMD_Stock<table> <tr><td> <a href="https://www.reddit.com/r/AMD_Stock/comments/1ukl9my/technical_analysis_for_amd_712nd_half_of_the_year/"> <img alt="Technical Analysis for AMD 7/1----------2nd half of the year" src="https://preview.redd.it/fy0j88d0amah1.png?width=140&height=86&au
reddit:AMD_Stock<!-- SC_OFF --><div class="md"><p>For the last few months I've been building a conversational CRM called Kessio. Instead of manually updating contacts, deals and follow ups after every meeting, you just type what happened and it takes care of the rest.</p> <p>A few days ago I fin
reddit:startups- 2026-07-01Founders Institute, i will not promote
<!-- SC_OFF --><div class="md"><p>For a company with co-founders, they evaluate application per founder, ask each founder to give 900$ and ~3% equity. </p> <p>My co-founder got accepted but then i am asked to write another application. </p> <p>That sounds really weird, how can yo
reddit:startups <!-- SC_OFF --><div class="md"><p>CEO of a $2M ARR AI startup that sell to tech midmarket/enterprises. most of it was added in the last 18 months.</p> <p>We're struggling with pipeline generation right now. Our ICP changed and we see larger companies come to our demos and actuall
reddit:startups<!-- SC_OFF --><div class="md"><p>I’m trying to sell software (AI agents for specific problems) to franchise dealerships in the US.</p> <p>Just trying to get our first paid pilot now</p> <p>However, I’m not based in the US.</p> <p>And especially with the automotive industry, it o
reddit:startups<!-- SC_OFF --><div class="md"><p>My life goal has been and will always to build a successful start up, whether that is as a founder or early stage member etc</p> <p>I want to break into the fintech, specially fraud-prevention, but i don't have any experience outside of my legacy
reddit:startups<!-- SC_OFF --><div class="md"><p>Does there need to be 3 or more direct competitors in to the problem/s I’m planning to revolve my startup around?</p> <p>It’s AI agents for a few specific problems inside service departments for automotive technicians</p> <p>There’s a whole bunch
reddit:startups- 2026-06-30Tried building my own startup, it failed, now I want to learn from someone else's i will not promote
<!-- SC_OFF --><div class="md"><p>After trying to build my own start up and failing I'm beginning and I would like to start at interning as a product analyst intern or customer Insights intern in a start up or work with a founder. Happy to have a conversation with anyone, also if
reddit:startups <!-- SC_OFF --><div class="md"><p>I built an event/experience discovery platform. Before users can use a marketplace, it helps if there are things in it. </p> <p>I solved the cold start problem by deciding to make profiles for the events and businesses that I plan to feature with
reddit:startups<!-- SC_OFF --><div class="md"><p>I’m currently building in the ai/ml infrastructure space, specifically around automating model training and fine-tuning workflows: from dataset preparation to training, evaluation, and deployment.</p> <p>we’re still pre-revenue / early build stag
reddit:startups<!-- SC_OFF --><div class="md"><p>I’m currently building in the ai/ml infrastructure space, specifically around automating model training and fine-tuning workflows: from dataset preparation to training, evaluation, and deployment.</p> <p>we’re still pre-revenue / early build stag
reddit:startups