Keeps: post title, subreddit, score
- 2026-08-09RTX 5090 96GB spotted on Alibaba?
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vjcljq/rtx_5090_96gb_spotted_on_alibaba/"> <img alt="RTX 5090 96GB spotted on Alibaba?" src="https://preview.redd.it/rzunb9uu39ih1.png?width=640&crop=smart&auto=webp&s=8a871629ca059e9bc1e608f1db7e
reddit:LocalLLaMA - 2026-08-09Intel Optane - Potential?
<!-- SC_OFF --><div class="md"><p>Intel's Optane technology as we know it has been discontinued. But while that's the case, I see it being highly promising still due to its role (or where it was meant to be) right in between system RAM and NVME SSDs. As far as I understand, it si
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Pasted the same HTML/JS code (330 lines) into Qwen 35B A3B and Gemma 26B A4B.</p> <p>Qwen: tokenized the input to 1609 tokens</p> <p>Gemma: tokenized the input to 4258 tokens.</p> <p>Damn. I've never noticed this before and I haven't seen people
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Firstly a big thanks to the poster "hellohazine", he basically only removed the multi-lingual fat of the model and just kept the English language intact. It is the exact model, and the rest of the model still intact with all of its high
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Hey everyone, I could use some advice on setting up speculative decoding correctly with <code>llama-server</code>.</p> <p><strong>My Hardware:</strong></p> <ul> <li><strong>GPUs:</strong> RTX 4090 + RTX 6000 Pro (120GB total VRAM)</li> <li><stron
reddit:LocalLLaMA- 2026-08-08Is Microsoft-Phi dead?
<!-- SC_OFF --><div class="md"><p>Phi was one of my favorite models with a bit of a mixed reputation with some claiming it's benchmaxxed and others seeing its potential and usecases. I was a big fan of Phi but the last major release was in december 2024 with every other release b
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Disclaimer - no LLM was used to write this post/note</p> <p>As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs.</p> <p>So I have pretty meaty server (8 channel AMD EPYC, ~150GB/s RAM bw) an
reddit:LocalLLaMA- 2026-08-08Comparing 4bit quants for MLX
<!-- SC_OFF --><div class="md"><p>Curious what people think are the ideal 4-bit quantization types on MLX</p> <p>These quants seem to be the most popular, at least for Gemma4 and Qwen3.6:<br /> - OptiQ 4bit (<a href="https://huggingface.co/mlx-community/Qwen3.6-27B-OptiQ-4bit">ml
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 million tokens?</p> </div><!-- SC_ON -->  
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>I’m planning a dedicated home AI server, mainly for local LLM inference, agents/tool use, Docker services, and eventually larger MoE models with CPU offload.</p> <p>My plan is to start with <strong>2x 16GB GPUs = 32GB VRAM</strong>, but I want to
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vj3b7m/i_built_a_local_realtime_voice_stack_for_ollama/"> <img alt="I built a local realtime voice stack for Ollama: Parakeet STT → Qwen 2.5 7B → Qwen3-TTS" src="https://external-preview.redd.it/aGo1MDYxZXgxN
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>I made gemma4 12B write timestamp-anchored summaries of youtube video transcripts. I tested if the summaries have significant qualitative variance and if the SLM can pick the best one by itself. Below is the prompt texts I used.</p> <pre><code>&q
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Over the past few months I have been building a CPU-first inference engine from scratch in pure C99 (no Python, no CUDA, no BLAS, just GCC and make). The focus has been running 1.58-bit ternary models natively without heavy runtime overhead.</p>
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vj18h4/showoff_saturday_local_4x_6000_pro_multiyear/"> <img alt="Showoff Saturday: Local 4x 6000 Pro (multi-year progression)" src="https://preview.redd.it/n9061cgul6ih1.jpg?width=140&height=105&auto=
reddit:LocalLLaMA- 2026-08-08My first run of Kimi K3 locally.
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vj0hil/my_first_run_of_kimi_k3_locally/"> <img alt="My first run of Kimi K3 locally." src="https://preview.redd.it/uah37umch6ih1.png?width=140&height=97&auto=webp&s=86aa0ba10df329017df484eb6b4b79b
reddit:LocalLLaMA <table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1viyn0c/we_tested_the_impact_of_ssd_speed_on_gaming/"> <img alt="We tested the impact of SSD speed on gaming performance in 11 titles — we analyzed from SATA to PCIe 5.0 to see whether upgrading to a faster NVMe
reddit:hardware- 2026-08-08Tesla V100 Qwen3.6 27B Performance
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vixl0f/tesla_v100_qwen36_27b_performance/"> <img alt="Tesla V100 Qwen3.6 27B Performance" src="https://preview.redd.it/bynog4baw5ih1.png?width=140&height=71&auto=webp&s=1bf2313037350d989bee6dfdc62
reddit:LocalLLaMA <table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vit8mg/hub_how_much_vram_can_gamers_get_away_with_4gb_vs/"> <img alt="HUB - How Much VRAM Can Gamers Get Away With? 4GB vs. 8GB" src="https://external-preview.redd.it/Z4Edrj700dL7F_W0V_Bv1Q5BP7PQQ-ni68UoRxzE6rI
reddit:hardware<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1viqwqm/gpu_viewadaptive_crackfree_subdivision_of_bézier/"> <img alt="GPU view-adaptive crack-free subdivision of Bézier surfaces" src="https://external-preview.redd.it/2GGMaSdGCDOPRRW6Bn21SuTSooa8gpP9WZgjVSWWEh
reddit:hardware<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1viqtgm/2027_memory_capacity_is_reportedly_sold_out/"> <img alt="2027 Memory Capacity Is Reportedly Sold Out" src="https://external-preview.redd.it/RULANo01uf3INcx62jxJ05Z56Nn9ube37FArEuVZSGw.jpeg?width=640&am
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1viqrp8/ray_tracing_massive_amounts_of_animated_geometry/"> <img alt="Ray tracing massive amounts of animated geometry using tetrahedral cages" src="https://external-preview.redd.it/hgd7EwulMc7UUDnMdgoo8E2c-xDhr
reddit:hardware- 2026-08-08DeepSeek V4 Flash 0731 appreciation post
<!-- SC_OFF --><div class="md"><p>I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real.</p> <p>Everyday tasks with Hermes agent? Effortless.</p> <p>Coding tasks with OpenCode? I’m genuinely amazed at what it can handle.
reddit:LocalLLaMA - 2026-08-08Anyone else amped up over Qwen 3.8?
<!-- SC_OFF --><div class="md"><p>I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a year, I can see home LLMs being served at home much like streaming music was introduced. A simple browser extensi
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about <strong>3.9× faster (~116 vs ~30 tok/s)</strong>, but the coding-quality difference was m
reddit:LocalLLaMA- 2026-08-084x2080Ti 22GB tensor parallel possible?
<!-- SC_OFF --><div class="md"><p>Hi there, I have a setup with 4x 2080Ti 22GB, but am not able to see much benefits in running models larger than 16-20GB in speed. I did some research and saw something like tensor parallels are possible, but also that because of flash attention
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Basically what the title says. For me, it would always crash and burn trying to use tensor split.</p> <p>Apparently, there's some bug where GPU memory gets corrupted with the default microbatch (512) or higher. I will be opening an issue report o
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vilcil/i_got_tired_of_my_300gb_model_loads_taking_5min/"> <img alt="I got tired of my 300GB model loads taking 5min on RPC. PR 26291 speeds it 300% to 1min30sec (4060ti+ddr4) + (4060ti+ddr5)" src="https://pre
reddit:LocalLLaMA- 2026-08-08BeeLLama issues
<!-- SC_OFF --><div class="md"><p>Tried to use Beellama , and using the kvarn6 flag, i notice that its in llama-server --help but its not working. I must be doing something wrong. </p> <p><strong>trying to run the following:</strong><br /> llama-server.exe ^</p> <p>--model "
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p><em>(I am not a native speaker, written by myself, so please bear with me)</em></p> <p>I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence benchmarks. And those flaws render
reddit:LocalLLaMA  submitted by   <a href="https://www.reddit.com/user/johnnyApplePRNG"> /u/johnnyApplePRNG </a> <br /> <span><a href="https://genesisopenmodels.anl.gov/">[link]</a></span>   <span><a href="https://www.reddit.com/r/LocalLLaMA/comments/1vijp8y/us_department_of_energy_la
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vijbrg/why_did_amd_just_buy_this_really_weird_chip/"> <img alt="Why did AMD just buy this REALLY WEIRD chip company? - TechTechPotato" src="https://external-preview.redd.it/RjmcTuxn4FOKQsb21MhYb8cPwBq0prWeKQS_E
reddit:hardware<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vih8vs/posted_a_video_where_i_repaste_and_benchmark_the/"> <img alt="Posted a video where i repaste and benchmark the GTX 970 in 2026 with modern games" src="https://external-preview.redd.it/8u_0AxYpGduscsApEvI
reddit:hardware<!-- SC_OFF --><div class="md"><p>Hey all, I'm serving DSv4Flash 0731 on a cluster of 2x DGX Sparks but am running into constant issues with having almost no RAM (unified memory) left for the OS/cache and I'd love to hear the community feedback on what I could do to get more RAM
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vifnby/amazon_cracks_down_on_cpu_waste_among_engineers/"> <img alt="Amazon cracks down on 'CPU waste' among engineers as agentic AI crunch intensifies — CPU demand makes low-utilization EC2 instances a hot comm
reddit:hardware<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1viffzo/sk_hynix_to_invest_54_trillion_won_in/"> <img alt="SK Hynix to Invest 54 Trillion Won in Semiconductor Facilities. New Plants in Yongin, Cheongju to Produce Next-Gen DRAM, NAND by 2029" src="https://exte
reddit:hardware<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vieca5/nanya_technology_invests_107_billion_in_euv_dram/"> <img alt="Nanya Technology Invests $10.7 Billion in EUV DRAM Expansion" src="https://external-preview.redd.it/nkOxJ9diCKDC3aXj0a_bLU3rO-grJKLy95nAcsW0C
reddit:hardware- 2026-08-07Qwen 3.6 27B flags/settings in llama.cpp
<!-- SC_OFF --><div class="md"><p>I run the following on a 5090 and have been okay with its performance, it does most things somewhere 80-100 t/s, though that can slow down at full 262k context - more like 40 t/s at times. I use it primarily in appdev tasks. This just barely fits
reddit:LocalLLaMA <table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1viamnr/intel_just_matched_apple_silicon_seriously/"> <img alt=" Intel just matched Apple Silicon. Seriously." src="https://external-preview.redd.it/SYaGXS_NgdBhRKjW12BVb8_-L57EUqjRvmbw7QY5iUY.jpeg?width=320&
reddit:hardware<!-- SC_OFF --><div class="md"><p>Hey all,</p> <p>Are there any Delphi developers in the crowd? If so, which models would you say are best at doing development in Delphi? What are your thoughts/suggestions here, and is there a good GUI client/harnass you like for doing delphi spe
reddit:LocalLLaMA- 2026-08-07DeepSeek V4 Flash 0731 - ARC-AGI Results
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vi9zls/deepseek_v4_flash_0731_arcagi_results/"> <img alt="DeepSeek V4 Flash 0731 - ARC-AGI Results" src="https://external-preview.redd.it/sWhpb1GjRlbd3knWV_xC1C2WMQX5RRFImjvBgSF_7ZI.png?width=640&crop=sma
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Hey everyone, I just wanted to share my journey here for some motivation.</p> <p>Three years ago, I saw the sudden spike in AI and realized it was the future of tech. My goal at the time was to be an indie game dev, and seeing that AI could write
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vi7eqk/exclusive_apple_is_scrambling_for_dram_weeks/"> <img alt="Exclusive: Apple is Scrambling for DRAM Weeks Ahead of Ultra, iPhone 18 Launch" src="https://external-preview.redd.it/MMZEuxTmjqXyrLJVbf32Hbik93P
reddit:hardware<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vi77dr/parakeetwgsl_fast_accurate_asr_in_the_browser_via/"> <img alt="parakeet.wgsl – Fast, accurate ASR in the browser, via raw WebGPU & SIMD WASM" src="https://external-preview.redd.it/aGFyZHk4cnZuemhoM
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vi6hmw/llamacpp_pr_reports_up_to_169_faster_quantizedkv/"> <img alt="llama.cpp PR reports up to 169% faster quantized-KV decode at 118K context on Intel Battlemage from one SYCL kernel switch" src="https://pr
reddit:LocalLLaMA- 2026-08-07What happened to the Mobidapter?
<!-- SC_OFF --><div class="md"><p>People used Mobidapters for old phones. It wasn’t a card reader, it was an adapter so you could plug a USB thumb drive into the phone. This could have been really good for Raspberry Pis or Orange Pi 5 Pluses because Microsds are getting more expe
reddit:hardware <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vi1r6t/wananimate2_pushing_the_application_boundaries_of/"> <img alt="Wan-Animate-2: Pushing the Application Boundaries of Character Animation Models" src="https://preview.redd.it/7t00lk39nyhh1.png?width=640&
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vi0d4i/lfm2526b_modelkv_cache_quantization_report/"> <img alt="LFM2.5-2.6B model+KV cache quantization report" src="https://preview.redd.it/3lr47ialdyhh1.png?width=140&height=92&auto=webp&s=e5244a
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vhz989/a_llamacpp_pr_makes_q2_0_3036x_faster_on_x86_cpus/"> <img alt="A llama.cpp PR makes Q2_0 3.0–3.6x faster on x86 CPUs, 8B decode goes 2.39 → 8.20 tok/s" src="https://preview.redd.it/pyim0m155yhh1.jpeg?w
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vhy2e6/rtx_5090_owner_built_an_opensource_tool_that/"> <img alt="RTX 5090 Owner Built An Open-Source Tool That Shuts Down PC If It Detects The 12VHPWR Cable Drawing Too Much Power, But It Can Only Work On Spe
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vhwqih/price_hikes_may_be_coming_for_pc_motherboards_next/"> <img alt="Price Hikes May Be Coming for PC Motherboards Next" src="https://external-preview.redd.it/OuRNAj98ocMRlkcLHuwjKZ6me96KJ_CF95mgg_z8UDI.jpeg?
reddit:hardware