Keeps: post title, subreddit, score
<!-- SC_OFF --><div class="md"><p>Firstly a big thanks to the poster "hellohazine", he basically only removed the multi-lingual fat of the model and just kept the English language intact. It is the exact model, and the rest of the model still intact with all of its high
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Hey everyone, I could use some advice on setting up speculative decoding correctly with <code>llama-server</code>.</p> <p><strong>My Hardware:</strong></p> <ul> <li><strong>GPUs:</strong> RTX 4090 + RTX 6000 Pro (120GB total VRAM)</li> <li><stron
reddit:LocalLLaMA- 2026-08-08Is Microsoft-Phi dead?
<!-- SC_OFF --><div class="md"><p>Phi was one of my favorite models with a bit of a mixed reputation with some claiming it's benchmaxxed and others seeing its potential and usecases. I was a big fan of Phi but the last major release was in december 2024 with every other release b
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Disclaimer - no LLM was used to write this post/note</p> <p>As larger post about my setup will come later, want to give heads-up to folks who use VLLM and >= 2 GPUs.</p> <p>So I have pretty meaty server (8 channel AMD EPYC, ~150GB/s RAM bw) an
reddit:LocalLLaMA- 2026-08-08Comparing 4bit quants for MLX
<!-- SC_OFF --><div class="md"><p>Curious what people think are the ideal 4-bit quantization types on MLX</p> <p>These quants seem to be the most popular, at least for Gemma4 and Qwen3.6:<br /> - OptiQ 4bit (<a href="https://huggingface.co/mlx-community/Qwen3.6-27B-OptiQ-4bit">ml
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 million tokens?</p> </div><!-- SC_ON -->  
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>I’m planning a dedicated home AI server, mainly for local LLM inference, agents/tool use, Docker services, and eventually larger MoE models with CPU offload.</p> <p>My plan is to start with <strong>2x 16GB GPUs = 32GB VRAM</strong>, but I want to
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vj3b7m/i_built_a_local_realtime_voice_stack_for_ollama/"> <img alt="I built a local realtime voice stack for Ollama: Parakeet STT → Qwen 2.5 7B → Qwen3-TTS" src="https://external-preview.redd.it/aGo1MDYxZXgxN
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>I made gemma4 12B write timestamp-anchored summaries of youtube video transcripts. I tested if the summaries have significant qualitative variance and if the SLM can pick the best one by itself. Below is the prompt texts I used.</p> <pre><code>&q
reddit:LocalLLaMA<!-- SC_OFF --><div class="md"><p>Over the past few months I have been building a CPU-first inference engine from scratch in pure C99 (no Python, no CUDA, no BLAS, just GCC and make). The focus has been running 1.58-bit ternary models natively without heavy runtime overhead.</p>
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vj18h4/showoff_saturday_local_4x_6000_pro_multiyear/"> <img alt="Showoff Saturday: Local 4x 6000 Pro (multi-year progression)" src="https://preview.redd.it/n9061cgul6ih1.jpg?width=140&height=105&auto=
reddit:LocalLLaMA- 2026-08-08My first run of Kimi K3 locally.
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vj0hil/my_first_run_of_kimi_k3_locally/"> <img alt="My first run of Kimi K3 locally." src="https://preview.redd.it/uah37umch6ih1.png?width=140&height=97&auto=webp&s=86aa0ba10df329017df484eb6b4b79b
reddit:LocalLLaMA <table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1viyn0c/we_tested_the_impact_of_ssd_speed_on_gaming/"> <img alt="We tested the impact of SSD speed on gaming performance in 11 titles — we analyzed from SATA to PCIe 5.0 to see whether upgrading to a faster NVMe
reddit:hardware- 2026-08-08Tesla V100 Qwen3.6 27B Performance
<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vixl0f/tesla_v100_qwen36_27b_performance/"> <img alt="Tesla V100 Qwen3.6 27B Performance" src="https://preview.redd.it/bynog4baw5ih1.png?width=140&height=71&auto=webp&s=1bf2313037350d989bee6dfdc62
reddit:LocalLLaMA <table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vit8mg/hub_how_much_vram_can_gamers_get_away_with_4gb_vs/"> <img alt="HUB - How Much VRAM Can Gamers Get Away With? 4GB vs. 8GB" src="https://external-preview.redd.it/Z4Edrj700dL7F_W0V_Bv1Q5BP7PQQ-ni68UoRxzE6rI
reddit:hardware<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1viqwqm/gpu_viewadaptive_crackfree_subdivision_of_bézier/"> <img alt="GPU view-adaptive crack-free subdivision of Bézier surfaces" src="https://external-preview.redd.it/2GGMaSdGCDOPRRW6Bn21SuTSooa8gpP9WZgjVSWWEh
reddit:hardware<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1viqtgm/2027_memory_capacity_is_reportedly_sold_out/"> <img alt="2027 Memory Capacity Is Reportedly Sold Out" src="https://external-preview.redd.it/RULANo01uf3INcx62jxJ05Z56Nn9ube37FArEuVZSGw.jpeg?width=640&am
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1viqrp8/ray_tracing_massive_amounts_of_animated_geometry/"> <img alt="Ray tracing massive amounts of animated geometry using tetrahedral cages" src="https://external-preview.redd.it/hgd7EwulMc7UUDnMdgoo8E2c-xDhr
reddit:hardware- 2026-08-08DeepSeek V4 Flash 0731 appreciation post
<!-- SC_OFF --><div class="md"><p>I’m running DSV4F 0731 on dual spark, and honestly… wow. It’s an absolute workhorse, and the benchmarks are real.</p> <p>Everyday tasks with Hermes agent? Effortless.</p> <p>Coding tasks with OpenCode? I’m genuinely amazed at what it can handle.
reddit:LocalLLaMA - 2026-08-08Anyone else amped up over Qwen 3.8?
<!-- SC_OFF --><div class="md"><p>I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a year, I can see home LLMs being served at home much like streaming music was introduced. A simple browser extensi
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about <strong>3.9× faster (~116 vs ~30 tok/s)</strong>, but the coding-quality difference was m
reddit:LocalLLaMA- 2026-08-084x2080Ti 22GB tensor parallel possible?
<!-- SC_OFF --><div class="md"><p>Hi there, I have a setup with 4x 2080Ti 22GB, but am not able to see much benefits in running models larger than 16-20GB in speed. I did some research and saw something like tensor parallels are possible, but also that because of flash attention
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p>Basically what the title says. For me, it would always crash and burn trying to use tensor split.</p> <p>Apparently, there's some bug where GPU memory gets corrupted with the default microbatch (512) or higher. I will be opening an issue report o
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vilcil/i_got_tired_of_my_300gb_model_loads_taking_5min/"> <img alt="I got tired of my 300GB model loads taking 5min on RPC. PR 26291 speeds it 300% to 1min30sec (4060ti+ddr4) + (4060ti+ddr5)" src="https://pre
reddit:LocalLLaMA- 2026-08-08BeeLLama issues
<!-- SC_OFF --><div class="md"><p>Tried to use Beellama , and using the kvarn6 flag, i notice that its in llama-server --help but its not working. I must be doing something wrong. </p> <p><strong>trying to run the following:</strong><br /> llama-server.exe ^</p> <p>--model "
reddit:LocalLLaMA <!-- SC_OFF --><div class="md"><p><em>(I am not a native speaker, written by myself, so please bear with me)</em></p> <p>I really want to like DeepSeek-V4-Flash-0731. But it has serious flaws that don't align with the high score on intelligence benchmarks. And those flaws render
reddit:LocalLLaMA  submitted by   <a href="https://www.reddit.com/user/johnnyApplePRNG"> /u/johnnyApplePRNG </a> <br /> <span><a href="https://genesisopenmodels.anl.gov/">[link]</a></span>   <span><a href="https://www.reddit.com/r/LocalLLaMA/comments/1vijp8y/us_department_of_energy_la
reddit:LocalLLaMA<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vijbrg/why_did_amd_just_buy_this_really_weird_chip/"> <img alt="Why did AMD just buy this REALLY WEIRD chip company? - TechTechPotato" src="https://external-preview.redd.it/RjmcTuxn4FOKQsb21MhYb8cPwBq0prWeKQS_E
reddit:hardware<table> <tr><td> <a href="https://www.reddit.com/r/hardware/comments/1vih8vs/posted_a_video_where_i_repaste_and_benchmark_the/"> <img alt="Posted a video where i repaste and benchmark the GTX 970 in 2026 with modern games" src="https://external-preview.redd.it/8u_0AxYpGduscsApEvI
reddit:hardware