---
title: "Qwen 3.6 27B Dense — The Real Local AI Starting Point"
date: 2026-07-14T12:30:00-04:00
account: braincramps
type: thread
tweet_url: "https://x.com/yume_arasaki/status/2077042336522207337"
published: true
---
@yume_arasaki (Yume_X) — 70 followers
夢 . Chasing free AI. Do not let the Epstein Class win. Decoding post AI human workflows. Covering what can be done with Hermes @NousResearch and Local LLMs
Posted: Mon Jul 14 2026
---
Everyone's "local ai gateway drug" is the Qwen 3.6 27B dense. It was mine.
The model that made local AI real. Fits on a 24GB card, 40 tok/s, flagship-level coding.
It's the starting point.
Afterwards the questions becomes "what's next"?
People recommend the 35B-A3B MoE as "the next step." It's not.
It's 4x faster but hallucinates in agent loops.
Tool call failures at 18-76% rates.
Context hallucinations at 37% usage. Qwen's own benchmarks show 27B beating it on agentic coding.
For agents, it's a downgrade.
The real upgrade paths depend on which constraint you're breaking: speed, quality, model size, or sovereignty.
It depends largely on how much VRAM you want to play with imo
Here's the map.
--
PATH 1: Same card, enable MTP (free)
The upgrade most people miss. Unsloth shipped MTP GGUFs. Multi-token prediction.
The draft head predicts 2 tokens in parallel, main model verifies all 3 in one forward pass. 83% acceptance rate.
Your 27B at Q4 running 40 tok/s? With MTP it hits 60-80 tok/s. Same model, same card, same quality. On an RTX 6000 Ada it hits ~160 tok/s.
Cost: $0.
--
PATH 2: RTX 5090, 32GB Blackwell (~$4-5K)
32GB GDDR7 at 1,792 GB/s. Same bandwidth class as the $12K RTX PRO 6000.
Native NVFP4 tensor cores.
What changes: 27B at Q8 fits cleanly with room for long context. NVFP4 quants run 2.5x faster than Q4_K_M on the same silicon.
sudoingX benched the 27B on a 5090 laptop: 35.3 tok/s generation, 1,509 tok/s prefill. The prefill is where Blackwell pulls ahead.
This is the supposed to be the cheapest card that runs the newest quant format., but it's not, supply chain means you are unlikely to find one at the RSP of $2.5k .
Cost: ~$4000-5,000 - Probably not worth it IMO, at its retail price of $2.5 it's more interesting. Supply issues pushing the card price up makes it unatttractive
--
PATH 3: Strix Halo / AI Max+ 395, 128GB unified (~$2.5K)
The budget Spark alternative. Geforce VS AMD all over again lol. AMD cheaper but…it's a huge underperformer without NVIDIA's efficiency gains due to large community on NVIDIA.
AMD's Strix Halo gives you 128GB unified at roughly $2K.
Stepfun officially lists it as a Step 3.7 Flash target.
Less bandwidth than the Spark. No NVFP4 tensor cores.
But half the price for the same memory capacity.
If you want to run 198B models and can't afford a Spark, maybe this is a choice
Cost: ~$2,500 --- To be honest, it's not worth paying less for a less serviced ecosystem, a DGX spark is a much better buy.
--
PATH 4: DGX Spark, 128GB (~$4,699)
Not a dense model speed play, this is where the DGX Spark is weak. Where DGX sparks excel is running MOE models, here it kinda becomes like a KING.
It's not a good GPU in terms of memory bandwidth (something dense models need)
273 GB/s bandwidth vs your 3090's 936.
Unsloth's NVFP4 quants (released July 10) halve bytes per parameter on Blackwell, which narrows the gap.
The Spark IS Blackwell.
What it's good at: running models too big for any consumer GPU, and serving many users at once.
The trick with it is having mixture of experts. Speed runs DSF4 (easily basically the MOST DIRECT "next step up" from Qwen 27B Dense.
Step 3.7 Flash (198B MoE, 11B active): sudoingX's top Spark pick. 198 billion parameters with vision. Full 256K context at 25 tok/s. Knowledge breadth is 7x the 27B.
Nemotron-Labs-3-Puzzle-75B-A9B: NVIDIA compressed their 120B Super down to 75B total / 9.3B active. NVIDIA's own post markets it as "perfect for your single GB10." 97%+ of parent quality. NVFP4 weights are 44.5GB. Released July 9.
Nemotron-3-Super 120B-A12B (NVFP4): 21-25 tok/s (NVIDIA developer forum). 1M context. 12B active. Native MTP. Built for multi-agent.
gpt-oss-120B (MXFP4): 57-60 tok/s pure decode. Fastest single-stream on the Spark.
Concurrency: WescheNex1q ran 64 concurrent users on one Spark. 700+ tok/s aggregate.
Cost: $4,699 --- Really good buy, but really understand the issues of low memory bandwidth with this one, and it's trade-offs
--
PATH 5: RTX 6000 Ada, 48GB (~$6,800)
Single card. 960 GB/s. Same 48GB as dual 3090s but no PCIe overhead, no heat nightmare.
Nemotron-3-Super 120B at Q3 (~58GB) fits with offload.
Honest take: 5x the cost of a single 3090 for double the VRAM. Rough price per GB.
But cleanest single-card 48GB path. Cost: ~$6,800
--
PATH 6: 2 DGX Sparks, 256GB (~$9,500)
"Near" Frontier-scale models fully offline: Qwen3-235B-A22B: 17 tok/s batch=1, 36 aggregate at batch=4. Beats NVIDIA's own published number by 45%.
MiniMax-M3 (428B MoE + vision): Fits at AutoRound 3.2-bit mixed quant (188.6GB across two nodes). 13.7 tok/s prose, 15 tok/s code, peaks of 20 with EAGLE3 speculative decoding. Vision tower works. Not full precision. 3.2-bit is aggressive. But it runs.
GLM-5.2: Needs 4 Sparks for full quality. 2 Sparks is the down payment.
This is why I bought mine. 256GB unified holds weights AND parallel agent KV caches. No rate limits. No provider seeing your prompts.
Cost: ~$9,500 with DAC cable ---
--
PATH 7: RTX PRO 6000 Blackwell, 96GB (~$12-14.5K)
1,792 GB/s. Native NVFP4.
This is what @Hikari_07_jp runs (two of them, 192GB total). 6.7x faster than the Spark on the same models.
Nemotron-3-Super 120B at Q4 (~70GB) fits on one card with room. Dense models fly. The bandwidth king.
$138 per GB of VRAM vs $33 for a 3090. Density buy, not a generalist one.
It has a strong upgrade path, but the rig you build if you want high VRAM e.g. 4-6x of these cards, will allow you to literally run the largest models, but cost will run to $70-90k, AND as others pointed out there will be a lot of HVAC things.
Cost: $12,000-14,500
--
The decision:
Want free speed? Enable MTP. 1.4-2.2x on same hardware.
Want speed demon Blackwell on a prosumer budget? RTX 5090. 32GB, NVFP4, 4.5K. Not recommended imo (tiny Vram, very overpriced), might as well just jump higher, or stick with a 3090 or 4090. (and save that $4k)
Want frontier MoE offline? Spark. Step 3.7 Flash, Puzzle-75B. $4,699.
Want the budget version? Strix Halo. 128GB unified, $2.5K. Not recommended
Want frontier scale on your desk? 2 Sparks. Qwen3-235B, MiniMax-M3. $9,500.
Want maximum dense speed? RTX PRO 6000 Blackwell. $12K+.
Don't buy a Spark to run dense models fast.
Don't buy a 3090 to run 235B. Know which axis you're upgrading.
Reference links in reply for your own technical DD 👇