Why Qwen 3.6 35B-A3B (MoE) pressures system RAM
Qwen 3.6 35B-A3B (MoE) is Mixture-of-Experts: inference activates 3B active/token, but VRAM/RAM must usually hold the full ~35B expert set for fast routing. At Q4 the weight slab is ~19.7GB before KV (~0.01GB at 8K) and ~6GB OS/runtime overhead — totaling ~25.7GB raw, rounded to a 32GB kit. Stretching toward the full 128K-token window multiplies KV far faster than weights; that is the usual “I bought enough RAM for the model but still OOM” failure on Alibaba Qwen MoE pages.
What RAM kit to buy
A 32GB dual-channel kit is enough for quantized Qwen 3.6 35B-A3B (MoE) at modest context. Still prefer 2× matched SO-DIMM/UDIMM sticks; 1x RTX 3090 or RTX 4090 (24GB VRAM) covers the Flagship Consumer GPU GPU profile. If you chat with long pastes, jump a tier before the KV cache forces paging.
Workload notes
Qwen-family models like Qwen 3.6 35B-A3B (MoE) often ship strong coding/agent variants; leave RAM for tool runners and browser IDEs beside the weights. At 35B, Qwen 3.6 35B-A3B (MoE) is a practical mid-size local model — sweet spot for single-GPU Q4/Q8 experimenters who still want headroom for IDE + Docker. Release window noted as April 2026; always re-check the official source before buying hardware for a specific checkpoint.






