Skip to main content
2026 open-weight model registry

Local AI & LLM RAM Sizing Hub

Planning to run LLMs locally? Browse 96 open-weight models with verified parameter counts. Each calculator supports inference vs training modes, GPU / CPU-offload / Mac paths, and a kit compatibility gate before you buy.

Official-source overrides for flagshipsGGUF Q4 / Q8 / FP16 estimatesLive Amazon kit picks

Tracked Models

96

Quantizations

4-bit, 8-bit, FP16

Key Providers

Kimi, DeepSeek, Meta, Qwen

Popular kits for local AI

In-stock value picks at the capacities most local LLM builds actually buy.

Disclosure: As an Amazon Associate I earn from qualifying purchases. Rankings use price and spec data only — not paid placement. How we rank products

A-Tech 64GB (2x32GB) DDR4 2666 MHz UDIMM PC4-21300 (PC4-2666V) CL19 DIMM 2Rx8 Non-ECC Desktop RAM Memory Modules

$426.78$6.67/GB

CORSAIR DOMINATOR PLATINUM RGB DDR5 RAM 64GB (2x32GB) 5600MHz CL40 Intel XMP iCUE Compatible Computer Memory - White (CMT64GX5M2B5600C40W)

$969.99$15.16/GB

CORSAIR Vengeance RGB DDR5 RAM 96GB (2x48GB) 6000MHz CL30 Intel XMP iCUE Compatible Computer Memory - Black (CMH96GX5M2B6000C30)

$189.99$1.98/GB

A-Tech 96GB Kit (2x48GB) DDR5 5600MHz PC5-44800 CL46 SODIMM 2Rx8 Dual Rank 1.1V Non-ECC Unbuffered SO-DIMM 262-Pin Laptop Computer RAM Memory Upgrade Modules

$1530.87$15.95/GB

128GB kits

All prices →

A-Tech 128GB Kit (4x32GB) RAM for Apple iMac 2019 & 2020 27 inch Retina 5K | DDR4 2666 MHz SODIMM PC4-21300 / PC4-21333 260-Pin SO-DIMM Max Memory Upgrade

$825.02$6.45/GB

A-Tech 128GB Kit (4x32GB) DDR5 4800MHz PC5-38400 CL40 SODIMM 2Rx8 Dual Rank 1.1V Non-ECC Unbuffered SO-DIMM 262-Pin Laptop Computer RAM Memory Upgrade Modules

$1402.16$10.95/GB

Showing 96 of 96 open-weights models

Moonshot Kimi
MoEWeights pending

Kimi K3 (2.8T MoE)

Total Params:2800B
Active Params:50B est.
Context Window:1M tokens
Release:July 2026

Moonshot's open 2.8T-parameter frontier MoE with native vision and a 1M-token context. Official blog: activates 16 of 896 experts (Stable LatentMoE); API live now; full weights expected by July 27, 2026. Active params (~50B) estimated as 2800×16/896 until the technical report publishes an exact figure. RAM sizing uses total parameters for weight footprint.

Size RAM & shop kits
DeepSeek
MoE

DeepSeek-V4-Pro (1.6T MoE)

Total Params:1600B
Active Params:49B
Context Window:1M tokens
Release:April 2026

Flagship open reasoning MoE: 1.6T total / 49B active with Compressed Sparse Attention for long-context efficiency. RAM estimates assume GGUF-style quantization; native FP4/FP8 footprints can differ.

Size RAM & shop kits
Moonshot Kimi
MoE

Kimi K2 0905 (1T MoE)

Total Params:1000B
Active Params:32B
Context Window:262K tokens
Release:September 2025

Kimi K2 0905 checkpoint: 1T MoE with 32B active parameters and extended 256K context.

Size RAM & shop kits
Moonshot Kimi
MoE

Kimi K2 (1T MoE)

Total Params:1000B
Active Params:32B
Context Window:131K tokens
Release:July 2025

Original Kimi K2 open-weight MoE: ~1T total parameters with 32B activated per token. Foundation for the later K2.5 / K2.6 / K2.7-Code family.

Size RAM & shop kits
Moonshot Kimi
MoE

Kimi K2 Thinking (1T MoE)

Total Params:1000B
Active Params:32B
Context Window:262K tokens
Release:November 2025

Kimi K2 Thinking reasoning variant on the K2 MoE stack: 1T total / 32B active with long-horizon agentic reasoning.

Size RAM & shop kits
Moonshot Kimi
MoE

Kimi K2.5 (1T MoE)

Total Params:1000B
Active Params:32B
Context Window:262K tokens
Release:January 2026

Moonshot multimodal MoE: 1T total / 32B active (384 experts, 8 selected + 1 shared), 256K context, MoonViT vision encoder. Open weights on Hugging Face.

Size RAM & shop kits
Moonshot Kimi
MoE

Kimi K2.6 (1T MoE)

Total Params:1000B
Active Params:32B
Context Window:262K tokens
Release:March 2026

Kimi K2.6 agentic coding and planning flagship: 1T MoE with 32B active parameters and 256K context.

Size RAM & shop kits
Moonshot Kimi
MoE

Kimi K2.7 Code (1T MoE)

Total Params:1000B
Active Params:32B
Context Window:262K tokens
Release:June 2026

Coding-focused Kimi K2.7 open-weight MoE (1T / 32B active, 256K context). Tuned for long-horizon software engineering with ~30% fewer thinking tokens vs K2.6.

Size RAM & shop kits
Zhipu GLM
MoE

GLM-5.2 (744B MoE)

Total Params:744B
Active Params:40B
Context Window:1M tokens
Release:June 2026

Z.ai / Zhipu flagship open-weight MoE (744B-A40B) with 1M context and IndexShare sparse attention. MIT license; optimized for long-horizon agentic coding.

Size RAM & shop kits
Zhipu GLM
MoE

GLM-5 (744B MoE)

Total Params:744B
Active Params:40B
Context Window:256K tokens
Release:February 2026

Zhipu AI open-weights MoE flagship (744B total / 40B active). Strong general reasoning and multi-turn planning under MIT license.

Size RAM & shop kits
DeepSeek
MoE

DeepSeek V3 0324

Total Params:685B
Active Params:85.6B
Context Window:164K tokens
Release:2025/2026

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...

Size RAM & shop kits
DeepSeek
Dense

DeepSeek V3.1

Total Params:671B
Active Params:Dense
Context Window:164K tokens
Release:2025/2026

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

Size RAM & shop kits
DeepSeek
MoE

DeepSeek-R1-0528 (671B MoE)

Total Params:671B
Active Params:37B
Context Window:128K tokens
Release:May 2025

May 2025 DeepSeek-R1 refresh: same 671B MoE / 37B active architecture with updated post-training.

Size RAM & shop kits
DeepSeek
MoE

DeepSeek-R1 (671B MoE)

Total Params:671B
Active Params:37B
Context Window:128K tokens
Release:January 2025

DeepSeek-R1 reasoning MoE built on DeepSeek-V3: 671B total / 37B activated per token, 128K context. Open weights on Hugging Face.

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 Coder 480B A35B

Total Params:480B
Active Params:35B
Context Window:262K tokens
Release:2025/2026

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

Size RAM & shop kits
MiniMax
Dense

MiniMax-01

Total Params:456B
Active Params:Dense
Context Window:1M tokens
Release:2025/2026

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context...

Size RAM & shop kits
MiniMax
MoE

MiniMax-M3 (428B MoE)

Total Params:428B
Active Params:23B
Context Window:1M tokens
Release:June 2026

Native multimodal MiniMax MoE: ~428B total / ~23B active, 1M context, MiniMax Sparse Attention (MSA). Open weights on Hugging Face.

Size RAM & shop kits
Nous Research
Dense

Hermes 4 405B

Total Params:405B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...

Size RAM & shop kits
Nous Research
Dense

Hermes 3 405B Instruct

Total Params:405B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

Size RAM & shop kits
Meta Llama
MoE

Llama 4 Maverick (400B MoE)

Total Params:400B
Active Params:17B
Context Window:1M tokens
Release:April 2025

Meta Llama 4 Maverick open-weight MoE: 400B total / 17B active with long-context multimodal capability.

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3.5 397B A17B

Total Params:397B
Active Params:17B
Context Window:262K tokens
Release:2025/2026

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

Size RAM & shop kits
DeepSeek
MoE

DeepSeek-V4-Flash (284B MoE)

Total Params:284B
Active Params:13B
Context Window:1M tokens
Release:April 2026

High-efficiency DeepSeek-V4 MoE variant: 284B total / 13B active with 1M context for lower-latency local inference.

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 VL 235B A22B Thinking

Total Params:235B
Active Params:22B
Context Window:131K tokens
Release:2025/2026

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 VL 235B A22B Instruct

Total Params:235B
Active Params:22B
Context Window:262K tokens
Release:2025/2026

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 235B A22B Thinking 2507

Total Params:235B
Active Params:22B
Context Window:262K tokens
Release:2025/2026

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 235B A22B Instruct 2507

Total Params:235B
Active Params:22B
Context Window:262K tokens
Release:2025/2026

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 235B A22B

Total Params:235B
Active Params:22B
Context Window:131K tokens
Release:2025/2026

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

Size RAM & shop kits
Mistral AI
MoE

Mixtral 8x22B Instruct

Total Params:176B
Active Params:44B
Context Window:66K tokens
Release:2025/2026

Mistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b). It uses 39B active parameters out of 141B, offering unparalleled cost efficiency for its size. Its strengths include: - strong math, coding,...

Size RAM & shop kits
Microsoft
MoE

WizardLM-2 8x22B

Total Params:176B
Active Params:44B
Context Window:66K tokens
Release:2025/2026

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary models, and it consistently outperforms all existing state-of-the-art opensource models. It is...

Size RAM & shop kits
Mistral AI
Dense

Mistral Medium 3.5

Total Params:128B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Size RAM & shop kits
Mistral AI
Dense

Devstral 2 2512

Total Params:123B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3.5-122B-A10B

Total Params:122B
Active Params:10B
Context Window:262K tokens
Release:2025/2026

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...

Size RAM & shop kits
Mistral AI
MoE

Mistral Small 4 (119B MoE)

Total Params:119B
Active Params:6.5B
Context Window:256K tokens
Release:March 2026

Mistral Small 4 production MoE unifying instruction following, multimodal inputs, and agentic workflows with a low active footprint.

Size RAM & shop kits
Cohere
Dense

Command A

Total Params:111B
Active Params:Dense
Context Window:256K tokens
Release:2025/2026

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...

Size RAM & shop kits
Meta Llama
MoE

Llama 4 Scout (109B MoE)

Total Params:109B
Active Params:17B
Context Window:10M tokens
Release:April 2025

Meta Llama 4 Scout: 109B MoE / 17B active with a native 10M-token context window for long-document workloads.

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 Coder Next

Total Params:80B
Active Params:10B
Context Window:262K tokens
Release:2025/2026

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 Next 80B A3B Thinking

Total Params:80B
Active Params:3B
Context Window:262K tokens
Release:2025/2026

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 Next 80B A3B Instruct

Total Params:80B
Active Params:3B
Context Window:262K tokens
Release:2025/2026

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen2.5 VL 72B Instruct

Total Params:72B
Active Params:Dense
Context Window:128K tokens
Release:2025/2026

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen2.5 72B Instruct

Total Params:72B
Active Params:Dense
Context Window:33K tokens
Release:2025/2026

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

Size RAM & shop kits
Nous Research
Dense

Hermes 4 70B

Total Params:70B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either...

Size RAM & shop kits
DeepSeek
Dense

R1 Distill Llama 70B

Total Params:70B
Active Params:Dense
Context Window:8K tokens
Release:2025/2026

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

Size RAM & shop kits
Meta Llama
Dense

Llama 3.3 70B Instruct

Total Params:70B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

Size RAM & shop kits
Nous Research
Dense

Hermes 3 70B Instruct

Total Params:70B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

Size RAM & shop kits
Meta Llama
Dense

Llama 3.1 70B Instruct

Total Params:70B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...

Size RAM & shop kits
Mistral AI
MoE

Mistral Large 3 2512

Total Params:41B
Active Params:5.1B
Context Window:262K tokens
Release:2025/2026

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3.6 35B A3B

Total Params:35B
Active Params:3B
Context Window:262K tokens
Release:2025/2026

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3.5-35B-A3B

Total Params:35B
Active Params:3B
Context Window:262K tokens
Release:2025/2026

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen 3.6 35B-A3B (MoE)

Total Params:35B
Active Params:3B
Context Window:128K tokens
Release:April 2026

Qwen 3.6 sparse MoE activating 3B parameters per token for efficient coding throughput.

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen3 VL 32B Instruct

Total Params:32B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen3 32B

Total Params:32B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen2.5 Coder 32B Instruct

Total Params:32B
Active Params:Dense
Context Window:33K tokens
Release:2025/2026

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reasoning**...

Size RAM & shop kits
Google Gemma
Dense

Gemma 4 31B

Total Params:31B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Size RAM & shop kits
Google Gemma
Dense

Gemma 4 31B (Dense)

Total Params:31B
Active Params:Dense
Context Window:131K tokens
Release:April 2026

Google DeepMind Gemma 4 31B dense flagship for single-GPU / high-RAM consumer local inference.

Size RAM & shop kits
Cohere
MoE

North Mini Code (free)

Total Params:30B
Active Params:3.8B
Context Window:256K tokens
Release:2025/2026

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 VL 30B A3B Thinking

Total Params:30B
Active Params:3B
Context Window:262K tokens
Release:2025/2026

Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 VL 30B A3B Instruct

Total Params:30B
Active Params:3B
Context Window:262K tokens
Release:2025/2026

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 30B A3B Thinking 2507

Total Params:30B
Active Params:3B
Context Window:82K tokens
Release:2025/2026

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 Coder 30B A3B Instruct

Total Params:30B
Active Params:3B
Context Window:262K tokens
Release:2025/2026

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 30B A3B Instruct 2507

Total Params:30B
Active Params:3B
Context Window:262K tokens
Release:2025/2026

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...

Size RAM & shop kits
Alibaba Qwen
MoE

Qwen3 30B A3B

Total Params:30B
Active Params:3B
Context Window:131K tokens
Release:2025/2026

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen3.6 27B

Total Params:27B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen3.5-27B

Total Params:27B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...

Size RAM & shop kits
Google Gemma
Dense

Gemma 3 27B

Total Params:27B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

Size RAM & shop kits
Google Gemma
Dense

Gemma 2 27B

Total Params:27B
Active Params:Dense
Context Window:8K tokens
Release:2025/2026

Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini). Gemma models are well-suited for a variety of...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen 3.6 27B (Dense)

Total Params:27B
Active Params:Dense
Context Window:128K tokens
Release:April 2026

Alibaba Qwen 3.6 dense 27B developer flagship for multilingual reasoning and structured local workflows.

Size RAM & shop kits
Google Gemma
MoE

Gemma 4 26B A4B

Total Params:26B
Active Params:4B
Context Window:262K tokens
Release:2025/2026

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Size RAM & shop kits
Google Gemma
MoE

Gemma 4 26B (MoE)

Total Params:26B
Active Params:3.8B
Context Window:131K tokens
Release:April 2026

Ultra-efficient Gemma 4 sparse MoE activating ~3.8B parameters per token for fast local inference.

Size RAM & shop kits
Mistral AI
Dense

Voxtral Small 24B 2507

Total Params:24B
Active Params:Dense
Context Window:32K tokens
Release:2025/2026

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding. Input audio...

Size RAM & shop kits
Cognitive Computations
Dense

Uncensored

Total Params:24B
Active Params:Dense
Context Window:128K tokens
Release:2025/2026

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...

Size RAM & shop kits
Mistral AI
Dense

Mistral Small 3.2 24B

Total Params:24B
Active Params:Dense
Context Window:256K tokens
Release:2025/2026

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...

Size RAM & shop kits
Mistral AI
Dense

Mistral Small 3.1 24B

Total Params:24B
Active Params:Dense
Context Window:128K tokens
Release:2025/2026

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities. It provides state-of-the-art performance in text-based reasoning and...

Size RAM & shop kits
Mistral AI
Dense

Saba

Total Params:24B
Active Params:Dense
Context Window:33K tokens
Release:2025/2026

Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually relevant responses while maintaining efficient performance. Trained on curated regional...

Size RAM & shop kits
Mistral AI
Dense

Mistral Small 3

Total Params:24B
Active Params:Dense
Context Window:33K tokens
Release:2025/2026

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed...

Size RAM & shop kits
Mistral AI
Dense

Ministral 3 14B 2512

Total Params:14B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen3 14B

Total Params:14B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

Size RAM & shop kits
Microsoft
Dense

Phi 4

Total Params:14B
Active Params:Dense
Context Window:16K tokens
Release:2025/2026

[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...

Size RAM & shop kits
Meta Llama
Dense

Llama Guard 4 12B

Total Params:12B
Active Params:Dense
Context Window:1M tokens
Release:2025/2026

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...

Size RAM & shop kits
Google Gemma
Dense

Gemma 3 12B

Total Params:12B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

Size RAM & shop kits
Mistral AI
Dense

Mistral Nemo

Total Params:12B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,...

Size RAM & shop kits
MiniMax
Dense

MiniMax M2.1

Total Params:10B
Active Params:Dense
Context Window:205K tokens
Release:2025/2026

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...

Size RAM & shop kits
MiniMax
Dense

MiniMax M2

Total Params:10B
Active Params:Dense
Context Window:205K tokens
Release:2025/2026

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen3.5-9B

Total Params:9B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

Size RAM & shop kits
Mistral AI
Dense

Ministral 3 8B 2512

Total Params:8B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen3 VL 8B Thinking

Total Params:8B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across complex scenes, documents, and temporal sequences. It integrates enhanced multimodal alignment and...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen3 VL 8B Instruct

Total Params:8B
Active Params:Dense
Context Window:262K tokens
Release:2025/2026

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen3 8B

Total Params:8B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...

Size RAM & shop kits
Meta Llama
Dense

Llama 3.1 8B Instruct

Total Params:8B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

Size RAM & shop kits
Cohere
Dense

Command R7B (12-2024)

Total Params:7B
Active Params:Dense
Context Window:128K tokens
Release:2025/2026

Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiring complex reasoning...

Size RAM & shop kits
Alibaba Qwen
Dense

Qwen2.5 7B Instruct

Total Params:7B
Active Params:Dense
Context Window:33K tokens
Release:2025/2026

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

Size RAM & shop kits
Google Gemma
Dense

Gemma 3n 4B

Total Params:4B
Active Params:Dense
Context Window:33K tokens
Release:2025/2026

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets. It supports multimodal inputs—including text, visual data, and audio—enabling diverse tasks...

Size RAM & shop kits
Google Gemma
Dense

Gemma 3 4B

Total Params:4B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

Size RAM & shop kits
Microsoft
Dense

Phi-4-mini (3.8B)

Total Params:3.8B
Active Params:Dense
Context Window:128K tokens
Release:February 2025

Microsoft Phi-4-mini dense model for fast on-device and low-RAM local text processing.

Size RAM & shop kits
Mistral AI
Dense

Ministral 3 3B 2512

Total Params:3B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

Size RAM & shop kits
Meta Llama
Dense

Llama 3.2 3B Instruct

Total Params:3B
Active Params:Dense
Context Window:131K tokens
Release:2025/2026

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...

Size RAM & shop kits
Meta Llama
Dense

Llama 3.2 1B Instruct

Total Params:1B
Active Params:Dense
Context Window:60K tokens
Release:2025/2026

Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis. Its smaller size allows it to operate...

Size RAM & shop kits

Understanding Local AI Memory Requirements

⚖️

Model Weights Size

The parameters of the model determine its baseline RAM/VRAM footprint. Every billion parameters requires 2GB in standard 16-bit float (`FP16`). Quantizing parameters to 4-bit (`Q4_K_M`) compresses this to ~0.56GB per billion, trading a tiny fraction of accuracy for a 72% memory savings.

Context Window KV Cache

As the context window scales, the key-value cache (KV cache) expands. Running massive 1M token windows (like DeepSeek V4) requires enormous RAM allocations strictly for context memory. Our calculators use model-specific parameters to compute this dynamically.

🚀

Memory Bandwidth Sizing

Because local LLM generation requires loading weights from RAM on every single step, generation speeds scale directly with memory bandwidth. Mainstream DDR5 configurations yield far superior tokens/s compared to DDR4, making memory frequency critical for local AI speed.

Frequently Asked Questions

How do you calculate RAM size for a local LLM?

The physical RAM requirement is the sum of three components: 1. **Model Weights Size**: calculated as `(Parameter Count * Bits per Weight) / 8` (MoE architectures require the total parameter weights loaded, even if only a subset are active per token). 2. **Context KV Cache Size**: determined by active parameters and target context length (models utilizing Grouped-Query Attention drastically reduce this footprint). 3. **OS Overhead**: typically 6GB to 12GB allocation depending on workstation or multi-GPU configurations.

Why do Mixture-of-Experts (MoE) models require so much RAM?

While Mixture-of-Experts models (like DeepSeek V4 or Llama 4 Maverick) only activate a small number of parameters per token during calculation (which keeps compute costs low), **the entire weights database of all experts must reside in physical RAM/VRAM** for fast expert switching during inference. Therefore, RAM sizing calculations must target the *total* parameters rather than just the *active* ones.

Can I run a 70B model on 32GB of RAM?

Not at full unquantized 16-bit weight precision (which requires ~140GB). However, you can run a 70B model at 4-bit quantization (e.g., `Q4_K_M`) which compresses weights down to ~39GB. When factoring in system overhead and context, a **64GB system memory kit** is required to run a 4-bit 70B model stably without out-of-memory crashes.

What is the speed bottleneck for running local LLMs on CPU/RAM?

The primary bottleneck is **system memory bandwidth**, not CPU cores. Large language models are highly memory-bandwidth bound. A typical dual-channel DDR5-6000 configuration achieves ~96 GB/s, generating tokens at around 2.4 tokens/s for a 40GB model. Upgrading to high-speed dual-channel CUDIMM (DDR5-8400 yields ~134 GB/s) or quad-channel workstation layouts is the best way to directly scale local inference speeds.