LLM VRAM Requirements by Model & Quantization

Compare source-backed FP16, 8-bit, and 4-bit VRAM planning targets for the same LLM profiles used by the ComputeAtlas AI Hardware Estimator. The table separates published model facts from ComputeAtlas planning headroom so you can see the weight floor, the safer planning target, and the hardware tier each model reaches.

6 current-generation LLM profiles · 10 reference profiles · source verification current through 2026-10-06

Decision visualPrecision changes the model-weight floor before runtime overhead
4-bitlower weight footprint
8-bitmiddle planning tier
FP16highest weight footprint
Conceptual relationship only. ComputeAtlas uses the canonical workload source of truth for model-specific values, runtime reserve, context, batching, and diffusion-pipeline handling.

Independence & evidence

Planning logic first, commercial links second

This reference is generated from the same governed model registry used by the ComputeAtlas estimator. Model-fit tables and calculator outputs are not ranked or adjusted based on retailer or affiliate compensation. Published model facts are tied to first-party sources, while ComputeAtlas-derived values are labeled as planning guidance rather than measured benchmark results.

Review the methodology and affiliate disclosure.

What These Numbers Mean

The values below are standard inference planning targets, not claims that every runtime will consume exactly that amount of memory. For dense and MoE language models, ComputeAtlas first calculates a model-weight floor from the published resident parameter count: approximately 2 bytes per parameter at FP16, 1 byte at 8-bit, and 0.5 byte at 4-bit. A runtime reserve is then included for practical planning. Light workloads are never allowed to fall below the model-weight floor. Heavy workloads add headroom for context, batching, concurrency, KV cache, and framework overhead.

The estimator's Fine-Tuning option represents parameter-efficient LoRA/QLoRA-style adaptation; it is not a full-parameter training-memory estimate. Multiple GPUs also do not automatically create one transparent memory pool: a runtime must explicitly shard or place model state across devices before aggregate VRAM becomes useful to one workload.

4-bit Planning Targets by VRAM Tier

This quick-fit view includes current-generation models only and shows the standard 4-bit inference planning target, not a promise that every context length, runtime, or concurrent workload will fit. Longer context, larger KV cache, batching, and framework overhead can require more memory.

≤ 16 GB VRAM

Gemma 4 26B-A4B, gpt-oss-20b

≤ 24 GB VRAM

Qwen3.8-27B, Gemma 4 31B, Gemma 4 26B-A4B, gpt-oss-20b

≤ 32 GB VRAM

Qwen3.8-27B, Gemma 4 31B, Gemma 4 26B-A4B, gpt-oss-20b

≤ 48 GB VRAM

Qwen3.8-27B, Gemma 4 31B, Gemma 4 26B-A4B, gpt-oss-20b

≤ 80 GB VRAM

Qwen3.8-27B, Gemma 4 31B, Gemma 4 26B-A4B, Mistral Small 4 119B-A6B, gpt-oss-20b, gpt-oss-120b

Full LLM Inference Planning Reference

Side-by-side model VRAM requirements table
ModelStatusFP16 planning8-bit planning4-bit planningFP16 model-weight floorTypical planning setupCalculator
Qwen3.8-27BCurrent generation60 GB32 GB17 GB55.562855904 GB1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bitOpen 4-bit estimate
Gemma 4 31BCurrent generation69 GB36 GB19 GB66 GB1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bitOpen 4-bit estimate
Gemma 4 26B-A4BCurrent generation56 GB29 GB16 GB54 GB1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 16–24 GB-class hardware at 4-bitOpen 4-bit estimate
Mistral Small 4 119B-A6BCurrent generation262 GB137 GB72 GB238 GB1× 80 GB-class GPU at 4-bit; higher-precision deployment requires explicit multi-GPU sharding or offloadOpen 4-bit estimate
gpt-oss-20bCurrent generation47 GB24 GB16 GB41.82 GB1× 16–24 GB-class GPU for the native MXFP4 checkpoint; higher-precision conversions require more memoryOpen 4-bit estimate
gpt-oss-120bCurrent generation258 GB135 GB80 GB233.66 GB1× 80 GB-class GPU for the native MXFP4 checkpoint; higher-precision representations move into multi-GPU territoryOpen 4-bit estimate
Llama 3 8BReference18 GB10 GB5 GB16 GB1× 24 GB-class GPU for comfortable local inference headroomOpen 4-bit estimate
Llama 3 70BReference154 GB81 GB42 GB140 GB2× 80 GB-class GPUs for FP16; quantization can materially reduce the memory floorOpen 4-bit estimate
Mixtral 8x7BReference104 GB55 GB29 GB94 GB2× 80 GB-class GPUs at FP16, or a smaller footprint with quantizationOpen 4-bit estimate
DeepSeek LLM 67BReference148 GB78 GB41 GB134 GB2× 80 GB-class GPUs for FP16; quantization can materially reduce the memory floorOpen 4-bit estimate
Qwen3 8BReference18 GB10 GB5 GB16.4 GB1× 16–24 GB-class GPU for local inference, with more headroom for long contexts or concurrencyOpen 4-bit estimate
Qwen3 32BReference73 GB38 GB20 GB65.6 GB1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bitOpen 4-bit estimate
Mistral Small 3.1 24BReference55 GB28 GB15 GB48 GB1× 80 GB-class GPU at FP16 or 1× 24–48 GB-class GPU when quantizedOpen 4-bit estimate
Gemma 3 27BReference60 GB32 GB17 GB54 GB1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bitOpen 4-bit estimate
DeepSeek-R1-Distill-Qwen-32BReference71 GB37 GB20 GB64 GB1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bitOpen 4-bit estimate
DeepSeek-R1 671BReference1477 GB772 GB403 GB1342 GBMulti-GPU server or cluster deployment; even 4-bit planning exceeds a single high-memory GPUOpen 4-bit estimate

Model Architecture Notes

  • Qwen3.8-27B: Dense multimodal model with a 27B language model and 262K native context. The official checkpoint metadata reports 27.781427952B total resident parameters including multimodal weights, which ComputeAtlas uses for the raw model-weight floor.
  • Gemma 4 31B: Dense Gemma 4 multimodal model. The model card reports a 30.7B text model plus a vision encoder, while the published safetensors checkpoint is reported as 33B parameters; ComputeAtlas uses the full checkpoint size for the resident-weight floor.
  • Gemma 4 26B-A4B: Sparse Gemma 4 MoE model with a 25.2B text MoE, about 3.8B active parameters, and a vision encoder. The published safetensors checkpoint is reported as 27B parameters, which ComputeAtlas uses for the resident-weight floor.
  • Mistral Small 4 119B-A6B: Sparse MoE model with 119B resident parameters, about 6.5B active parameters per token, and 256K context. ComputeAtlas plans memory from resident weights; active parameters describe compute sparsity, not the checkpoint placement requirement.
  • gpt-oss-20b: Sparse MoE reasoning model with 20.91B total parameters and about 3.61B active parameters per token. OpenAI distributes the model in MXFP4 and documents local operation in about 16 GB of memory; ComputeAtlas preserves that published floor for the 4-bit planning reference.
  • gpt-oss-120b: Sparse MoE reasoning model with 116.83B total parameters and about 5.13B active parameters per token. OpenAI distributes the model in MXFP4 and documents operation on a single 80 GB GPU, so ComputeAtlas uses 80 GB as the conservative 4-bit planning reference.
  • Llama 3 8B: Dense 8B model. ComputeAtlas derives the model-weight floor from the published parameter count, then adds a runtime planning reserve.
  • Llama 3 70B: Dense 70B model. The full checkpoint must be resident unless the runtime explicitly offloads or shards it.
  • Mixtral 8x7B: Sparse MoE model with about 47B resident parameters and about 13B active parameters per token. Memory planning uses resident weights, not only active parameters.
  • DeepSeek LLM 67B: Dense 67B model. ComputeAtlas uses the published 67B parameter count for the model-weight floor before runtime reserve.
  • Qwen3 8B: Dense 8.2B model with hybrid thinking/non-thinking operation. ComputeAtlas uses the published total parameter count for the model-weight floor and keeps long-context KV/cache costs outside the raw-weight calculation.
  • Qwen3 32B: Dense 32.8B model. ComputeAtlas derives the model-weight floor from the published total parameter count, then adds planning headroom for runtime state, context, and framework overhead.
  • Mistral Small 3.1 24B: Dense 24B multimodal model with up to 128K context. Mistral documents about 55 GB of GPU RAM for BF16/FP16 inference, so ComputeAtlas anchors the FP16 planning target to that published deployment guidance.
  • Gemma 3 27B: Dense 27B multimodal model with a 128K context window. ComputeAtlas uses the published parameter count for raw model-weight floors and treats context/image/runtime memory as additional planning overhead.
  • DeepSeek-R1-Distill-Qwen-32B: Dense 32B distilled reasoning model based on Qwen2.5-32B. ComputeAtlas plans from the full resident checkpoint size rather than the larger teacher model used to generate reasoning traces.
  • DeepSeek-R1 671B: Sparse MoE reasoning model with 671B total resident parameters and about 37B activated parameters per token. Memory planning uses resident weights, not only activated parameters, because the full checkpoint must be placed across the deployment unless the runtime offloads or shards it.

Quantization and Runtime Effects

  • FP16: Approximately 2 bytes per resident parameter for the language-model weight floor, before runtime allocations.
  • 8-bit: Approximately 1 byte per resident parameter for the language-model weight floor; implementation metadata and runtime allocations still consume additional memory.
  • 4-bit: Approximately 0.5 byte per resident parameter for the theoretical language-model weight floor; practical quantized formats carry additional metadata and runtime overhead.
  • Context and concurrency: KV cache, batch size, concurrent sessions, framework allocations, and temporary buffers can materially raise peak VRAM beyond model weights.

Source Provenance

Parameter counts and architecture facts are grounded in first-party model documentation. ComputeAtlas-derived planning targets are deliberately identified as planning guidance rather than manufacturer-measured peak VRAM. Source verification is current through 2026-10-06.

Plan Your Build with ComputeAtlas

Apply the shared reference to your model, precision, workload, and adaptation goal.

Estimate Hardware for Your Model