≤ 16 GB VRAM
Gemma 4 26B-A4B, gpt-oss-20b
Compare source-backed FP16, 8-bit, and 4-bit VRAM planning targets for the same LLM profiles used by the ComputeAtlas AI Hardware Estimator. The table separates published model facts from ComputeAtlas planning headroom so you can see the weight floor, the safer planning target, and the hardware tier each model reaches.
6 current-generation LLM profiles · 10 reference profiles · source verification current through 2026-10-06
Independence & evidence
This reference is generated from the same governed model registry used by the ComputeAtlas estimator. Model-fit tables and calculator outputs are not ranked or adjusted based on retailer or affiliate compensation. Published model facts are tied to first-party sources, while ComputeAtlas-derived values are labeled as planning guidance rather than measured benchmark results.
Review the methodology and affiliate disclosure.
The values below are standard inference planning targets, not claims that every runtime will consume exactly that amount of memory. For dense and MoE language models, ComputeAtlas first calculates a model-weight floor from the published resident parameter count: approximately 2 bytes per parameter at FP16, 1 byte at 8-bit, and 0.5 byte at 4-bit. A runtime reserve is then included for practical planning. Light workloads are never allowed to fall below the model-weight floor. Heavy workloads add headroom for context, batching, concurrency, KV cache, and framework overhead.
The estimator's Fine-Tuning option represents parameter-efficient LoRA/QLoRA-style adaptation; it is not a full-parameter training-memory estimate. Multiple GPUs also do not automatically create one transparent memory pool: a runtime must explicitly shard or place model state across devices before aggregate VRAM becomes useful to one workload.
This quick-fit view includes current-generation models only and shows the standard 4-bit inference planning target, not a promise that every context length, runtime, or concurrent workload will fit. Longer context, larger KV cache, batching, and framework overhead can require more memory.
Gemma 4 26B-A4B, gpt-oss-20b
Qwen3.8-27B, Gemma 4 31B, Gemma 4 26B-A4B, gpt-oss-20b
Qwen3.8-27B, Gemma 4 31B, Gemma 4 26B-A4B, gpt-oss-20b
Qwen3.8-27B, Gemma 4 31B, Gemma 4 26B-A4B, gpt-oss-20b
Qwen3.8-27B, Gemma 4 31B, Gemma 4 26B-A4B, Mistral Small 4 119B-A6B, gpt-oss-20b, gpt-oss-120b
| Model | Status | FP16 planning | 8-bit planning | 4-bit planning | FP16 model-weight floor | Typical planning setup | Calculator |
|---|---|---|---|---|---|---|---|
| Qwen3.8-27B | Current generation | 60 GB | 32 GB | 17 GB | 55.562855904 GB | 1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bit | Open 4-bit estimate |
| Gemma 4 31B | Current generation | 69 GB | 36 GB | 19 GB | 66 GB | 1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bit | Open 4-bit estimate |
| Gemma 4 26B-A4B | Current generation | 56 GB | 29 GB | 16 GB | 54 GB | 1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 16–24 GB-class hardware at 4-bit | Open 4-bit estimate |
| Mistral Small 4 119B-A6B | Current generation | 262 GB | 137 GB | 72 GB | 238 GB | 1× 80 GB-class GPU at 4-bit; higher-precision deployment requires explicit multi-GPU sharding or offload | Open 4-bit estimate |
| gpt-oss-20b | Current generation | 47 GB | 24 GB | 16 GB | 41.82 GB | 1× 16–24 GB-class GPU for the native MXFP4 checkpoint; higher-precision conversions require more memory | Open 4-bit estimate |
| gpt-oss-120b | Current generation | 258 GB | 135 GB | 80 GB | 233.66 GB | 1× 80 GB-class GPU for the native MXFP4 checkpoint; higher-precision representations move into multi-GPU territory | Open 4-bit estimate |
| Llama 3 8B | Reference | 18 GB | 10 GB | 5 GB | 16 GB | 1× 24 GB-class GPU for comfortable local inference headroom | Open 4-bit estimate |
| Llama 3 70B | Reference | 154 GB | 81 GB | 42 GB | 140 GB | 2× 80 GB-class GPUs for FP16; quantization can materially reduce the memory floor | Open 4-bit estimate |
| Mixtral 8x7B | Reference | 104 GB | 55 GB | 29 GB | 94 GB | 2× 80 GB-class GPUs at FP16, or a smaller footprint with quantization | Open 4-bit estimate |
| DeepSeek LLM 67B | Reference | 148 GB | 78 GB | 41 GB | 134 GB | 2× 80 GB-class GPUs for FP16; quantization can materially reduce the memory floor | Open 4-bit estimate |
| Qwen3 8B | Reference | 18 GB | 10 GB | 5 GB | 16.4 GB | 1× 16–24 GB-class GPU for local inference, with more headroom for long contexts or concurrency | Open 4-bit estimate |
| Qwen3 32B | Reference | 73 GB | 38 GB | 20 GB | 65.6 GB | 1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bit | Open 4-bit estimate |
| Mistral Small 3.1 24B | Reference | 55 GB | 28 GB | 15 GB | 48 GB | 1× 80 GB-class GPU at FP16 or 1× 24–48 GB-class GPU when quantized | Open 4-bit estimate |
| Gemma 3 27B | Reference | 60 GB | 32 GB | 17 GB | 54 GB | 1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bit | Open 4-bit estimate |
| DeepSeek-R1-Distill-Qwen-32B | Reference | 71 GB | 37 GB | 20 GB | 64 GB | 1× 80 GB-class GPU at FP16, 1× 48 GB-class GPU at 8-bit, or 24 GB-class hardware at 4-bit | Open 4-bit estimate |
| DeepSeek-R1 671B | Reference | 1477 GB | 772 GB | 403 GB | 1342 GB | Multi-GPU server or cluster deployment; even 4-bit planning exceeds a single high-memory GPU | Open 4-bit estimate |
Parameter counts and architecture facts are grounded in first-party model documentation. ComputeAtlas-derived planning targets are deliberately identified as planning guidance rather than manufacturer-measured peak VRAM. Source verification is current through 2026-10-06.
Apply the shared reference to your model, precision, workload, and adaptation goal.