GPU memory
48GB GDDR6 ECC
48GB local-LLM model-fit guide
The RTX 6000 Ada provides 48GB of ECC GDDR6 on one professional workstation GPU. That larger single-card memory envelope can keep some quantized 70B-class and sparse-MoE planning targets on one board, but capacity alone does not guarantee a specific runtime, context length, or tokens-per-second result.
GPU memory
48GB GDDR6 ECC
Maximum power
300W
PCIe interface
PCI Express Gen 4 x16
Form factor
Dual-slot
Thermal
Active
Manufacturer specification verified 2026-10-03. The ComputeAtlas catalog is build-gated against the 48GB / 300W professional-workstation baseline.
This table is generated from the same MODEL_PROFILES registry used by the ComputeAtlas AI Hardware Estimator. A green result means the standard inference planning target is at or below 48GB; it is not a runtime guarantee, benchmark claim, or promise that every context length and software stack will fit.
| Model | Precision | ComputeAtlas planning target | 48GB result | Capacity margin |
|---|---|---|---|---|
| Llama 3 8B | FP16 | 18GB | Within 48GB planning reference | 30GB remaining versus the planning target |
| Llama 3 8B | 8-bit | 10GB | Within 48GB planning reference | 38GB remaining versus the planning target |
| Llama 3 8B | 4-bit | 5GB | Within 48GB planning reference | 43GB remaining versus the planning target |
| Llama 3 70B | FP16 | 154GB | Exceeds 48GB single-GPU reference | 106GB above one-card capacity |
| Llama 3 70B | 8-bit | 81GB | Exceeds 48GB single-GPU reference | 33GB above one-card capacity |
| Llama 3 70B | 4-bit | 42GB | Within 48GB planning reference | 6GB remaining versus the planning target |
| Mixtral 8x7B | FP16 | 104GB | Exceeds 48GB single-GPU reference | 56GB above one-card capacity |
| Mixtral 8x7B | 8-bit | 55GB | Exceeds 48GB single-GPU reference | 7GB above one-card capacity |
| Mixtral 8x7B | 4-bit | 29GB | Within 48GB planning reference | 19GB remaining versus the planning target |
| DeepSeek LLM 67B | FP16 | 148GB | Exceeds 48GB single-GPU reference | 100GB above one-card capacity |
| DeepSeek LLM 67B | 8-bit | 78GB | Exceeds 48GB single-GPU reference | 30GB above one-card capacity |
| DeepSeek LLM 67B | 4-bit | 41GB | Within 48GB planning reference | 7GB remaining versus the planning target |
| Qwen3 8B | FP16 | 18GB | Within 48GB planning reference | 30GB remaining versus the planning target |
| Qwen3 8B | 8-bit | 10GB | Within 48GB planning reference | 38GB remaining versus the planning target |
| Qwen3 8B | 4-bit | 5GB | Within 48GB planning reference | 43GB remaining versus the planning target |
| Qwen3 32B | FP16 | 73GB | Exceeds 48GB single-GPU reference | 25GB above one-card capacity |
| Qwen3 32B | 8-bit | 38GB | Within 48GB planning reference | 10GB remaining versus the planning target |
| Qwen3 32B | 4-bit | 20GB | Within 48GB planning reference | 28GB remaining versus the planning target |
| Mistral Small 3.1 24B | FP16 | 55GB | Exceeds 48GB single-GPU reference | 7GB above one-card capacity |
| Mistral Small 3.1 24B | 8-bit | 28GB | Within 48GB planning reference | 20GB remaining versus the planning target |
| Mistral Small 3.1 24B | 4-bit | 15GB | Within 48GB planning reference | 33GB remaining versus the planning target |
| Gemma 3 27B | FP16 | 60GB | Exceeds 48GB single-GPU reference | 12GB above one-card capacity |
| Gemma 3 27B | 8-bit | 32GB | Within 48GB planning reference | 16GB remaining versus the planning target |
| Gemma 3 27B | 4-bit | 17GB | Within 48GB planning reference | 31GB remaining versus the planning target |
| DeepSeek-R1-Distill-Qwen-32B | FP16 | 71GB | Exceeds 48GB single-GPU reference | 23GB above one-card capacity |
| DeepSeek-R1-Distill-Qwen-32B | 8-bit | 37GB | Within 48GB planning reference | 11GB remaining versus the planning target |
| DeepSeek-R1-Distill-Qwen-32B | 4-bit | 20GB | Within 48GB planning reference | 28GB remaining versus the planning target |
| DeepSeek-R1 671B | FP16 | 1477GB | Exceeds 48GB single-GPU reference | 1429GB above one-card capacity |
| DeepSeek-R1 671B | 8-bit | 772GB | Exceeds 48GB single-GPU reference | 724GB above one-card capacity |
| DeepSeek-R1 671B | 4-bit | 403GB | Exceeds 48GB single-GPU reference | 355GB above one-card capacity |
| gpt-oss-20b | FP16 | 47GB | Within 48GB planning reference | 1GB remaining versus the planning target |
| gpt-oss-20b | 8-bit | 24GB | Within 48GB planning reference | 24GB remaining versus the planning target |
| gpt-oss-20b | 4-bit | 16GB | Within 48GB planning reference | 32GB remaining versus the planning target |
| gpt-oss-120b | FP16 | 258GB | Exceeds 48GB single-GPU reference | 210GB above one-card capacity |
| gpt-oss-120b | 8-bit | 135GB | Exceeds 48GB single-GPU reference | 87GB above one-card capacity |
| gpt-oss-120b | 4-bit | 80GB | Exceeds 48GB single-GPU reference | 32GB above one-card capacity |
If the complete standard inference planning target is below 48GB, the RTX 6000 Ada has a credible single-card capacity case. Quantization can materially change that result, especially for 67B–70B-class models.
A planning target in the low-40GB range leaves limited margin for larger contexts, alternate runtimes, adapters, or concurrency. Validate the exact software stack before procurement.
If the governed target exceeds 48GB, move to offload, a software-aware multi-GPU topology, or a higher-memory single GPU instead of pretending the workload cleanly fits.
The RTX 6000 Ada pairs 48GB ECC memory with a 300W dual-slot professional form factor, which changes workstation density and power planning compared with large open-air consumer cards.
Two boards provide 96GB aggregate VRAM, but that is not one transparent 96GB memory pool. Separate workers can use one GPU each; a single model spanning both cards needs software-aware model parallelism or sharding that your chosen runtime actually supports.
Primary source for 48GB ECC GDDR6, 300W maximum power, PCIe Gen 4 x16, dual-slot form factor, and active thermal design.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-09-09.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-09-09.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-09-09.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-10-03.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-10-03.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-10-03.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-10-03.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-10-03.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-10-03.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-10-03.
Model evidence used by the shared ComputeAtlas workload registry. Verified 2026-10-03.