2× RTX 6000 Ada Rackmount RAG Server — 96GB Aggregate VRAM

This planning baseline targets an on-prem RAG node that can separate ingestion, embedding, reranking, and generation work across two professional GPUs. Each RTX 6000 Ada supplies 48GB of ECC graphics memory; the two boards provide 96GB aggregate capacity, not one transparent 96GB memory pool.

Configuration at a glance

GPU layout
2 × RTX 6000 Ada
VRAM per GPU
48GB ECC GDDR6
Aggregate VRAM
96GB total
GPU board power
600W combined
Chassis plan
4U rackmount
Platform
Xeon W790 + ECC RDIMM

Memory boundary: 96GB is aggregate board memory. A normal process still sees 48GB on one GPU; spanning both cards requires explicit framework and model-parallel support.

Where this dual-GPU RAG plan fits

  • Assign embedding or reranking services to one GPU and generation services to the other.
  • Run multiple model workers or evaluation jobs with explicit per-GPU process placement.
  • Use a model-parallel runtime when one workload must span both cards; support and scaling efficiency depend on the software stack.

Platform, storage, power, and cooling checks

  • PCIe: ASUS documents four PCIe 5.0 x16 slots with a Xeon W-2400 processor on this W790 board. Verify the final two-card slot placement, riser topology, and chassis clearance against the board and chassis manuals.
  • Storage: ASUS documents M.2_1 and M.2_2 as disabled with Xeon W-2400. The modeled second NVMe therefore needs a documented alternate path, such as a supported SlimSAS or add-in adapter, before procurement.
  • Power: The two GPUs account for 600W of maximum board power before CPU, memory, drives, fans, and conversion losses. ComputeAtlas models a 1350W target and selects EVGA SuperNOVA 1600 P+ (1600W Platinum); validate connectors, rail distribution, and transient behavior with the final PSU vendor.
  • Cooling: Each RTX 6000 Ada is an active dual-slot card, but a dense 4U deployment still requires an unobstructed front-to-back airflow path and a chassis-specific thermal review.

Related planning tools

RAG and vector-search workstation guidanceCompare the broader CPU, memory, storage, and GPU tradeoffs for RAG deployments.RTX 6000 Ada 48GB local-LLM model fitSee which governed model and precision planning targets fit inside one 48GB RTX 6000 Ada before deciding whether two GPUs are justified.Compare GPU specificationsCompare VRAM, power, cooling, and deployment class across the current GPU catalog.24GB vs 48GB vs 96GB VRAMUnderstand per-GPU memory limits, aggregate capacity, and when software must span devices.PCIe lanes and slot spacingValidate electrical width, physical clearance, and CPU-dependent platform behavior.Multi-GPU airflow and coolingPlan intake, exhaust, spacing, and thermal behavior before selecting a chassis.Estimate your AI hardwareStart with workload requirements, then open an eligible configuration in Builder.

Specification sources and evidence boundary

Manufacturer sources support the component specifications below. The complete build fit, PSU target, and workload guidance remain ComputeAtlas planning models—not benchmark evidence, live pricing, or procurement approval.

Use case

RAG server

Example system budget

$17,500 modeled planning estimate.

Source: ComputeAtlas generated-build budget model (generated-build-budget-v1) • freshness: model-derived • uncertainty: high • not live. Heuristic planning budget derived from ComputeAtlas build-tier formulas. It is not a summed retailer cart, seller quote, or live market observation.

Hardware breakdown

  • GPU: 2x RTX 6000 Ada 48GB
  • CPU: Xeon w7-2495X
  • RAM: Kingston Server Premier DDR5-5600 ECC RDIMM 128GB (4x32GB KSM56R46BD8-32HA)
  • Motherboard: ASUS Pro WS W790E-SAGE SE
  • Storage: 2x WD Black SN850X 2TB NVMe + 8TB enterprise SSD tier
  • PSU: EVGA SuperNOVA 1600 P+ (1600W Platinum)

What This Build Includes

Includes:

  • GPU(s)
  • CPU
  • RAM
  • Storage
  • Motherboard

Not Included:

  • Case / chassis
  • Cooling system
  • Power cables / adapters
  • Peripherals

Deployment Notes

  • High-power multi-GPU systems require proper airflow
  • Ensure PSU headroom for GPU transient spikes
  • Verify motherboard PCIe lane and spacing compatibility
  • Suitable for workstation or rack environments
Build this system

Builder component amounts are legacy internal planning estimates with unknown freshness, not live market quotes. Validate current market pricing, seller terms, taxes/shipping, and availability before purchasing hardware.