Every build below is computed from the catalog data and the sizing rules; use them to check arithmetic. Prices in USD as of 2026-09-09; final price confirmed on order. Continental US; Canada quoted case by case. Orders and quotes: book a meeting. Every page: llms.txt (map) · homepage (all tables) · catalog.json (data).
GPUs → nodes: 32 ÷ 8 = 4 nodes. Multi-node training means the GPUs act as one, so HGX class (B300, 288 GB per GPU, 2.3 TB per node); a PCIe server would not serve it. Host memory stays at the listed 2TB (no offload stated).
GPU fabric: 1 port per GPU = 32 × 800G. 32 ≤ 32, so one 32× 800G OSFP switch is non-blocking at 1:1 (single point of failure, accepted; growth past 4 nodes needs the 64-port switch or a fat-tree: at 32 nodes, 8 leaves + 4 spines on 64-port switches). Links: all 32 GPU ports priced optical (2× SR8 + MTP-12 fibre, $1,908 per link) — at 2 nodes per 40 kW rack the cluster spans 2 racks, so most links leave the switch's rack and 800G passive DAC (3 m reach) cannot be assumed; swapping the 16 genuinely in-rack links to DAC is a final-engineering saving, not part of the estimate. No storage on this fabric.
Front-end (converged with storage): each node takes the listed +2× 200G option and connects one port to each switch of an MLAG pair of 24× 200G + 8× 400G switches, 3:1 or better; the storage nodes' 2× 100G ports land on the same pair through 200G→2×100G breakouts (5 nodes × 2 ports = 10 ports, 3 breakout cables per switch); peer link 2× 400G DAC. Uplinks to the buyer's core are not priced. Storage stays converged because 16 GB/s across 4 nodes is far below the 40 GB/s-per-node dedicated-fabric threshold.
Storage supply: 2U all-NVMe storage node, dual socket, 24x 7.68TB NVMe (184TB raw) = 184 TB raw, 20 GB/s planning per node. Replica 3 on N nodes plans to 184 × N × (N−1)/N × 0.333 × 0.85: N = 4 gives 156.2 TB (short), N = 5 gives 208.3 TB ≥ 204.5 TB. Throughput 5 × 20 = 100 GB/s ≥ 16. Erasure coding 4+2 would need 6 hosts and plan to 521.6 TB, 2.6× the need, so replica 3 on 5 nodes is the answer. Software (Ceph) is buyer-supplied and not priced.
Racks at 40 kW: B300 = 8 U and 16.96 kW provisioned, so 2 per rack (power binds). Rack A: 2× B300 + GPU switch + OOB switch = 35.24 kW; rack B: 2× B300 = 33.92 kW; rack C: 5× storage + front-end pair = 9.75 kW, 12 U. Racks, PDUs and cabling labour are quoted separately.
Power and colo: 78.91 kW provisioned in total; 5-year colo $175/kW/month × 78.91 kW = $13,809/month. Management nodes (3× the listed 1U compute node) are recommended at this size and not included here.
Worked design: 2 PB usable archive on 5× HDD storage nodes (RAID 6), 100G front-end pair, 5-year colo
Per node: 2U HDD storage node, single socket, 24x 24TB SAS (576TB raw) = 576 TB raw. RAID 6 as two 12-drive groups (10+2) = 0.833 efficiency; plannable = 576 × 0.833 × 0.85 (fill ceiling) = 407.8 TB per node.
Nodes: ceil(2,000 ÷ 407.8) = 5 nodes = 2,039.0 TB plannable (1.02× the need, inside the 2× overshoot check). Throughput 5 × 5 GB/s ≈ 25 GB/s streaming (estimate); no stated demand, so capacity binds. Each node is one failure domain: a node outage takes its 407.8 TB offline until repaired.
Alternative, distributed across nodes with Ceph erasure coding 4+2 (0.667 efficiency, 6-host floor): plannable = 576 × N × (N−1)/N × 0.667 × 0.85; N = 7 gives 1,959.4 TB (short), N = 8 gives 2,285.9 TB, so 8 nodes ($413,947.60 for the nodes) plus NVMe DB/WAL drives at 1–4% of HDD capacity via the rate card. Survives a whole node.
Network: each node has 2× 100G; one port to each switch of an MLAG pair of 32× 100G QSFP28 switches (10 ports of 32 per switch), peer link 2× 100G DAC, uplinks to the buyer's core not priced; the listed −$874 option swaps to 2× 25G if the core is 25G. OOB: 5 nodes + 2 switches on one 48× 1G RJ45 + 4× 10G switch.
Rack and power: 5 × 2 U + 3 U of switches = 13 U, 4.59 kW provisioned: one rack, one 208V 1-ph 30A circuit pair. 5-year colo $803/month.