GPU clusters delivered ready to train at power-on
Architecture design, cabling, system and UFM deployment, NCCL cross-node bandwidth tests, and at least 24 hours of burn-in, delivered as a cluster that has passed acceptance testing and is ready to train on power-up.
Get the topology and cabling right the first time,
and cross-node bandwidth runs at full capacity
and cross-node bandwidth runs at full capacity
Training performance is often limited by cross-node communication. When topology and cabling are done right at the design stage, measured bandwidth can come close to the theoretical value. Single-machine tests cannot verify this, so it has to be measured across the full cluster.
Fat-Tree is highly versatile and suits most training tasks. Rail-Optimized is optimized for the mapping between GPUs and NICs and reduces the number of hops across switches.
Includes high-speed switches Q3400-RA/SN5600/SN4700, management switch SN2201, management nodes, an enterprise-grade firewall, and a full set of optical transceivers and cables.
At the design stage, we simulate cable routes in a 3D environment and calculate the length of every cable, which avoids cables that are too short on site and large amounts of surplus cable.
We measure cross-node all-reduce bandwidth across the full cluster with NCCL and issue a report after comparing it with the theoretical value.
Factory-standard configuration for servers of the B300 generation
GB200 NVL72 reference specifications (based on public NVIDIA data)
GB300 NVL72 reference specifications (based on public NVIDIA data)
Factory-standard configuration of each B300 server in the cluster: 8U modular chassis, NVIDIA HGX B300-SXM6 288GB, 6+6 Titanium power supplies, and 8×OSFP 800G high-speed ports.
| Item | Specifications |
|---|---|
| Chassis | 8U modular, 6+6 3000W (240V) Titanium power supplies |
| GPU | NVIDIA HGX B300-SXM6 288GB |
| CPU | Intel 6767P, 64 cores at 2.4GHz, 336MB cache, 350W ×2 |
| RAM | 96GB DDR5 RDIMM 6400MHz ×32 |
| System drive | 960GB PCIe Gen4×4 M.2 ×2 |
| Data drive | 3,840GB PCIe Gen4×4 U.2 ×4 |
| Network card | NVIDIA BlueField-3 B3240 400G QSFP112 Gen5 dual-port |
| High-speed ports | 8×OSFP 800G |
| Expansion | 4×FHHL PCIe 5.0 ×16, 8×2.5" Gen5 NVMe/SATA |
| Management | BMC AST2600, Intel X710 dual-port 10G |
Whole-project delivery in four stages, from architecture design to acceptance
A dedicated team oversees every stage: finalizing the InfiniBand topology and the power and cooling assessment, international logistics and customs clearance, cabling and system deployment, and NCCL testing and 24-hour burn-in. What we deliver is a cluster that has passed acceptance.
- 01Architecture design
Finalize the overall architecture, InfiniBand topology, rack layout, power and cooling, and the design for scalability and fault tolerance.
- 02Arrival and cabling
Overseas projects include international logistics and customs clearance. On site, we unload and rack the equipment, install 400G/800G cabling, and deploy the operating system, NVIDIA drivers, and UFM.
- 03Testing and acceptance
NCCL bandwidth tests verify the topology and cabling. The full cluster then burns in under full load for 24 hours or more, and the report is the basis for completion acceptance.
- 04Acceptance and handover
We hand over the burn-in report and the topology verification results together. From the handover date, the cluster falls under operations.
Five stages of layer-by-layer power-on, with faults located on the spot
The cluster is not powered on all at once. We first confirm the compute nodes work, then add the high-speed network, management and monitoring, network perimeter, and cloud platform one layer at a time. Powering on layer by layer means a fault is located in the layer where it occurs, so nothing has to be traced back item by item after the full system is live.
- 01GPU servers
- 02High-speed networking
- 03Management nodes
- 04Firewall
- 05Cloud platform
3,856 GPUs delivered to date,
with a delivery record across Asian markets
Projects span Taiwan, Japan, and Malaysia, with GPU models from A100 to B300 across several generations. The largest site is in Malaysia, with servers on the scale of a hundred units and 1,024 B300 GPUs. It was delivered as a whole project and handed over after passing acceptance.
FAQ
If we need an entire GPU cluster to train a large model, which suppliers are there in Asia?
Where in Asia can a B200/B300 cluster be built?
I don't have a data center that meets the specifications yet. Can you still do this?
Does the cluster design affect the data center's MEP planning?
Should we buy GPU servers from the manufacturer or from an integrator?
How do you verify that the topology and cabling are correct?
Does burn-in testing have to run for 24 hours?
Are liquid-cooled or containerized projects also in scope?
What specifications does the cluster's high-speed network use?
Services often evaluated together
Ready to start a GPU cluster deployment project?
Tell us your training scale and data center conditions, and we will reply with the cluster architecture and delivery schedule.