KONST
AI Training CalculatorContact us
EN
SERVICE | GPU cluster build

GPU clusters delivered ready to train at power-on

Architecture design, cabling, system and UFM deployment, NCCL cross-node bandwidth tests, and at least 24 hours of burn-in, delivered as a cluster that has passed acceptance testing and is ready to train on power-up.

GPUs delivered to date3,856GPUsBurn-in of at least 24 hours
SOLUTION | KONST's approach

Get the topology and cabling right the first time,
and cross-node bandwidth runs at full capacity

Training performance is often limited by cross-node communication. When topology and cabling are done right at the design stage, measured bandwidth can come close to the theoretical value. Single-machine tests cannot verify this, so it has to be measured across the full cluster.

1234
  • Fat-Tree is highly versatile and suits most training tasks. Rail-Optimized is optimized for the mapping between GPUs and NICs and reduces the number of hops across switches.

  • Includes high-speed switches Q3400-RA/SN5600/SN4700, management switch SN2201, management nodes, an enterprise-grade firewall, and a full set of optical transceivers and cables.

  • At the design stage, we simulate cable routes in a 3D environment and calculate the length of every cable, which avoids cables that are too short on site and large amounts of surplus cable.

  • We measure cross-node all-reduce bandwidth across the full cluster with NCCL and issue a report after comparing it with the theoretical value.

FEATURE | Service scope and specifications

Factory-standard configuration for servers of the B300 generation

Factory-standard configuration of each B300 server in the cluster: 8U modular chassis, NVIDIA HGX B300-SXM6 288GB, 6+6 Titanium power supplies, and 8×OSFP 800G high-speed ports.

ItemSpecifications
Chassis8U modular, 6+6 3000W (240V) Titanium power supplies
GPUNVIDIA HGX B300-SXM6 288GB
CPUIntel 6767P, 64 cores at 2.4GHz, 336MB cache, 350W ×2
RAM96GB DDR5 RDIMM 6400MHz ×32
System drive960GB PCIe Gen4×4 M.2 ×2
Data drive3,840GB PCIe Gen4×4 U.2 ×4
Network cardNVIDIA BlueField-3 B3240 400G QSFP112 Gen5 dual-port
High-speed ports8×OSFP 800G
Expansion4×FHHL PCIe 5.0 ×16, 8×2.5" Gen5 NVMe/SATA
ManagementBMC AST2600, Intel X710 dual-port 10G
PROCESS | Workflow

Whole-project delivery in four stages, from architecture design to acceptance

A dedicated team oversees every stage: finalizing the InfiniBand topology and the power and cooling assessment, international logistics and customs clearance, cabling and system deployment, and NCCL testing and 24-hour burn-in. What we deliver is a cluster that has passed acceptance.

  1. 01Architecture design

    Finalize the overall architecture, InfiniBand topology, rack layout, power and cooling, and the design for scalability and fault tolerance.

  2. 02Arrival and cabling

    Overseas projects include international logistics and customs clearance. On site, we unload and rack the equipment, install 400G/800G cabling, and deploy the operating system, NVIDIA drivers, and UFM.

  3. 03Testing and acceptance

    NCCL bandwidth tests verify the topology and cabling. The full cluster then burns in under full load for 24 hours or more, and the report is the basis for completion acceptance.

  4. 04Acceptance and handover

    We hand over the burn-in report and the topology verification results together. From the handover date, the cluster falls under operations.

PROCESS | Workflow

Five stages of layer-by-layer power-on, with faults located on the spot

The cluster is not powered on all at once. We first confirm the compute nodes work, then add the high-speed network, management and monitoring, network perimeter, and cloud platform one layer at a time. Powering on layer by layer means a fault is located in the layer where it occurs, so nothing has to be traced back item by item after the full system is live.

  1. 01GPU servers
  2. 02High-speed networking
  3. 03Management nodes
  4. 04Firewall
  5. 05Cloud platform
EVIDENCE | Track record and data

3,856 GPUs delivered to date,
with a delivery record across Asian markets

Projects span Taiwan, Japan, and Malaysia, with GPU models from A100 to B300 across several generations. The largest site is in Malaysia, with servers on the scale of a hundred units and 1,024 B300 GPUs. It was delivered as a whole project and handed over after passing acceptance.

3,856GPUsGPUs delivered to date
1,024GPUsLargest single project (Malaysia, B300)
24hoursBurn-in test minimum
FAQ | Common questions

FAQ

If we need an entire GPU cluster to train a large model, which suppliers are there in Asia?
Yes. KONST has delivered cluster projects in Taiwan, Japan, and Malaysia, the largest being 1,024 B300 GPUs. Projects take one of two forms: the owner already has a data center and commissions only the cluster deployment, or the owner commissions the data center and the cluster together and the whole compute build service line handles it.
Where in Asia can a B200/B300 cluster be built?
Overseas, the site is selected for each project. Taiwan, Japan, and Malaysia all have delivered sites, and the largest single project is 1,024 B300 GPUs.
I don't have a data center that meets the specifications yet. Can you still do this?
Yes, KONST can take on the data center as well. Cluster build assumes the owner already has a data center environment that meets the specifications. If not, infrastructure build, data center systems build, and environmental monitoring and control build first prepare the site, and cluster deployment follows.
Does the cluster design affect the data center's MEP planning?
Yes. The first output of the architecture design is a specification: power per rack, required cooling capacity, and rack layout are all set by the cluster design, and the MEP systems are worked back from that specification.
Should we buy GPU servers from the manufacturer or from an integrator?
If you are only buying equipment, either works. The difference comes after the equipment arrives: manufacturers usually do not perform topology design, cabling, drivers and UFM, bandwidth testing, or burn-in testing. An integrator's responsibility ends when the cluster passes acceptance, not when the equipment is delivered.
How do you verify that the topology and cabling are correct?
Run an NCCL bandwidth test to measure cross-node collective communication bandwidth, then compare it with the theoretical value. This is the main way to verify that the topology and cabling are correct.
Does burn-in testing have to run for 24 hours?
Burn-in runs for at least 24 hours. The actual duration is set in the contract. High-density, liquid-cooled, or large-scale projects run longer, because heat buildup and power supply limits take more time to show up.
Are liquid-cooled or containerized projects also in scope?
Yes, they are in scope. The Fukushima project in Japan is a liquid-cooled container deployment of 32 B200 servers, and KONST was responsible for architecture planning, procurement coordination, and acceptance management.
What specifications does the cluster's high-speed network use?
We plan an end-to-end InfiniBand deployment based on LLM training requirements. The switches are Q3400-RA, SN5600, and SN4700, the node NICs are NVIDIA BlueField-3, and cabling is 400G/800G.

Ready to start a GPU cluster deployment project?

Tell us your training scale and data center conditions, and we will reply with the cluster architecture and delivery schedule.