Diagram of a containerized modular data centerFlagship project · Cyberjaya, Malaysia
Cyberjaya, Malaysia1,024 B300 GPU cluster compute service
KONST Group, in partnership with ADATA, operates its own cluster of 1,024 B300 GPUs in Cyberjaya, Malaysia, and offers compute services to external customers as bare metal.
- Location
- Cyberjaya, Selangor, Malaysia
- Cluster size
- 1,024 NVIDIA B300 GPUs
- Go-live
- Expected March 2027
131 B300 liquid-cooled nodes with 1:1 non-blocking interconnect
In this project KONST holds both the site capacity and the GPU cluster. One team is responsible end to end, covering AIDC design, power and capacity planning, GPU procurement, rack deployment, performance validation, and cluster operations.
The cluster is delivered as bare metal. The customer gets root access to the operating system and out-of-band management access, and decides the scheduler layer and software stack. The cluster is physically dedicated to the customer and shares no racks or compute resources with other tenants.
- Project
- Cyberjaya GPU cluster, Malaysia
- Location
- Cyberjaya, Selangor, Malaysia
- Partners
- ADATA Technology
- Service type
- Bare-metal GPU cluster compute service
- Cluster size
- 1,024 NVIDIA B300 GPUs | 131 liquid-cooled nodes
- Compute network
- InfiniBand 1:1 non-blocking
- Go-live
- Expected March 2027
Site and compute, covered at once
KONST holds both the site capacity and the GPU cluster, so customers do not need to buy hardware, lease a data center, or take on build risk. Compute, network, and operations are covered in a single contract.
Non-blocking InfiniBand
Each node has 8 OSFP XDR links at 800 Gb/s in a 1:1 non-blocking Spine-Leaf topology, suited to the east-west bandwidth needs of large-scale distributed training.
High-speed interconnect across eight GPUs in one domain
Inside a node, NVLink and NVLink Switch provide 1.8 TB/s of GPU-to-GPU bandwidth across the 8 GPUs of one domain.
Hot spares take over immediately
Hot spare nodes are provided in addition to the production cluster. They take over when a node fails, failures are handled under the manufacturer's RMA procedure, and spare parts are stored on site.
One team is responsible from site design to cluster operations
Most compute projects can only get involved in one stage: they stop after the AIDC is built, or they take over a ready-made cluster from someone else and resell it. In this project KONST is responsible for all six stages.
AIDC design
Site, power, and cooling architecture
Capacity and power planning
2,061 kW contracted capacity allocation
GPU procurement
131 B300 liquid-cooled nodes
Rack deployment
Rack cabling and liquid cooling pipe connections
Testing and validation
CFD thermal simulation and NCCL bandwidth validation
Cluster operations
24×7 monitoring and on-site staff
* Production GPUs are 128 nodes × 8 GPUs, for a total of 1,024. The 3 hot spare nodes (24 GPUs) do not count toward production capacity. The actual configuration follows the order signed by both parties.
Full specifications of the Cyberjaya cluster
Each node is an NVIDIA HGX B300 eight-GPU liquid-cooled platform. The production cluster has 128 nodes, plus 3 hot spares.

Container module dedicated to a single tenant
The Cyberjaya site is a prefabricated containerized modular AI data center. The modules are integrated at the factory, pass factory acceptance, and are then delivered to the site and lifted into place. Site infrastructure is already in place beforehand, so the cluster can be delivered in a powered-on state. The IT pod is dedicated to a single tenant and is not shared with other tenants.
- Power
- 11 kV dual-feeder utility supply, main switching station and transformers, UPS and generator backup
- Cooling
- Direct-to-chip liquid cooling for CPUs and GPUs with leak detection; residual rack-level heat is handled by an in-row airflow system
- Thermal validation
- CFD thermal simulation is complete, and all racks meet the inlet air temperature limit of 27 °C or below
- Physical security
- Two layers of fencing around the campus and data center, a full set of CCTV and sensors, access control, and 24×7 monitoring
| GPU nodes | NVIDIA HGX B300 eight-GPU liquid-cooled platform (NVL8 form), 131 nodes in total: 128 in production and 3 hot spares |
|---|---|
| Total GPUs | 1,024 (128 nodes × 8 GPUs) * |
| Node processors | 2 × Intel Xeon 6776P, 64 cores, 2.3 GHz, 350 W |
| Node memory | 3 TB DDR5 RDIMM 6400 MHz |
| Node local storage | 8 × 3.84 TB Gen5 NVMe (2.5-inch), plus 2 × 960 GB M.2 for the operating system |
| Node NIC | NVIDIA ConnectX-8 SuperNIC, compatible with BlueField-3 DPU |
| Management nodes | 12 nodes, AMD EPYC 9334, 32 cores, 2.7 GHz; ConnectX-7 200 G × 2 |
| Rack layout | 25 racks: 17 GPU racks (16 fully loaded with 8 nodes each, 1 with 3 hot spare nodes) and 8 network and management racks |
| Shared storage | Capacity and aggregate bandwidth are configured to customer requirements and planned separately |
| Compute network | InfiniBand, 8 × 800 Gb/s OSFP XDR per node, provided through the NVIDIA ConnectX-8 SuperNIC |
|---|---|
| Topology | 1:1 non-blocking Spine-Leaf; switch models to be confirmed against the final configuration |
| GPU-to-GPU interconnect | NVLink and NVLink Switch, 8 GPUs per domain, 1.8 TB/s GPU to GPU |
| GPUDirect RDMA | Supported; RDMA between nodes is internal cluster traffic and is not billed separately |
| Storage network | Separate 200 GbE network connected through the backbone, not shared with the compute network |
| Management network | Separate segment providing out-of-band management (BMC/IPMI) access |
| Network isolation | The compute, storage, and management networks are separate. The cluster is dedicated to a single tenant and is not shared with other tenants |
| External connectivity | 2 × 10G dual links, on different carriers and different routes as mutual backup; private line (EPL) and VPN access supported |
| Connectivity handoff | External connectivity is handed off at the patch panel in the MMR, and KONST handles in-rack patching |
| Location | Cyberjaya, Selangor, Malaysia |
|---|---|
| Build type | Prefabricated containerized modular AI data center; IT container pod dedicated to a single tenant |
| Data center tier | Tier III design standard |
| Contracted capacity | 2,061 kW, of which 2,016 kW is the billing basis and 45 kW is hot spare, listed separately * |
| Power architecture | 11 kV dual-feeder utility supply with main switching station and transformers; UPS and generator backup |
| Cooling architecture | Direct-to-chip liquid cooling (DLC) for CPUs and GPUs with leak detection; residual rack-level heat is handled by an in-row airflow system |
| Thermal validation | CFD thermal simulation is complete. The acceptance criterion is a rack inlet air temperature of 27 °C or below, and all simulation results meet it |
| Physical security | Two layers of fencing around the campus and data center, a full set of CCTV and sensors, access control, and 24×7 monitoring |
| Environmental conditions | Temperature and humidity are monitored around the clock; environmental control and fire protection systems follow data center standards |
| Delivery type | Bare metal delivery. The customer gets root access to the operating system; the scope of out-of-band management access through BMC and IPMI is stated in the order |
|---|---|
| Dedicated resources | The cluster is physically dedicated to the customer and does not share racks or compute resources with other tenants |
| Site operations | The data center operator is responsible for facility operations, with staff on site 24 hours a day |
| Cluster operations | KONST sends its own GPU operations staff to be stationed on site, responsible for the GPU nodes, network, and hardware layer |
| Hot spares | 3 hot spare nodes are provided and take over when a node fails; failures are handled under the manufacturer's RMA procedure |
| Spare parts | Spare parts are stored on site in Malaysia |
| Monitoring and reporting | Real-time monitoring metrics at the node, storage, and GPU level, and monthly availability reports |
| Planned maintenance | 7 days' notice; GPU driver updates are scheduled, with the time agreed after confirming customer requirements |
| Acceptance criteria | Can include NCCL all-reduce bandwidth, NVLink bandwidth, GEMM throughput, HBM bandwidth, storage IOPS, and burn-in testing |
| Service level | Cluster availability and tiered incident response terms follow the compute service master agreement (MSA) and the order |
* Production GPUs are 128 nodes × 8 GPUs, for a total of 1,024. The 3 hot spare nodes do not count toward production capacity. Site design parameters follow the current design version and will be updated separately if adjusted after the design freeze.
Three separate networks for compute, storage, and management
The compute network is InfiniBand in a 1:1 non-blocking topology. The storage network is a separate fabric and is not shared with the compute network.
Three-network separation architecture
DiagramThe three networks are separate from each other. The cluster is dedicated to a single tenant and not shared with other tenants
1:1 non-blocking compute network
8 × 800 Gb/s OSFP XDR per node in a Spine-Leaf topology. An NCCL all-reduce bandwidth baseline can be included as an acceptance criterion at the customer's request.
Separate storage network
200 GbE connected through the backbone and not shared with the compute network, so storage traffic does not use up the east-west bandwidth that training needs.
Separate management network segment
Provides out-of-band management access through BMC and IPMI. The scope of access is stated in the order.
In large-scale distributed training, the bottleneck is usually east-west bandwidth during the all-reduce phase, not single-GPU compute. Once the oversubscription ratio goes above 1:1, gradient synchronization gets worse as the node count grows. This cluster adopted 1:1 non-blocking at the design stage.
KONST operates the cluster, ADATA provides the site
ADATA Technology is responsible for site and facility operations. KONST operates the GPU cluster itself and stations its own operations staff on site.
Partnership structure
DiagramADATA Technology is responsible for site and facility operations, with staff on site 24 hours a day. KONST operates the GPU cluster itself and sends its own operations staff to be stationed on site. Operations are split into two parts, and the division of responsibilities is stated in the compute service master agreement.
Cluster operations
Handled by KONST's own staff, covering the GPU nodes, network, and hardware layer. Only KONST operations staff have access rights to the white space.
- Monitoring of GPU nodes and management nodes
- Hardware replacement, spare parts, and RMA handling
- 3 hot spare nodes replace failed nodes
- Monthly reports at the node, storage, and GPU level
Facility operations
Handled by the data center operator, covering the site, power, cooling, and physical security, with staff on site 24 hours a day to handle facility-level events.
- 11 kV dual-feeder utility supply with UPS and generator backup
- Direct-to-chip liquid cooling and leak detection
- Two layers of fencing, CCTV, and access control
- Resident engineers on site 24 hours a day
The build has started and is not affected by when the contract is signed
The container modules are in factory integration. The cluster is progressing on the current build schedule, in parallel with commercial discussions.
Factory integration and factory acceptanceFAT
November 2026
First modules arrive on site
December 2026
Module lifting and MEP connection
January 2027
Rack installation and cluster commissioning
January to February 2027
POC and go-live
March 2027
* The schedule above is a working estimate of the current build plan, provided for evaluation purposes. The contractually binding delivery date is the delivery date agreed in the contract signed by both parties.
KONST is responsible for the hardware layer; the customer has full control above the operating system
The customer gets root access to the operating system. The cluster is physically dedicated to the customer and shares no racks or compute resources with other tenants.
KONST responsibility
- Data center space, power, and cooling, including direct-to-chip liquid cooling water loops
- GPU nodes, management nodes, racks, and in-rack cabling
- Build of the InfiniBand compute network, storage network, and management network
- External network access and handoff point
- Hardware maintenance and replacement, spare parts, and RMA handling
- 3 hot spare nodes that take over on failure
- 24×7 monitoring and tiered incident response
Customer responsibility
- Operating system, firmware versions, and driver installation
- Scheduler layer (Slurm or Kubernetes) deployment and maintenance
- CUDA version and AI framework software stack
- Own workloads and data
- Account, key, and access credential management
- Workload-level performance tuning
Can be delegated to KONST
- Managed orchestration: Slurm or K8s control plane deployment and lifecycle management
- Cluster performance validation (NCCL/HPL)
- Custom software environment build
* The final scope of responsibilities follows the compute service master agreement (MSA) and the order form signed by both parties.
The seven questions asked most often when evaluating the Cyberjaya cluster
Answers reflect project terms as of September 2026.
How much compute can this cluster provide?
The production cluster has 128 NVIDIA HGX B300 liquid-cooled nodes with eight GPUs each, for a total of 1,024 GPUs. It also has 3 hot spare nodes and 12 management nodes, spread across 25 racks. Contracted capacity is 2,061 kW. Each node has 2 Intel Xeon 6776P CPUs, 3 TB of DDR5 memory, and 8 Gen5 NVMe drives of 3.84 TB each.
What is the cluster's network architecture?
The compute network uses InfiniBand. Each node has 8 OSFP XDR links at 800 Gb/s, provided by the NVIDIA ConnectX-8 SuperNIC, in a 1:1 non-blocking Spine-Leaf topology. Inside a node, GPUs connect over NVLink and NVLink Switch, at 1.8 TB/s across the 8 GPUs of one domain. The storage network is a separate 200 GbE network, and the management network is on its own segment. The three networks are kept apart and do not share links.
Why use a 1:1 non-blocking network instead of standard Ethernet?
In large-scale distributed training, the bottleneck is usually east-west bandwidth during the all-reduce phase, not single-GPU compute. Once the oversubscription ratio goes above 1:1, gradient synchronization gets worse as the node count grows. This cluster adopted 1:1 non-blocking at the design stage, and an NCCL all-reduce bandwidth baseline can be included as an acceptance criterion at the customer's request.
Is delivery bare metal or virtualized?
Bare metal delivery. The customer gets root access to the operating system. The scope of out-of-band management access through BMC and IPMI is stated in the order. The customer chooses the scheduler layer and software stack. The cluster is physically dedicated to the customer and shares no racks or compute resources with other tenants.
What happens if a node fails or the facility has an outage?
The cluster has 3 hot spare nodes that take over when a node fails. Failures are handled under the manufacturer's RMA procedure, and spare parts are stored on site in Malaysia. Operations are split into two parts. The data center operator is responsible for the site and facilities, with staff on site 24 hours a day. KONST sends its own operations staff to be stationed on site for the GPU nodes, network, and hardware layer. Cluster availability and tiered incident response terms follow the compute service master agreement and the order.
What is the partnership model and how is it billed?
Pricing is based on GPU usage hours, and the compute service fee is charged monthly. The standard contract term is 5 years, and terms shorter than 5 years are negotiated based on scale and conditions. A performance deposit is paid after signing and is offset during the service period as agreed. Billing starts on the service start date, when the customer confirms in writing that acceptance has passed, not on the equipment delivery date. Network connectivity, shared storage, managed orchestration, and performance validation are quoted separately. Contact the sales team for actual rates.
If we start talks now, when can we start using it?
The cluster build has started and runs in parallel with commercial discussions, and the schedule does not change with the signing date. The container modules finish factory integration and factory acceptance in November 2026, the first modules arrive on site in December 2026, module lifting and MEP connection take place in January 2027, rack installation and cluster commissioning run from January to February 2027, and POC and acceptance go-live are expected to be complete in March 2027. Going from discussions to signing takes about one month.
How many GPUs do you need, and when?
Tell us how many GPUs you need, how long you need them, and your storage and network requirements. KONST provides a recommended cluster configuration, a draft of acceptance criteria, and a formal quote.