KONST
AI Training CalculatorContact us
EN
Diagram of a containerized modular data centerDiagram of a containerized modular data center

Flagship project · Cyberjaya, Malaysia

Cyberjaya, Malaysia1,024 B300 GPU cluster compute service

KONST Group, in partnership with ADATA, operates its own cluster of 1,024 B300 GPUs in Cyberjaya, Malaysia, and offers compute services to external customers as bare metal.

Location
Cyberjaya, Selangor, Malaysia
Cluster size
1,024 NVIDIA B300 GPUs
Go-live
Expected March 2027
PartnersADATA Technology
1,024GPUsNVIDIA B300, 128 production nodes × 8 GPUs
Contracted capacity2,061kW2,016 kW billing basis and 45 kW hot spare, listed separately
25racks17 GPU racks and 8 network and management racks
March2027POC and acceptance go-live complete
Highlights

131 B300 liquid-cooled nodes with 1:1 non-blocking interconnect

In this project KONST holds both the site capacity and the GPU cluster. One team is responsible end to end, covering AIDC design, power and capacity planning, GPU procurement, rack deployment, performance validation, and cluster operations.

The cluster is delivered as bare metal. The customer gets root access to the operating system and out-of-band management access, and decides the scheduler layer and software stack. The cluster is physically dedicated to the customer and shares no racks or compute resources with other tenants.

Best forCompanies that need large-scale GPU compute but do not want to buy hardware, lease a data center, or take on build risk themselves
Project
Cyberjaya GPU cluster, Malaysia
Location
Cyberjaya, Selangor, Malaysia
Partners
ADATA Technology
Service type
Bare-metal GPU cluster compute service
Cluster size
1,024 NVIDIA B300 GPUs | 131 liquid-cooled nodes
Compute network
InfiniBand 1:1 non-blocking
Go-live
Expected March 2027
131nodesB300 eight-GPU liquid-cooled nodes

Site and compute, covered at once

KONST holds both the site capacity and the GPU cluster, so customers do not need to buy hardware, lease a data center, or take on build risk. Compute, network, and operations are covered in a single contract.

8×800Gb/sCompute network bandwidth per node

Non-blocking InfiniBand

Each node has 8 OSFP XDR links at 800 Gb/s in a 1:1 non-blocking Spine-Leaf topology, suited to the east-west bandwidth needs of large-scale distributed training.

1.8TB/sNVLink interconnect bandwidth

High-speed interconnect across eight GPUs in one domain

Inside a node, NVLink and NVLink Switch provide 1.8 TB/s of GPU-to-GPU bandwidth across the 8 GPUs of one domain.

3nodesHot spare nodes

Hot spares take over immediately

Hot spare nodes are provided in addition to the production cluster. They take over when a node fails, failures are handled under the manufacturer's RMA procedure, and spare parts are stored on site.

One team is responsible from site design to cluster operations

Most compute projects can only get involved in one stage: they stop after the AIDC is built, or they take over a ready-made cluster from someone else and resell it. In this project KONST is responsible for all six stages.

01

AIDC design

Site, power, and cooling architecture

02

Capacity and power planning

2,061 kW contracted capacity allocation

03

GPU procurement

131 B300 liquid-cooled nodes

04

Rack deployment

Rack cabling and liquid cooling pipe connections

05

Testing and validation

CFD thermal simulation and NCCL bandwidth validation

06

Cluster operations

24×7 monitoring and on-site staff

* Production GPUs are 128 nodes × 8 GPUs, for a total of 1,024. The 3 hot spare nodes (24 GPUs) do not count toward production capacity. The actual configuration follows the order signed by both parties.

Cluster specifications

Full specifications of the Cyberjaya cluster

Each node is an NVIDIA HGX B300 eight-GPU liquid-cooled platform. The production cluster has 128 nodes, plus 3 hot spares.

Cross-section diagram of the container module, showing the layout of racks, aisles, cable trays, and cooling equipment
Container module cross-section diagram

Container module dedicated to a single tenant

The Cyberjaya site is a prefabricated containerized modular AI data center. The modules are integrated at the factory, pass factory acceptance, and are then delivered to the site and lifted into place. Site infrastructure is already in place beforehand, so the cluster can be delivered in a powered-on state. The IT pod is dedicated to a single tenant and is not shared with other tenants.

Power
11 kV dual-feeder utility supply, main switching station and transformers, UPS and generator backup
Cooling
Direct-to-chip liquid cooling for CPUs and GPUs with leak detection; residual rack-level heat is handled by an in-row airflow system
Thermal validation
CFD thermal simulation is complete, and all racks meet the inlet air temperature limit of 27 °C or below
Physical security
Two layers of fencing around the campus and data center, a full set of CCTV and sensors, access control, and 24×7 monitoring
GPU nodesNVIDIA HGX B300 eight-GPU liquid-cooled platform (NVL8 form), 131 nodes in total: 128 in production and 3 hot spares
Total GPUs1,024 (128 nodes × 8 GPUs) *
Node processors2 × Intel Xeon 6776P, 64 cores, 2.3 GHz, 350 W
Node memory3 TB DDR5 RDIMM 6400 MHz
Node local storage8 × 3.84 TB Gen5 NVMe (2.5-inch), plus 2 × 960 GB M.2 for the operating system
Node NICNVIDIA ConnectX-8 SuperNIC, compatible with BlueField-3 DPU
Management nodes12 nodes, AMD EPYC 9334, 32 cores, 2.7 GHz; ConnectX-7 200 G × 2
Rack layout25 racks: 17 GPU racks (16 fully loaded with 8 nodes each, 1 with 3 hot spare nodes) and 8 network and management racks
Shared storageCapacity and aggregate bandwidth are configured to customer requirements and planned separately

* Production GPUs are 128 nodes × 8 GPUs, for a total of 1,024. The 3 hot spare nodes do not count toward production capacity. Site design parameters follow the current design version and will be updated separately if adjusted after the design freeze.

Network fabric

Three separate networks for compute, storage, and management

The compute network is InfiniBand in a 1:1 non-blocking topology. The storage network is a separate fabric and is not shared with the compute network.

Three-network separation architecture

Diagram
InfiniBand compute network8 × 800 Gb/s, 1:1 non-blocking
Storage network200 GbE
Management networkOut-of-band management, BMC/IPMI
128 production nodes · 3 hot spares · 12 management nodesNVIDIA HGX B300 eight-GPU liquid-cooled platform (NVL8)
NVLink and NVLink Switch inside each node8 GPUs per domain, 1.8 TB/s GPU to GPU

The three networks are separate from each other. The cluster is dedicated to a single tenant and not shared with other tenants

1:1 non-blocking compute network

8 × 800 Gb/s OSFP XDR per node in a Spine-Leaf topology. An NCCL all-reduce bandwidth baseline can be included as an acceptance criterion at the customer's request.

Separate storage network

200 GbE connected through the backbone and not shared with the compute network, so storage traffic does not use up the east-west bandwidth that training needs.

Separate management network segment

Provides out-of-band management access through BMC and IPMI. The scope of access is stated in the order.

Why 1:1 non-blocking

In large-scale distributed training, the bottleneck is usually east-west bandwidth during the all-reduce phase, not single-GPU compute. Once the oversubscription ratio goes above 1:1, gradient synchronization gets worse as the node count grows. This cluster adopted 1:1 non-blocking at the design stage.

Partnership

KONST operates the cluster, ADATA provides the site

ADATA Technology is responsible for site and facility operations. KONST operates the GPU cluster itself and stations its own operations staff on site.

Partnership structure

Diagram
Client
The software stack above the operating system and the customer's own workloads
KONSTSelf-operated 1,024 B300 GPU cluster · Single contracting contact
Delivery scope handled by KONST
Cluster hardwareGPU and management nodesRacks and in-rack cabling
Network buildInfiniBand compute networkStorage and management networks
Hardware maintenanceReplacement, spare parts, and RMA3 hot spares take over
Cluster operations24×7 monitoring and incident responseOn-site staff
Site partnerADATA Technology
Cyberjaya AIDC sitePower & CoolingFacility operationsStaff on site 24 hours a day

ADATA Technology is responsible for site and facility operations, with staff on site 24 hours a day. KONST operates the GPU cluster itself and sends its own operations staff to be stationed on site. Operations are split into two parts, and the division of responsibilities is stated in the compute service master agreement.

Cluster Operations

Cluster operations

Handled by KONST's own staff, covering the GPU nodes, network, and hardware layer. Only KONST operations staff have access rights to the white space.

  • Monitoring of GPU nodes and management nodes
  • Hardware replacement, spare parts, and RMA handling
  • 3 hot spare nodes replace failed nodes
  • Monthly reports at the node, storage, and GPU level
Facility Operations

Facility operations

Handled by the data center operator, covering the site, power, cooling, and physical security, with staff on site 24 hours a day to handle facility-level events.

  • 11 kV dual-feeder utility supply with UPS and generator backup
  • Direct-to-chip liquid cooling and leak detection
  • Two layers of fencing, CCTV, and access control
  • Resident engineers on site 24 hours a day
Speed to market | Build schedule

The build has started and is not affected by when the contract is signed

The container modules are in factory integration. The cluster is progressing on the current build schedule, in parallel with commercial discussions.

01

Factory integration and factory acceptanceFAT

November 2026

02

First modules arrive on site

December 2026

03

Module lifting and MEP connection

January 2027

04

Rack installation and cluster commissioning

January to February 2027

05

POC and go-live

March 2027

* The schedule above is a working estimate of the current build plan, provided for evaluation purposes. The contractually binding delivery date is the delivery date agreed in the contract signed by both parties.

Service model | Bare-metal delivery and responsibilities

KONST is responsible for the hardware layer; the customer has full control above the operating system

The customer gets root access to the operating system. The cluster is physically dedicated to the customer and shares no racks or compute resources with other tenants.

Infrastructure

KONST responsibility

  • Data center space, power, and cooling, including direct-to-chip liquid cooling water loops
  • GPU nodes, management nodes, racks, and in-rack cabling
  • Build of the InfiniBand compute network, storage network, and management network
  • External network access and handoff point
  • Hardware maintenance and replacement, spare parts, and RMA handling
  • 3 hot spare nodes that take over on failure
  • 24×7 monitoring and tiered incident response
Software Stack

Customer responsibility

  • Operating system, firmware versions, and driver installation
  • Scheduler layer (Slurm or Kubernetes) deployment and maintenance
  • CUDA version and AI framework software stack
  • Own workloads and data
  • Account, key, and access credential management
  • Workload-level performance tuning
Optional

Can be delegated to KONST

  • Managed orchestration: Slurm or K8s control plane deployment and lifecycle management
  • Cluster performance validation (NCCL/HPL)
  • Custom software environment build

* The final scope of responsibilities follows the compute service master agreement (MSA) and the order form signed by both parties.

FAQ | Common questions

The seven questions asked most often when evaluating the Cyberjaya cluster

Answers reflect project terms as of September 2026.

How much compute can this cluster provide?

The production cluster has 128 NVIDIA HGX B300 liquid-cooled nodes with eight GPUs each, for a total of 1,024 GPUs. It also has 3 hot spare nodes and 12 management nodes, spread across 25 racks. Contracted capacity is 2,061 kW. Each node has 2 Intel Xeon 6776P CPUs, 3 TB of DDR5 memory, and 8 Gen5 NVMe drives of 3.84 TB each.

What is the cluster's network architecture?

The compute network uses InfiniBand. Each node has 8 OSFP XDR links at 800 Gb/s, provided by the NVIDIA ConnectX-8 SuperNIC, in a 1:1 non-blocking Spine-Leaf topology. Inside a node, GPUs connect over NVLink and NVLink Switch, at 1.8 TB/s across the 8 GPUs of one domain. The storage network is a separate 200 GbE network, and the management network is on its own segment. The three networks are kept apart and do not share links.

Why use a 1:1 non-blocking network instead of standard Ethernet?

In large-scale distributed training, the bottleneck is usually east-west bandwidth during the all-reduce phase, not single-GPU compute. Once the oversubscription ratio goes above 1:1, gradient synchronization gets worse as the node count grows. This cluster adopted 1:1 non-blocking at the design stage, and an NCCL all-reduce bandwidth baseline can be included as an acceptance criterion at the customer's request.

Is delivery bare metal or virtualized?

Bare metal delivery. The customer gets root access to the operating system. The scope of out-of-band management access through BMC and IPMI is stated in the order. The customer chooses the scheduler layer and software stack. The cluster is physically dedicated to the customer and shares no racks or compute resources with other tenants.

What happens if a node fails or the facility has an outage?

The cluster has 3 hot spare nodes that take over when a node fails. Failures are handled under the manufacturer's RMA procedure, and spare parts are stored on site in Malaysia. Operations are split into two parts. The data center operator is responsible for the site and facilities, with staff on site 24 hours a day. KONST sends its own operations staff to be stationed on site for the GPU nodes, network, and hardware layer. Cluster availability and tiered incident response terms follow the compute service master agreement and the order.

What is the partnership model and how is it billed?

Pricing is based on GPU usage hours, and the compute service fee is charged monthly. The standard contract term is 5 years, and terms shorter than 5 years are negotiated based on scale and conditions. A performance deposit is paid after signing and is offset during the service period as agreed. Billing starts on the service start date, when the customer confirms in writing that acceptance has passed, not on the equipment delivery date. Network connectivity, shared storage, managed orchestration, and performance validation are quoted separately. Contact the sales team for actual rates.

If we start talks now, when can we start using it?

The cluster build has started and runs in parallel with commercial discussions, and the schedule does not change with the signing date. The container modules finish factory integration and factory acceptance in November 2026, the first modules arrive on site in December 2026, module lifting and MEP connection take place in January 2027, rack installation and cluster commissioning run from January to February 2027, and POC and acceptance go-live are expected to be complete in March 2027. Going from discussions to signing takes about one month.

How many GPUs do you need, and when?

Tell us how many GPUs you need, how long you need them, and your storage and network requirements. KONST provides a recommended cluster configuration, a draft of acceptance criteria, and a formal quote.