An AI data center must be designed around high-density GPU clusters. The design must account for how much power each rack receives, how heat is removed, how much weight the floor supports, and how power and network connections are arranged. It must also identify which systems could be affected by the same failure. Replacing CPU servers in a conventional data center with GPU servers is rarely enough.
Small or lower-density air-cooled GPU deployments may fit through reduced rack density, distributed placement, or targeted upgrades. Multi-node training, racks above 40 kW, and liquid-cooled rack-scale systems above 100 kW require power, cooling, networking, and GPU systems to be planned together as an AIDC.
What is an AI Data Center?
An AI Data Center, or AIDC, is built to keep high-density GPU systems operating reliably while continuously feeding them data and supporting expansion. It integrates power, heat rejection, high-speed networks, storage, and cluster operations. The label AI-ready is not sufficient; the facility must confirm power available at each rack, supported cooling, equipment weight and floor load limits, cable routes, room for maintenance, and what happens during failures.
AI data centers versus traditional data centers
| Dimension | Traditional IDC | AIDC / GPU-ready design | Question to verify |
|---|---|---|---|
| Workload | Web, databases, virtualization, general IT | Training, fine-tuning, batch and online inference | Single-node or synchronized multi-node work? |
| Compute architecture | Relatively independent CPU servers | GPU connections within a system (scale-up) and between systems (scale-out) | How are connections within servers, between servers, and to storage kept separate? |
| Rack power | Lower-density general racks | 20-40 kW racks through 100+ kW rack-scale systems | Agreed power capacity, brief demand spikes, and room to expand? |
| Power path | General dual-feed distribution | Matches equipment voltage, phases, rack power units, and N+1 or 2N backup capacity | Can IT equipment remain operational after any power-path or equipment failure? |
| Cooling | Perimeter air and hot/cold aisles | Air, rear-door heat exchangers, direct-to-chip, or hybrid liquid cooling | Required water temperature, flow, pressure, and redundancy? |
| Rack and space | U height and standard depth | Depth, weight, busway, piping, CDU, cabling, and service area | Can floor loading and access routes support the system? |
| Network and storage | Traffic entering or leaving the data center, plus shared storage | Fast connections between servers, plus storage accessed in parallel | Will communication or data supply stall GPUs? |
| Failure impact | Application or virtualization layer may absorb a server failure | One node, network, or power failure may stop a distributed job | How is job progress saved, how are jobs restarted, and which systems share a failure risk? |
Six questions before building an AIDC
- Workload: training, fine-tuning, batch inference, or online inference; work completed per second, response time, and whether jobs can be interrupted.
- GPU scale: model, quantity, node configuration, and expansion plan.
- Rack power: total requirements for servers, switches, storage, CDU, and redundancy.
- Cooling: air-side and liquid-side heat load, water temperature, flow, pressure, and failure scenarios.
- Cluster connections: which GPUs share an NVLink connection, how servers connect, and the number and length of cables.
- Failure and recovery: which power, cooling, network, server, and storage systems could fail together; how often job progress is saved and how quickly work must resume.
KONST Group can translate these six inputs into an infrastructure plan through Konstra AI, including comparison of IDC upgrades, AIDC zones, colocation, and rented compute.
Six common AIDC planning mistakes
- Buying GPUs before confirming where they can be installed.
- Looking only at the building's total MW instead of delivery through UPS, busway, PDU, and rack PDU.
- Treating liquid cooling as one interchangeable technology.
- Sizing a cluster only by GPU count and ignoring fabric, storage, and software.
- Checking rack units but not depth, weight, piping, cabling, and service access.
- Planning only for normal operation, without mapping how failures in power, cooling, or networking affect the same equipment and its ability to resume saved work.
The first AIDC constraint: power
A facility may have enough total power while lacking a path to deliver it to the target rack. Check every step from the utility supply and transformers to backup power (UPS), power distribution units (PDUs), distribution busways, rack PDUs, and server power supplies. For example, a DGX H100 has a maximum power of 10.2 kW; four systems in one rack create 40.8 kW of server demand. Rack-scale systems such as GB200 and GB300 NVL72 move the design well beyond 100 kW, so power requirements must be checked against the equipment manufacturer's exact configuration.
Power-distribution assessment covers the complete path
- Input voltage, phase, and connector requirements
- Capacity of UPS, PDU, busway, rack PDU, breakers, and conductors
- Safe capacity for sustained loads, usable power relative to supplied power (power factor), balanced loads across phases, and spare capacity
- Whether backup capacity extends to rack and equipment power supplies: one extra capacity unit (N+1) or two complete sets (2N)
- The effect of GPU load variation on facility capacity
- Whether the distributed workload remains operational during generator or UPS transfer
A four-DGX-H100 high-density rack uses a 415 VAC, 32 A, three-phase N+1 design in NVIDIA's reference. Other systems must be engineered from their own input and redundancy requirements. See the NVIDIA DGX SuperPOD electrical design guide.
Cooling options for high-density GPUs
Cooling must reach the rack and remove heat continuously. Airflow, pressure, containment, and heat-exchange paths matter as much as the facility's total cooling figure. Uptime Institute notes that optimized perimeter air cooling may support roughly 20-25 kW per rack, while older systems are often closer to 10-15 kW; there is no universal liquid-cooling threshold. See Uptime Institute's overview of AI cooling methods and capacities.
| Cooling method | How it works | Best fit | Main constraint |
|---|---|---|---|
| Enhanced air cooling | Containment, higher airflow, better air management | Air-cooled GPU equipment within facility limits | Airflow, noise, fan power, and hot spots |
| Rear-door heat exchanger | Removes exhaust heat at the rack rear | Targeted density upgrades in existing facilities | Piping, door weight, condensation, and service access |
| Direct-to-chip liquid cooling | Cold plates cool CPUs and GPUs; a coolant distribution unit (CDU) separates the building and equipment cooling circuits | High-density GPU and rack-scale systems | Water quality, temperature, flow, pressure, quick connects, leak detection, and redundancy |
| Immersion cooling | Compatible equipment is submerged in electrically non-conductive fluid | Specialized high-density deployments | Compatibility, maintenance, materials, warranty, and fluid management |
| Hybrid cooling | Liquid cools CPUs and GPUs; air cools network adapters, storage, and other components | Most current direct-to-chip racks | Both air-side and liquid-side capacity remain necessary |
High-density GPU rack design
Open rack space does not prove that equipment can be installed safely. Validate static and dynamic rack and floor loading, equipment depth, rail and cable clearance, vertical rPDU placement, airflow direction, CDU and manifold interfaces, leak detection, cable bend radius, and service access. A DGX H100 is 8U, up to 130.45 kg, and 897.1 mm deep; NVIDIA's current design guide calls for at least a 600 × 1,200 mm, 48U rack.
GPU architecture determines data-center design
Layer 1: Scale-up within a server
An 8-GPU server may connect its GPUs through NVLink and NVSwitch. When a model's calculations are split across GPUs (tensor parallelism), those GPUs exchange results frequently. Their connection layout therefore affects performance and should be part of the selection.
Layer 2: Rack-scale architecture
GB200 and GB300 NVL72 integrate compute trays, NVLink switch trays, power shelves, bus bars, and liquid-cooling components into a 72-GPU NVLink domain. The rack is planned as one system. See the NVIDIA DGX GB rack-scale documentation.
Layer 3: Scale-out across racks
Clusters spanning several servers need fast data exchange between them. This can use InfiniBand or high-speed Ethernet with RoCE, which allows direct data transfer between server memory. Network distance and cable routing must be recalculated when rack density or footprint changes.
High-speed networking and storage are part of the AIDC
GPUs produce useful work only when they receive data and synchronize continuously. Plan connections within GPU systems, connections between systems, and the management and storage networks separately. Test storage capacity and speed when data is read or written in order or from scattered locations. Also check combined transfer speed, handling of file information, saving job progress, recovery, and simultaneous reads by several servers.
Can a traditional IDC be upgraded for AI?
Yes, when the GPU fleet is small, air cooling remains viable, rack density can be reduced or distributed, power is deliverable through the existing chain, floor loading and access routes are adequate, and network and storage performance can be validated. A dedicated AIDC zone, major upgrade, or new build is usually more suitable when the system requires direct-to-chip cooling, rack demand greatly exceeds current distribution, low-latency multi-node topology is fixed, facility water cannot be added, or floor and access constraints cannot support the equipment.
AI Data Center planning and construction with Konstra AI
Konstra AI provides AI data-center infrastructure construction, covering electrical and mechanical systems, environmental controls, rack layout, and integrated GPU-cluster deployment, as well as equipment maintenance and compute delivery.
For high-density deployments, KONST can plan air- and liquid-cooled solutions above 20 kW per rack and provide services from turnkey AIDC construction through compute delivery and ongoing operations. Depending on project conditions, operations can begin in as little as four months, while operations follow an ISO 27001:2022 information security management system. New or expanded utility power requires a separate lead time for application and electrical works.
KONST also covers colocation and compute rental, allowing the deployment model to reflect site conditions, investment scale, and launch schedule.
Contact KONST. The KONST team will respond within three business days.
Related service: AI compute infrastructure construction.
FAQ
What is the biggest difference between an AI data center and a traditional data center?
The design order. An AIDC starts with the GPU system, cluster topology, and workload, then derives power, cooling, rack, network, and storage requirements. A traditional facility more often starts with general-purpose capacity.
Can a traditional IDC host GPU servers?
Often yes for a small number of air-cooled systems, using reduced density, distributed racks, or targeted upgrades. Feasibility still depends on rack-level power, cooling, floor loading, depth, networking, and expansion limits.
Above what rack density is liquid cooling mandatory?
There is no universal threshold. The answer depends on the component's rated heat load (TDP), server cooling design, rack density, inlet conditions, and facility cooling performance. Follow the equipment vendor's requirements and a project-specific thermal analysis.
Does every AIDC need NVLink and InfiniBand?
No. NVLink and NVSwitch serve specific scale-up architectures. Scale-out networks may use InfiniBand or high-speed Ethernet with RoCE. A single-server inference service and a large distributed training cluster have different requirements.
Should a company upgrade an existing facility or build a new AIDC?
Upgrade can be faster when the existing site can deliver the target rack power, cooling, loading, piping, and network. A dedicated zone or new build is usually more appropriate for high-density liquid-cooled rack-scale systems or when upstream power and space cannot expand. Compare deliverable GPU capacity, construction disruption, expansion ceiling, schedule, and total cost of ownership.



