Three compute delivery modes
Bare metal, virtualization and containers, token metering
Bare metal, virtualization
KONST Group offers three compute delivery options: bare-metal clusters supplied directly by KONST, the token-metered platform of Horizon AI, which the group developed in-house, and virtualization and container resources from partner Glows.ai. GPU models include B300, B200, H200, and H100.
Three delivery modes,
from thousand-GPU clusters to a token API key
Start from the workload: training that runs at full load for long periods takes bare metal, work that needs flexible quota takes virtualization and containers, and work that only needs model output takes token metering. The three modes can be combined in one project.
Bare-metal delivery
Exclusive hardware for the whole batch, from a single node to a thousand-GPU cluster, with InfiniBand or RoCE v2 high-speed interconnect. The equipment is housed in KONST data centers, and KONST keeps ownership.
Virtualization and container delivery
Use GPUs through virtual machines or containers, with snapshots and VPC isolation built in and nothing to build or maintain yourself. The platform is provided by partner Glows.ai.
Token-metered delivery
The group's in-house ATP Token platform: one API key calls more than 70 open-source models, metered per request, with no hardware to obtain.
From workload assessment to activation,
delivery completed in four stages
Confirm the workload type first, then work back to the delivery mode and GPU model. Activation differs by mode: bare metal hands over access, the platform activates quota, and token metering issues a project key.
- 01Workload assessment
Confirm whether the workload is training, fine-tuning, or inference, along with expected usage hours and data storage requirements.
- 02Choose mode and GPU model
Work back from the workload to the delivery mode and GPU model. One project can span modes, for example training on bare metal and inference on token metering.
- 03Activation and handover
Bare metal is handed over for use through a dedicated VPN after the initial operating system installation. For virtualization and containers, quota is activated on the platform. For token metering, a project key is issued and takes effect immediately.
- 04Usage and operations
Bare metal falls under operations from the handover date. Token usage is attributed by organization, workspace, and project, and shown project by project on the usage dashboard.
3,856 GPUs delivered in total,
largest single project 1,024 GPUs
The GPU count and the largest single project use the same basis as the GPU cluster build page. Projects cover Taiwan, Japan, and Malaysia. The model count is the union of the bare-metal and virtualization modes.
Bare-metal delivery
KONST · GPU Cluster
Exclusive hardware for the whole batch, supplied directly by KONST. The equipment is housed in KONST data centers, access is handed over through a dedicated VPN once the initial operating system installation is complete, and KONST keeps ownership of the equipment.
Single node to thousand-GPU cluster
Delivered as a whole batch, from a single server up to a cluster at the thousand-GPU level
InfiniBand · RoCE v2
High-speed interconnect across nodes, selected by cluster size
High-speed storage
Clusters are delivered together with AI high-speed storage
Dedicated VPN handover
Separate network and dedicated connection, with data staying at the designated site
- HGX B300Latest generation
- HGX B200
- HGX H200
- HGX H100
- SXM A100
Single cluster scale · Largest single project 1,024 B300 GPUs
- InfiniBand
- RoCE v2
- AI high-speed storage
Virtualization and container delivery
Glows.ai · GPU Cloud
Use GPUs through virtual machines or containers. The platform is ready on activation, quota is allocated per project, billing is by usage time, and there is nothing to build or maintain yourself. The platform is provided by partner Glows.ai.
Virtual machine and container environments
Two environments, virtual machines and containers, chosen by workload
Automated deployment
Compute environments deploy automatically and are ready on activation
Datadrive and snapshots
Datadrive persistent storage, and environments can be saved as snapshots
VPC and Matrix0 database
Private network isolation, with a database service built in
- HGX B200
- HGX H200
- HGX H100
- RTX PRO 6000
- SXM A100
- L40S
- RTX 6000 Ada
- VMs
- Container environments
- Automated deployment
- Datadrive
- Snapshots
- Matrix0 database
- VPC
Token-metered delivery
Horizon AI · ATP Token
One Key, Every Model. The group's in-house ATP Token platform lets one project API key call more than 70 open-source models, metered per request. Existing code only needs its endpoint changed, and no hardware is required.
One Key, Every Model
Compatible with OpenAI, Anthropic, and Gemini interfaces
Project quotas and credits
Prepaid, no monthly fee, billed by request usage
Usage dashboard
View usage level by level: organization, workspace, project
Request-level audit log
Every request is traceable and auditable
Organization›Workspace›Project›Key
- Multi-model access
- MCP tool integration
- A2A protocol
- Agentic RAG
- Agent workflows
- Human confirmation
- Permission boundaries
- Usage governance
- Coding agent
- Deep research agent
- Browser agent
- Multi-agent orchestrator
Comparing the three GPU compute delivery modes by use case
The deciding factor is the load curve: choose bare metal when usage runs close to capacity for long periods, virtualization and containers when usage swings widely, and token metering when you only need model output.
| Comparison item | Bare-metal delivery | Virtualization and container delivery | Token-metered delivery |
|---|---|---|---|
| What you get | GPU servers, exclusive for the whole batch | Virtual machines or containers used within a quota | Model output, no hardware included |
| Workload type | Full load over the long term, running for months without stopping | Intermittent or peak loads, with usage rising and falling with demand | Sporadic or irregular inference requests |
| Typical uses | Large-model pretraining and long fine-tuning runs; multi-node distributed training; high-throughput batch inference | Scientific computing such as molecular dynamics; quantitative trading strategy backtesting; teaching, research, and course environments | Enterprise AI application and service development; agent workflows and automation; multi-model comparison and evaluation |
| Hardware control | Full control; you install frameworks and drivers yourself | You configure within the platform environment | No hardware or model deployment to manage |
| Data storage | Stays at the designated site, dedicated VPN connection | VPC isolation within the platform | Tiered permission control by project key |
| Billing basis | By GPU count and contract term | By usage time | By request usage, prepaid, no monthly fee |
| Activation method | Handed over by VPN after the initial operating system installation | Ready on activation through the platform | Issuing a project key takes effect immediately |
All three modes share the same efficiency, compliance, and security
Efficiency, compliance, and security apply the same way in every delivery mode. The benchmark is the operations and audit capability a company would have to supply itself when building and maintaining equivalent compute in-house.
Cost optimization
- Monitoring platform
- Global compute supply
- Range of choices
Enterprise-grade operations
- ISO 27001
- Request-level audit log
- 24x7 operations
Permissions and isolation
- Tiered key permissions
- Team edition permission isolation
- Private cloud / hybrid cloud
Related services
Besides compute delivery, data center build, equipment colocation, and enterprise AI adoption each have their own service line and can be commissioned together in one project.
FAQ
Can one project use two or more delivery modes at the same time?
Yes. For example, training can run on bare metal and inference on token metering, with resources and billing kept separate. The three modes belong to different service lines, so they do not have to be tied to one contract.
If I only need model output, do I still need to rent GPUs first?
No. Token-metered delivery bills per request, so you do not need to get a whole machine or deploy and operate anything yourself. Bare metal or virtualization and container resources are needed only when you require exclusive hardware or want to deploy your own models.
Which GPU models are available through these three modes?
Bare metal covers HGX B300, B200, H200, H100, and SXM A100. Virtualization and containers cover HGX B200, H200, H100, RTX PRO 6000, SXM A100, L40S, and RTX 6000 Ada. Token metering is billed by model and needs no GPU model specified.
How many GPUs can a bare-metal cluster have at most?
From a single node up to a thousand-GPU cluster. The largest single project is 1,024 B300 GPUs in Malaysia. The high-speed interconnect is InfiniBand or RoCE v2, selected by cluster size and training needs.
How does virtualization and container delivery differ from bare-metal delivery?
The difference is whether the hardware is exclusive. Bare metal is exclusive for the whole batch, which suits training that runs at full load for months. Virtualization and containers allocate quota per project and bill by usage time, which suits intermittent loads and short-term validation.
How does token-metered delivery control each team's usage?
Usage is attributed across four levels: organization, workspace, project, and key. Each project can have a quota and a credit limit. The usage dashboard shows usage project by project, and the request-level audit log lets you check every request individually.
Can it be deployed on a private or hybrid cloud?
Yes. When data must not leave a designated site, a private cloud or hybrid cloud architecture can be used, with tiered key permissions and team edition permission isolation.
Still evaluating which delivery mode to use?
Tell us your workload type, expected usage hours, and data storage requirements, and we will reply with a recommended delivery mode and GPU configuration.