KONST
AI Training CalculatorContact us
EN
Compute Delivery

Three compute delivery modes
Bare metal, virtualization and containers, token metering

KONST Group offers three compute delivery options: bare-metal clusters supplied directly by KONST, the token-metered platform of Horizon AI, which the group developed in-house, and virtualization and container resources from partner Glows.ai. GPU models include B300, B200, H200, and H100.

GPUs delivered to date3,856GPUsLargest single project: 1,024 GPUs
SOLUTION | KONST's approach

Three delivery modes,
from thousand-GPU clusters to a token API key

Start from the workload: training that runs at full load for long periods takes bare metal, work that needs flexible quota takes virtualization and containers, and work that only needs model output takes token metering. The three modes can be combined in one project.

01KONST · GPU Cluster

Bare-metal delivery

Exclusive hardware for the whole batch, from a single node to a thousand-GPU cluster, with InfiniBand or RoCE v2 high-speed interconnect. The equipment is housed in KONST data centers, and KONST keeps ownership.

Bare-metal delivery specs
02Glows.ai · GPU Cloud

Virtualization and container delivery

Use GPUs through virtual machines or containers, with snapshots and VPC isolation built in and nothing to build or maintain yourself. The platform is provided by partner Glows.ai.

Virtualization and container specs
03Horizon AI · ATP Token

Token-metered delivery

The group's in-house ATP Token platform: one API key calls more than 70 open-source models, metered per request, with no hardware to obtain.

Token metering specs
PROCESS | Workflow

From workload assessment to activation,
delivery completed in four stages

Confirm the workload type first, then work back to the delivery mode and GPU model. Activation differs by mode: bare metal hands over access, the platform activates quota, and token metering issues a project key.

  1. 01Workload assessment

    Confirm whether the workload is training, fine-tuning, or inference, along with expected usage hours and data storage requirements.

  2. 02Choose mode and GPU model

    Work back from the workload to the delivery mode and GPU model. One project can span modes, for example training on bare metal and inference on token metering.

  3. 03Activation and handover

    Bare metal is handed over for use through a dedicated VPN after the initial operating system installation. For virtualization and containers, quota is activated on the platform. For token metering, a project key is issued and takes effect immediately.

  4. 04Usage and operations

    Bare metal falls under operations from the handover date. Token usage is attributed by organization, workspace, and project, and shown project by project on the usage dashboard.

EVIDENCE | Track record and data

3,856 GPUs delivered in total,
largest single project 1,024 GPUs

The GPU count and the largest single project use the same basis as the GPU cluster build page. Projects cover Taiwan, Japan, and Malaysia. The model count is the union of the bare-metal and virtualization modes.

3,856GPUsGPUs delivered to date
1,024GPUsLargest single project (Malaysia, B300)
8modelsAvailable GPU models
70+modelsOpen-source models on the token platform
FEATURE | Core feature 01

Bare-metal delivery

KONST · GPU Cluster

Exclusive hardware for the whole batch, supplied directly by KONST. The equipment is housed in KONST data centers, access is handed over through a dedicated VPN once the initial operating system installation is complete, and KONST keeps ownership of the equipment.

Single node to thousand-GPU cluster

Delivered as a whole batch, from a single server up to a cluster at the thousand-GPU level

InfiniBand · RoCE v2

High-speed interconnect across nodes, selected by cluster size

High-speed storage

Clusters are delivered together with AI high-speed storage

Dedicated VPN handover

Separate network and dedicated connection, with data staying at the designated site

Available GPUs
  • HGX B300Latest generation
  • HGX B200
  • HGX H200
  • HGX H100
  • SXM A100
1 node1,000+ GPU

Single cluster scale · Largest single project 1,024 B300 GPUs

  • InfiniBand
  • RoCE v2
  • AI high-speed storage
FEATURE | Core feature 02

Virtualization and container delivery

Glows.ai · GPU Cloud

Use GPUs through virtual machines or containers. The platform is ready on activation, quota is allocated per project, billing is by usage time, and there is nothing to build or maintain yourself. The platform is provided by partner Glows.ai.

Virtual machine and container environments

Two environments, virtual machines and containers, chosen by workload

Automated deployment

Compute environments deploy automatically and are ready on activation

Datadrive and snapshots

Datadrive persistent storage, and environments can be saved as snapshots

VPC and Matrix0 database

Private network isolation, with a database service built in

Available GPUs
  • HGX B200
  • HGX H200
  • HGX H100
  • RTX PRO 6000
  • SXM A100
  • L40S
  • RTX 6000 Ada
Platform services
  • VMs
  • Container environments
  • Automated deployment
  • Datadrive
  • Snapshots
  • Matrix0 database
  • VPC
FEATURE | Core feature 03

Token-metered delivery

Horizon AI · ATP Token

One Key, Every Model. The group's in-house ATP Token platform lets one project API key call more than 70 open-source models, metered per request. Existing code only needs its endpoint changed, and no hardware is required.

One Key, Every Model

Compatible with OpenAI, Anthropic, and Gemini interfaces

Project quotas and credits

Prepaid, no monthly fee, billed by request usage

Usage dashboard

View usage level by level: organization, workspace, project

Request-level audit log

Every request is traceable and auditable

Organization structure

OrganizationWorkspaceProjectKey

70+Connected open-source models
Agent capabilities
  • Multi-model access
  • MCP tool integration
  • A2A protocol
  • Agentic RAG
  • Agent workflows
  • Human confirmation
  • Permission boundaries
  • Usage governance
Supported applications
  • Coding agent
  • Deep research agent
  • Browser agent
  • Multi-agent orchestrator
FEATURE | Use-case comparison

Comparing the three GPU compute delivery modes by use case

The deciding factor is the load curve: choose bare metal when usage runs close to capacity for long periods, virtualization and containers when usage swings widely, and token metering when you only need model output.

Comparison itemBare-metal deliveryVirtualization and container deliveryToken-metered delivery
What you getGPU servers, exclusive for the whole batchVirtual machines or containers used within a quotaModel output, no hardware included
Workload typeFull load over the long term, running for months without stoppingIntermittent or peak loads, with usage rising and falling with demandSporadic or irregular inference requests
Typical usesLarge-model pretraining and long fine-tuning runs; multi-node distributed training; high-throughput batch inferenceScientific computing such as molecular dynamics; quantitative trading strategy backtesting; teaching, research, and course environmentsEnterprise AI application and service development; agent workflows and automation; multi-model comparison and evaluation
Hardware controlFull control; you install frameworks and drivers yourselfYou configure within the platform environmentNo hardware or model deployment to manage
Data storageStays at the designated site, dedicated VPN connectionVPC isolation within the platformTiered permission control by project key
Billing basisBy GPU count and contract termBy usage timeBy request usage, prepaid, no monthly fee
Activation methodHanded over by VPN after the initial operating system installationReady on activation through the platformIssuing a project key takes effect immediately
BENEFIT | Common safeguards

All three modes share the same efficiency, compliance, and security

Efficiency, compliance, and security apply the same way in every delivery mode. The benchmark is the operations and audit capability a company would have to supply itself when building and maintaining equivalent compute in-house.

Efficiency

Cost optimization

  • Monitoring platform
  • Global compute supply
  • Range of choices
Compliance

Enterprise-grade operations

  • ISO 27001
  • Request-level audit log
  • 24x7 operations
Security

Permissions and isolation

  • Tiered key permissions
  • Team edition permission isolation
  • Private cloud / hybrid cloud
FAQ | Common questions

FAQ

Can one project use two or more delivery modes at the same time?

Yes. For example, training can run on bare metal and inference on token metering, with resources and billing kept separate. The three modes belong to different service lines, so they do not have to be tied to one contract.

If I only need model output, do I still need to rent GPUs first?

No. Token-metered delivery bills per request, so you do not need to get a whole machine or deploy and operate anything yourself. Bare metal or virtualization and container resources are needed only when you require exclusive hardware or want to deploy your own models.

Which GPU models are available through these three modes?

Bare metal covers HGX B300, B200, H200, H100, and SXM A100. Virtualization and containers cover HGX B200, H200, H100, RTX PRO 6000, SXM A100, L40S, and RTX 6000 Ada. Token metering is billed by model and needs no GPU model specified.

How many GPUs can a bare-metal cluster have at most?

From a single node up to a thousand-GPU cluster. The largest single project is 1,024 B300 GPUs in Malaysia. The high-speed interconnect is InfiniBand or RoCE v2, selected by cluster size and training needs.

How does virtualization and container delivery differ from bare-metal delivery?

The difference is whether the hardware is exclusive. Bare metal is exclusive for the whole batch, which suits training that runs at full load for months. Virtualization and containers allocate quota per project and bill by usage time, which suits intermittent loads and short-term validation.

How does token-metered delivery control each team's usage?

Usage is attributed across four levels: organization, workspace, project, and key. Each project can have a quota and a credit limit. The usage dashboard shows usage project by project, and the request-level audit log lets you check every request individually.

Can it be deployed on a private or hybrid cloud?

Yes. When data must not leave a designated site, a private cloud or hybrid cloud architecture can be used, with tiered key permissions and team edition permission isolation.

Still evaluating which delivery mode to use?

Tell us your workload type, expected usage hours, and data storage requirements, and we will reply with a recommended delivery mode and GPU configuration.