KONST
AI Training CalculatorContact us
EN
← Back to BlogBLOG | GPU compute · 2026-09-01

How to Choose GPU Bare-Metal Rental in Taiwan: H100, H200, B200, and B300

Compare H100, H200, B200, and B300 for GPU bare-metal rental in Taiwan, including memory, delivery models, pricing, benchmarks, and SLA considerations.

台灣 GPU 裸機租賃怎麼選:H100、H200、B200、B300 比較

When companies evaluate GPU bare-metal rental in Taiwan, the first decision is often whether to use H100, H200, B200, or B300. The same GPU may also be delivered as a single GPU, an 8-GPU HGX server, or a virtualized resource. A useful comparison therefore asks three questions: Does the model fit in memory? How many GPUs are required to finish the job? What is the total cost per training run or per million tokens?

H100 remains a practical choice for testing models and adapting them to a specific task through fine-tuning. H200 suits models limited by memory capacity or speed, including those that process long inputs. B200 targets training and inference that need to process more work in less time, once the software has been checked for Blackwell compatibility. B300 is designed for work that needs more than 180 GB per GPU, including reasoning tasks that use extra computation while generating an answer (test-time compute).

What is GPU bare-metal rental?

GPU bare-metal rental gives a company exclusive use of a physical server. The customer controls the GPU, CPU, system memory, local storage, drivers, containers, and cluster software without sharing the host with general multi-tenant virtual machines.

  • Predictable GPU performance and dedicated capacity
  • Full administrator access to configure GPU software and drivers, containers, and job-management tools such as CUDA, Docker, Kubernetes, or Slurm
  • Access to fast connections between GPUs or servers, such as NVLink, NVSwitch, or InfiniBand, where supported
  • Long-running training or steady inference workloads
  • Clear requirements for data isolation, software versions, and operational control

H100, H200, B200, and B300 specifications and delivery models

GPU / architectureGPU memoryCommon delivery unitTypical fit
NVIDIA H100 SXM / Hopper80 GB HBM3; about 3.35 TB/s1 GPU or 8-GPU HGXValidation, fine-tuning, mature training and inference
NVIDIA H200 SXM / Hopper141 GB HBM3e; 4.8 TB/s1 GPU or 8-GPU HGXLarge models, long context, memory-bound workloads
NVIDIA B200 SXM / Blackwell180 GB HBM3e; up to 8 TB/sUsually 8-GPU HGX; smaller units may be virtualizedLarge-scale training, high-throughput inference, multimodal models
NVIDIA B300 SXM / Blackwell Ultra288 GB HBM3e; up to 8 TB/s8-GPU HGX B300; high-power systems may require liquid coolingAI reasoning, very large models, test-time compute

How to choose among H100, H200, B200, B300, GB200, and RTX PRO 6000

  • H100: Mature software and deployment experience. A strong baseline for model testing, fine-tuning, retrieval-augmented generation (RAG, which uses retrieved information to help answer questions), image generation, and running models for enterprise applications.
  • H200: Keeps the Hopper software environment while increasing memory to 141 GB HBM3e. It is useful for longer inputs, processing more items together, and storing more of the intermediate information reused during text generation (the KV cache).
  • B200: Best evaluated with the exact model and framework after Blackwell compatibility testing. A higher headline performance figure does not automatically mean lower cost per completed job.
  • B300: Provides 288 GB HBM3e per GPU and targets memory-intensive reasoning and test-time compute. Power density is a deployment constraint: high-power 8-GPU nodes may exceed 14 kW and require data-center power and cooling validation.
  • GB200 NVL72: A liquid-cooled rack-scale system with 36 Grace CPUs and 72 Blackwell GPUs. It requires the rack, power, cooling, networking, and delivery model to be assessed as one system.
  • RTX PRO 6000: Provides 96 GB GDDR7 ECC memory and fits enterprise inference, fine-tuning, scientific computing, 3D rendering, and virtual workstation scenarios.

KONST advantages for companies deploying AI infrastructure in Asia

Global compute platforms offer large GPU fleets and mature cloud tooling. Companies in Taiwan must also evaluate data location, cross-border connectivity, Chinese-language technical communication, the contracting entity, and the operating model after deployment.

KONST Group is a next-generation AI compute operator headquartered in Taipei, Taiwan. Through the One KONST brand and strategic partner ecosystem, it brings GPU bare-metal capacity and data-center operations together under one team and one order. The service is designed for companies that need dedicated resources, customized environments, and stable long-term capacity. Depending on scale, the deployment can extend from a bare-metal server to a dedicated GPU cluster, colocation, or AIDC planning. Available sites, GPU models, and delivery schedules are confirmed against current capacity.

How to compare GPU rental pricing

ItemWhat to confirm
BillingSpot, on-demand, monthly, long-term contract, or reserved capacity
Minimum unitOne GPU, one server, an 8-GPU node, or a full rack
InterruptionWhether capacity can be reclaimed and whether long jobs can resume
Included resourcesCPU, RAM, local NVMe, shared storage, bandwidth, IP addresses, and support
AvailabilityFacility tier, availability target, measurement method, and service credits
ContractMinimum term, deposit, payment terms, early termination, and expansion commitments

Compare providers under the same model and test conditions, then calculate a cost per completed task. Cost per million tokens = total GPU, storage, network, and required service cost / completed tokens × 1,000,000. For training, compare the total cost to complete the same number of passes through the training data (epochs), the same dataset, or the same target training error (loss).

A four-step process for selecting GPU bare metal

Step 1: Identify training, fine-tuning, or inference

Training needs memory for the model's learned values (weights), the information used to update them (gradients and optimizer states), and intermediate results (activations). For inference—using a trained model—the main considerations are model memory, the KV cache used during generation, output speed, the wait for the first output token, and how many requests run at once. Fine-tuning requirements depend on whether all model parameters or only a small subset are updated.

Step 2: Estimate memory and GPU coun

Estimate the highest GPU memory use. If the model does not fit on one GPU, compare using fewer bits to store its values (quantization), splitting the model across GPUs (model parallelism), or adding GPUs. Each choice affects performance and how much work the system takes to manage.

Step 3: Benchmark the real model

Use the same test conditions for each option. Record the GPU model and count, storage, network, and model version. Also record numerical precision, peak memory use, input and output lengths, items processed together (batch size), and simultaneous requests (concurrency). Compare work completed per second, response or completion time, and average GPU utilization.

Step 4: Compare cost per completed task

GPU/hour is only one billing input. Include idle time, storage, data transfer, environment setup, operations labor, long-term discounts, and interruption risk. The final metric should be cost per training run, completed task, or million tokens.

How KONST supports GPU selection and delivery

KONST starts with the model scale, workload, and data-processing requirements, then maps them to the GPU model, minimum rental unit, network, deployment conditions, schedule, and estimated cost. The scope covers deployment, testing, and post-launch support. Bronze, Silver, and Gold service tiers address different availability targets, hardware replacement windows, and support response times; the guaranteed values and service-credit terms are defined in the contract.

Contact KONST. The KONST team will respond within three business days.

Related service: GPU bare-metal rental.

FAQ

How should a company choose between H100 and H200?

Choose H100 when software maturity and broad workload support are the priority. Consider H200 when the workload is constrained by 80 GB of memory, requires longer context, or needs a larger KV cache. Validate the decision with the same model and inference settings.

Is B200 always more cost-effective than H200?

No. B200 can be more efficient for workloads that use Blackwell capabilities, but H200 may deliver a lower cost per completed job when software compatibility or utilization prevents the workload from using B200 effectively.

What is the difference between B200 and B300?

Start with the workload: prioritize B200 for evaluation when scientific simulations or mixed AI/HPC tasks require FP64 double-precision computation. Prioritize B300 for evaluation when LLM training, large-scale inference, or long-context workloads need more memory and low-precision compute. B300 offers higher dense FP4 Tensor Core performance but lower native FP64 performance, so a newer generation does not mean better performance for every task. B200 provides 180 GB of HBM3e per GPU, compared with 288 GB on B300; an 8-GPU HGX node provides 1.44 TB and 2.3 TB, respectively. For multi-node training, also verify the server network adapters (NICs) and how the cluster network is configured. Compare completion time, power consumption, and cost per completed job using the same model, precision, and software versions, and confirm facility power and cooling requirements. See the NVIDIA enterprise reference architecture.

How is GPU bare metal different from GPU Cloud?

Bare metal provides exclusive control of a physical server and is well suited to full administrator access, predictable performance, and customized cluster environments. GPU Cloud usually offers faster provisioning and more flexible billing. Always confirm whether resources are dedicated, the minimum rental unit, and who owns each operational responsibility.

GPU computeGPUBare MetalH100H200Blackwell

Plan the next stage of compute together

Where Efficiency Leads.
Where AI Defines the Future.