For your model's next step,
how much GPU memory is needed?
Compare the capacity needs of inference and LoRA fine-tuning, with item-by-item estimates for model weights, KV Cache, and GPU count.
- Model Weights
- 16.0 GB
- KV Cache
- 0.5 GB
- Activations
- 3.2 GB
- Framework reserve
- 1.5 GB
Illustrative figures for capacity planning, not measured results.
Sites that can meet this requirement
The tool calculates the number of cards needed from the estimated GPU memory, then lists sites with enough cards and the earliest delivery.
How many GPUs do you need?
This table is based only on total GPU memory. Multi-card configurations still require model sharding plus framework and interconnect support, and are not guaranteed to run directly.
| GPU | GPU memory / card | Minimum cards | Total capacity | Capacity usage |
|---|---|---|---|---|
| B300Deliverable | 288 GB | 1 | 288 GB | 7.4% |
| B200 | 180 GB | 1 | 180 GB | 11.8% |
| H200Deliverable | 141 GB | 1 | 141 GB | 15.1% |
| H100 NVL | 94 GB | 1 | 94 GB | 22.6% |
| A100 80GB | 80 GB | 1 | 80 GB | 26.5% |
| H100 SXM5Deliverable | 80 GB | 1 | 80 GB | 26.5% |
| L40S | 48 GB | 1 | 48 GB | 44.2% |
| A100 40GB | 40 GB | 1 | 40 GB | 53.1% |
| RTX 4090 | 24 GB | 1 | 24 GB | 88.5% |
State the assumptions behind the estimate
View formulas and limitations
GB uses decimal units. Weights = parameters × precision bytes; KV Cache = 2 × layers × KV Heads × Head Dim × context length × Batch Size × 2 bytes. The cache is fixed at FP16.
Inference activation is estimated at 20% of the weights; for LoRA fine-tuning it is 25% with Gradient Checkpointing enabled and 60% with it disabled. LoRA Rank 8/16/32/64 corresponds to estimated ratios of 2%/4%/6%/8%, and LoRA state = parameters × ratio × 2 × 13 bytes. The framework reserve is fixed at 1.5 GB. Fine-tuning mode also includes the KV Cache, which is not the same as cache behavior during actual training.
A custom model's architecture is approximated from its parameter count and cannot replace the model configuration file.
The LoRA ratio and activation are both rough estimates. Actual usage depends on the framework, model architecture, quantization, and optimizer. INT8/INT4 fine-tuning requires a compatible quantized LoRA toolchain. This tool does not estimate training time or run training jobs.
Plan the next stage of compute together
Bring your estimate to the sales team and discuss specs