Export Compliant| GCC · MEA · APAC| Business Bay, Dubai, UAE

Clusters

Sizing an AI Cluster for LLM Training

Parameter count, token budget, and utilisation assumptions translated into a concrete node count.

Sizing an AI cluster for LLM training is one of the most important decisions when building an enterprise AI infrastructure. Choosing the right number of GPUs requires more than estimating model size. A properly designed cluster must account for GPU memory, compute requirements, training tokens, batch size, sequence length, GPU interconnects, networking, storage, power, cooling, and target training time.
Whether you are training a small language model, fine-tuning an existing foundation model, or building a large-scale LLM from scratch, the objective is the same: deploy enough GPU capacity to meet your performance target without creating an unnecessarily expensive or underutilized cluster.
## How Many GPUs Do You Need to Train an LLM?
There is no universal GPU count for LLM training.
A practical estimate depends on:
* Number of model parameters
* Total training tokens
* Sequence length
* Training precision
* Batch size
* Target training duration
* GPU performance
* GPU memory capacity
* Inter-GPU communication efficiency
* Number of training nodes
* Desired GPU utilization
A model with several billion parameters may be trained effectively on a relatively small GPU cluster, while a frontier-scale model can require hundreds or thousands of accelerators.
The correct approach is to start with the training workload, calculate the required compute and memory, and then work backward to the number of GPUs.
## 1. Start With Model Size
The first input is the number of parameters.
For example:
* 7B–8B models: suitable for smaller research and enterprise training environments
* 13B–14B models: require substantially more memory and compute
* 30B–70B models: typically require multi-GPU infrastructure
* 100B+ models: require large distributed GPU clusters
* Frontier-scale models: can require hundreds or thousands of GPUs
Model size alone, however, does not determine training requirements.
Training a 70B model for a relatively small dataset is very different from training the same model on hundreds of billions or trillions of tokens.
## 2. Calculate Training Compute
A commonly used first-order estimate for dense Transformer training is approximately:
Training FLOPs ≈ 6 × model parameters × training tokens
This provides a useful starting point for estimating the total compute required.
For example, a hypothetical 70B-parameter model trained on 1 trillion tokens would require approximately:
6 × 70B × 1T = 4.2 × 10²³ FLOPs
This is an enormous amount of computation, which is why large-scale LLM training requires highly parallel GPU infrastructure.
The real-world requirement will vary depending on architecture, implementation, numerical precision, sequence length, parallelism strategy, and achieved hardware utilization.
## 3. GPU Performance Is Not the Same as Real Training Performance
GPU specifications often advertise enormous theoretical AI compute performance.
However, you should not simply divide the total required FLOPs by the GPU's theoretical FLOPS.
Real training performance depends on:
* Tensor Core utilization
* Memory bandwidth
* GPU memory capacity
* Communication overhead
* Kernel efficiency
* Data loading
* Parallelism strategy
* Network bandwidth
* Software optimization
A cluster achieving 35–50% of theoretical peak compute can sometimes be perfectly reasonable depending on the workload and measurement methodology.
For cluster sizing, the most useful number is therefore measured training throughput, such as tokens per second per GPU or tokens per second per cluster.
## 4. GPU Memory Determines How Large a Model Can Fit
GPU memory is one of the first constraints to evaluate.
During training, GPU memory is used for more than model weights.
It may contain:
* Model parameters
* Gradients
* Optimizer states
* Activations
* Temporary tensors
* Communication buffers
For mixed-precision training with Adam-style optimizers, the total memory requirement can be several times larger than the model's parameter size.
A simplified estimate can be expressed as:
Training memory ≈ model states + gradients + optimizer states + activations + overhead
This is why a model that appears to fit easily in GPU memory during inference may require substantially more memory during training.
## 5. H200, B200, B300 and MI355X
GPU memory capacity can have a major effect on cluster design.
For example, high-end data-center accelerators include:
* NVIDIA H200
* NVIDIA B200
* NVIDIA B300
* AMD Instinct MI300X
* AMD Instinct MI355X
The additional HBM capacity available on newer accelerators can allow larger models, longer contexts, or larger batch sizes to fit more efficiently.
For example, NVIDIA B200 systems provide up to 180 GB of HBM3e per GPU, while B300 platforms provide substantially more GPU memory. AMD's MI355X provides 288 GB of HBM3E per GPU.
This makes memory capacity an important consideration when comparing GPUs with similar theoretical compute performance.
## 6. Example: Estimating a 70B Training Cluster
Suppose you want to train a 70B-parameter model on 1 trillion tokens.
Using the simplified compute estimate:
6 × 70B × 1T = 4.2 × 10²³ FLOPs
Now assume your selected GPU cluster achieves an effective sustained training performance of approximately 500 TFLOPS per GPU for the actual workload.
The theoretical GPU-hours required would be approximately:
4.2 × 10²³ ÷ 5 × 10¹⁴ ≈ 840 million GPU-seconds
or approximately:
233,000 GPU-hours
This is only a rough planning calculation. Real-world training time will depend heavily on hardware utilization and distributed-training efficiency.
For example, a cluster of:
256 GPUs
could theoretically complete the workload much faster than a 64-GPU cluster, but only if the additional GPUs can be efficiently utilized.
This illustrates why simply adding GPUs does not always produce linear speed improvements.
## 7. Scaling Efficiency Matters
A cluster with 512 GPUs is not necessarily twice as fast as one with 256 GPUs.
As GPU count increases, communication overhead becomes more significant.
The relationship can be expressed as:
Effective performance = GPU compute × scaling efficiency
For example:
| GPU Count | Ideal Scaling | Example 85% Efficiency |
| --------: | ------------: | ---------------------: |
| 64 | 64× | 54.4× |
| 128 | 128× | 108.8× |
| 256 | 256× | 217.6× |
| 512 | 512× | 435.2× |
The actual efficiency depends heavily on the model, parallelism strategy, network, GPU interconnect, and software stack.
This is why a smaller cluster with excellent communication can outperform a larger cluster with poor scaling.
## 8. GPU Interconnect Is Critical
For multi-GPU training, GPU-to-GPU communication is essential.
Inside high-end servers, technologies such as NVIDIA NVLink provide high-bandwidth communication between GPUs.
Between servers, the cluster requires a high-performance network such as:
* NVIDIA InfiniBand
* High-performance Ethernet
* RoCE
* AI-optimized Ethernet fabrics
Large distributed workloads commonly use communication operations such as:
* All-reduce
* All-gather
* Reduce-scatter
* All-to-all
The larger the cluster, the more important network performance becomes.
## 9. Choose the Right Node Size
Many enterprise AI clusters use 8-GPU servers as their fundamental compute node.
For example:
8 GPUs × 32 servers = 256 GPUs
or:
8 GPUs × 64 servers = 512 GPUs
This approach simplifies deployment, networking, scheduling, and maintenance.
However, the ideal node configuration depends on the GPU architecture and the training framework.
For very large models, tightly connected GPU systems can reduce communication overhead and improve scaling.
## 10. Storage Requirements
Training datasets can be extremely large.
The storage architecture should support:
* Training datasets
* Model checkpoints
* Tokenized datasets
* Experiment outputs
* Logs
* Evaluation data
* Model artifacts
Storage throughput is particularly important during:
Dataset loading + checkpointing + validation
If storage cannot keep up with the GPUs, expensive accelerators may remain idle.
For large clusters, organizations may use high-performance parallel file systems or object-storage architectures combined with local NVMe caching.
## 11. Networking Requirements
The network should be sized according to GPU count and communication intensity.
A small development cluster might use high-speed Ethernet, while a large distributed training environment may require:
* 400 Gb/s networking
* 800 Gb/s networking
* InfiniBand
* RoCE
* RDMA-capable NICs or SuperNICs
* Dedicated GPU cluster fabrics
The network should be benchmarked using actual collective communication workloads rather than relying exclusively on theoretical link bandwidth.
## 12. Training Parallelism
Large LLMs typically use several forms of parallelism.
### Data Parallelism
Different GPUs process different batches of training data.
### Tensor Parallelism
A model's tensors are divided across multiple GPUs.
### Pipeline Parallelism
Different sections of the model are assigned to different GPUs.
### Expert Parallelism
Used by Mixture-of-Experts architectures to distribute experts across GPUs.
Modern large-scale training often combines several of these strategies.
The cluster must therefore be designed around the parallelism approach that the model and training framework will use.
## 13. Power and Cooling
GPU count directly affects infrastructure requirements.
A 64-GPU cluster and a 512-GPU cluster are fundamentally different data-center projects.
Calculate:
GPU power + CPU power + memory + networking + storage + cooling overhead
High-density AI systems can require advanced cooling, including direct-to-chip liquid cooling, especially when rack power density becomes too high for conventional air cooling.
Cooling should be considered before selecting the final rack architecture.
## 14. Cluster Sizing by Workload
A useful high-level planning model is:
### Small AI Training Cluster
8–32 GPUs
Suitable for:
* Fine-tuning
* Smaller LLMs
* Research
* Enterprise AI development
* Prototyping
### Medium AI Training Cluster
32–128 GPUs
Suitable for:
* Larger model training
* Advanced fine-tuning
* Large-scale experiments
* Enterprise foundation models
### Large AI Training Cluster
128–512 GPUs
Suitable for:
* Large LLM training
* Multimodal models
* Large-scale generative AI
* Advanced research
### Frontier-Scale Cluster
512–1,000+ GPUs
Designed for:
* Very large foundation models
* Massive training datasets
* Advanced reasoning models
* Frontier AI research
These ranges are planning guidelines rather than strict requirements.
## 15. Don't Optimize for Maximum GPU Count
One of the most common mistakes is buying as many GPUs as the budget allows.
More GPUs only help if the workload can use them efficiently.
A better strategy is to determine:
Target training time → required throughput → required GPU count → network → storage → power/cooling
For example, if a training job requires 10 million tokens per second, determine how many tokens per second your selected GPU configuration can realistically deliver.
Then calculate:
Required GPUs = Target cluster throughput ÷ Measured throughput per GPU
Adjust the result for expected scaling efficiency.
## 16. Calculate Total Cost of Ownership
GPU acquisition is only part of the cost.
Your TCO should include:
* GPUs
* GPU servers
* CPUs
* Memory
* Networking
* Switches
* Optics
* Storage
* Racks
* Power distribution
* Cooling
* Data-center space
* Software
* Support
* Maintenance
* Electricity
* Replacement hardware
A cluster with fewer GPUs but better utilization can deliver lower cost per trained token than a larger cluster.
## 17. The Most Important Metric: Cost per Training Token
For LLM infrastructure, a useful business metric is:
Cost per trained token
rather than simply:
GPU price
Consider:
Total training cost ÷ number of training tokens
This captures the combined impact of:
* GPU performance
* GPU utilization
* Network efficiency
* Power
* Cooling
* Software
* Infrastructure costs
The same methodology can be applied to inference using metrics such as cost per million tokens or tokens per second per GPU.
## 18. A Practical AI Cluster Sizing Workflow
Use the following process when designing a new cluster:
### Step 1 — Define the Model
Determine:
* Parameter count
* Architecture
* Context length
* Number of experts
* Training tokens
### Step 2 — Calculate Memory
Estimate:
* Parameters
* Gradients
* Optimizer states
* Activations
* Communication buffers
### Step 3 — Calculate Compute
Estimate total training FLOPs.
### Step 4 — Select the GPU
Compare:
* HBM capacity
* HBM bandwidth
* AI compute
* Interconnect
* Power
* Price
### Step 5 — Benchmark
Measure actual tokens/second on representative workloads.
### Step 6 — Determine GPU Count
Calculate the number of GPUs required to reach the target training time.
### Step 7 — Validate Scaling
Benchmark the workload across multiple nodes.
### Step 8 — Design Networking
Size the InfiniBand or RoCE/Ethernet fabric.
### Step 9 — Design Storage
Ensure dataset and checkpoint throughput can keep up with the cluster.
### Step 10 — Validate Power and Cooling
Confirm the facility can support the final rack density.
## Final Takeaway
Sizing an AI cluster for LLM training is a systems-engineering problem, not simply a GPU purchasing exercise.
The right cluster balances GPU compute, HBM memory, GPU interconnects, networking, storage, software efficiency, power, cooling, and budget.
For smaller models, an 8-GPU server may be sufficient. For larger models and extensive training datasets, dozens or hundreds of GPUs may be required. At very large scales, the performance of the network and the efficiency of distributed training can become just as important as the GPU itself.
The best approach is to start with the target model and training workload, estimate the required compute and memory, benchmark the selected GPU, measure multi-node scaling, and then size the cluster around the required training throughput and completion time.
In 2026, platforms based on NVIDIA H200, B200, B300, GB200/GB300 and AMD Instinct MI300X/MI355X provide multiple paths for building LLM training infrastructure. The optimal choice ultimately depends on the model architecture, memory requirements, expected scale, software ecosystem, and total cost per trained token.

Ready to Start?

Build Your AI Infrastructure With Confidence

Talk to our enterprise infrastructure team. Get expert guidance, GPU pricing, and a custom deployment plan — no commitment required.

Live chat Enterprise infrastructure team
eCirclec