LLM Training
training workload
Training a large language model from scratch. Weights, gradients and optimizer state multiply memory demand; almost always multi-GPU and multi-node.
Classification
- Category
- training
- Compute profile
- compute-bound
- Scaling
- horizontal
Infrastructure considerations
Qualitative — no fabricated numbers- Interconnect sensitivity
- high
- Network sensitivity
- high
Curated · Memory · Weights + gradients + optimizer state (often 3-4x the weights); sharded across GPUs.
Curated · Storage · High-throughput shared storage for datasets and frequent checkpoints.
Typical frameworks
Curated — not a catalog relationshipPyTorchDeepSpeedMegatronRay
Relevant models
Models that run this workloadData class
Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.
