Skip to main content

Text-to-Speech

audio workload

Curated

Synthesizing speech from text. Latency-sensitive for interactive use; models are typically small relative to LLMs.

Classification

Category
audio
Compute profile
mixed
Scaling
horizontal

Infrastructure considerations

Qualitative — no fabricated numbers
Interconnect sensitivity
low
Network sensitivity
low

Curated · Memory · Small-to-mid model weights + activations; latency-driven rather than memory-driven.

Curated · Storage · Model weights plus audio output storage.

Typical frameworks

Curated — not a catalog relationship
PyTorchHugging Face

Data class

Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.