Embeddings
data workload
Encoding text or other inputs into vectors. Smaller models, high throughput; often CPU-viable but GPU-accelerated at scale.
Classification
- Category
- data
- Compute profile
- compute-bound
- Scaling
- batch
Infrastructure considerations
Qualitative — no fabricated numbers- Interconnect sensitivity
- low
- Network sensitivity
- low
Curated · Memory · Small model weights + batch activations; fits comfortably on modest GPUs.
Curated · Storage · Streaming input; vector output to a store or database.
Typical frameworks
Curated — not a catalog relationshipHugging FacePyTorchTensorRT
Data class
Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.
