RAG (Retrieval-Augmented Generation)
inference workload
Combining a retrieval step over a vector store with LLM generation. Infrastructure spans an embedding/retrieval path and an LLM inference path.
Classification
- Category
- inference
- Compute profile
- mixed
- Scaling
- horizontal
Infrastructure considerations
Qualitative — no fabricated numbers- Interconnect sensitivity
- low
- Network sensitivity
- medium
Curated · Memory · Dominated by the LLM inference component; embeddings are lightweight.
Curated · Storage · Vector database plus the document corpus; low-latency retrieval.
Typical frameworks
Curated — not a catalog relationshipvLLMHugging FaceRay
Relevant models
Models that run this workloadData class
Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.
