Skip to main content

RAG (Retrieval-Augmented Generation)

inference workload

Curated

Combining a retrieval step over a vector store with LLM generation. Infrastructure spans an embedding/retrieval path and an LLM inference path.

Classification

Category
inference
Compute profile
mixed
Scaling
horizontal

Infrastructure considerations

Qualitative — no fabricated numbers
Interconnect sensitivity
low
Network sensitivity
medium

Curated · Memory · Dominated by the LLM inference component; embeddings are lightweight.

Curated · Storage · Vector database plus the document corpus; low-latency retrieval.

Typical frameworks

Curated — not a catalog relationship
vLLMHugging FaceRay

Relevant models

Models that run this workload

Data class

Curated · The workload taxonomy is authored qualitative reference, not an ingested source. Sensitivities and profiles are classifications, not measurements; a workload's concrete VRAM comes from the specific model it runs.