MODULE 3 · DAY 1
Production Serving with vLLM
Continuous batching + PagedAttention, same /v1 API
Gourav Shah · School of DevOps & AI · Containers for GenAI & Agentic AI
M3·01
What you'll learn
Batching, the contract, the CPU track, quantization
M3·02
1 · Why vLLM: from one cup at a time to a busy café
M3·03
Ollama is great, until the crowd arrives
One espresso shot at a time
M3·04
Continuous batching, never let a head sit idle
A finished slot refills mid-flight
M3·05
PagedAttention, virtual memory for the KV cache
Paging the KV cache like your OS pages RAM
M3·06
The payoff: roughly three times the throughput
Same hardware, heads never idle
M3·07
2 · Same contract, bigger engine
M3·08
Same contract, bigger engine
The wall socket idea from M2, again
M3·09
3 · The CPU track (what you'll run)
M3·10
The CPU track, learn the machinery anywhere
No GPU passthrough — slow on purpose
M3·11
Why containers report 0 NUMA nodes
The container can’t see the floor plan
M3·12
CPU tuning knobs that keep the laptop usable
A few dials, not all your cores
M3·13
4 · The GPU track (throughput — documented, not run here)
M3·14
CPU track vs GPU track
Learn here, scale on an NVIDIA box
M3·15
5 · Quantization in practice
M3·16
Quantization, trade a little precision for a lot of room
Like a JPEG, but for model weights
M3·17
6 · Operational gotchas
M3·18
GPU operational gotchas
Two flags, one sizing rule
M3·19
TO THE LAB
Same socket. Bigger engine.
Serve SmolLM2 on CPU vLLM, same /v1 API
Next: Lab — Serve SmolLM2 on CPU vLLM · Gourav Shah · School of DevOps & AI
M3·20