MODULE 2 · DAY 1
Serving Local Models
Open engines behind one universal endpoint
Gourav Shah · School of DevOps & AI · Hands-on
M2·01
What you'll learn
Five ideas that make model serving portable and boring.
M2·02
1 · Demo First: Docker Model Runner
M2·03
The quick demo: Docker Model Runner
One command, but locked to Docker's toolchain.
M2·04
2 · The Open Engines: Ollama, llama.cpp, LocalAI
M2·05
Open engines: different machines, same cup
Different machines, same standard cup.
M2·06
Which engine, when
Match the engine to the job.
M2·07
3 · The OpenAI-Compatible Endpoint: The Universal Contract
M2·08
The problem: every engine speaks differently?
Three engines, three ways to break.
M2·09
The universal contract: the /v1 endpoint
Swap what's behind the socket, not the plug.
M2·10
Swap engines by changing one variable
One variable, not a code change.
M2·11
4 · GGUF and Model Selection for Laptops
M2·12
GGUF: the JPEG of model weights
RAW is huge, JPEG loads instantly.
M2·13
Picking a model for a 16 GB laptop
Small models for the labs, big ones for later.
M2·14
5 · Two Wiring Patterns, One App
M2·15
Two wiring patterns, one app
Same app code, one variable changes.
M2·16
BIG IDEA
The engine is a deployment choice, not a code choice
Next: containerize a client that speaks this contract.
Continue to the Serving lab. · Gourav Shah · School of DevOps & AI
M2·17