How LLMs Are Built
Complete lifecycle of large language models from pre-training through fine-tuning, RLHF, and deployment — with architecture diagrams and production considerations.
Last updated
After this section you can
- Trace a model from training data to a served endpoint, and say which costs are one-off and which recur per request
- Explain what pretraining learns, what post-training adds, and why a base model answers differently from a chat model
- Predict time to first token and total latency from input and output length, using prefill, decode and the KV cache
How LLMs Are Built: Training to Serving
Where the next-token probabilities come from, why the model knows some things and not others, and what happens on the server each time you call it.
Training happens once and produces frozen weights: pretraining gives the model its knowledge, post-training gives it its behaviour. Serving happens on every request, and it is where your latency and your bill come from.