1:1 mentoring with Big Tech AI engineers
LLM & Agentic

How LLMs Are Built

Complete lifecycle of large language models from pre-training through fine-tuning, RLHF, and deployment — with architecture diagrams and production considerations.

Last updated

Foundations8 min readFirst readTokens, Context Windows & Sampling

After this section you can

  • Trace a model from training data to a served endpoint, and say which costs are one-off and which recur per request
  • Explain what pretraining learns, what post-training adds, and why a base model answers differently from a chat model
  • Predict time to first token and total latency from input and output length, using prefill, decode and the KV cache
02

How LLMs Are Built: Training to Serving

Where the next-token probabilities come from, why the model knows some things and not others, and what happens on the server each time you call it.

Key idea

Training happens once and produces frozen weights: pretraining gives the model its knowledge, post-training gives it its behaviour. Serving happens on every request, and it is where your latency and your bill come from.

Four phases happen once, before you ever call the model. The fifth happens on every request.
1 · Data crawl, filter, dedup sets the cutoff date 2 · Pretrain predict the next token learns knowledge 3 · Post-train SFT, preferences, safety learns behaviour 4 · Evaluate benchmarks, red-teaming then weights freeze 5 · Serve prefill, then decode on every request ONE-OFF · THE LAB PAYS, YOU INHERIT THE WEIGHTS PER REQUEST you pay on every call Phases 1 to 3 decide what the model knows and how it behaves. Phase 5 decides what each call costs and how long it takes.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium