Building Agentic Systems
Short, opinionated posts on the primitives that make agentic systems reliable — hooks, loops, harness config, skills, subagents, MCP. Each with worked examples across engineering and non-engineering domains, because the same shapes apply wherever an LLM is in the loop.
The Last Mile Is Hiring — Become a Forward Deployed Engineer Now
The bottleneck moved from the model to the distance between a model that works and a company that can use it. Five diagrams on why that distance is widest right now, what agents took from the job and what they left, and the five questions that expose a bad version of the role.
LLM Evaluation Metrics That Actually Matter
Eval scores rising while users complain is rarely a broken metric — it is a metric that stopped describing production and kept returning a number anyway. Here are the three layers that answer three different questions, seven metrics and what each is blind to, why retrieval has to be scored separately from generation, how to validate an LLM judge like the classifier it is, and how to build the two hundred examples that decide whether any of it tells the truth.
AI System Design Interview — A Support Agent, Designed Out Loud
Everyone draws roughly the same diagram, which is why the diagram is not what separates candidates. Here is one problem worked end to end: the four scoping questions that actually change the design, the cost arithmetic that turns a triage box into an architectural argument, the five-stage architecture with the four points interviewers push on, what a strong answer sounds like at each, and the failure modes worth naming before you are asked.
How to Prevent Prompt Injection — A Defence Checklist
Jailbreaking is a user talking a model into saying something; injection is a third party getting your agent to do something — and once the agent has tools, that is a security problem, not a content one. Here is why the two are different, the five sources of untrusted text that all land in one window, the defence layers ranked by what they actually buy you, a checklist split into blast radius, input handling and detection, and how to red-team from the tool list rather than the prompt.
The Claude API in Practice — Tools, Thinking, Streaming, Caching
Tool use is not a separate API and neither is structured output — it is all one Messages call with different parameters attached. Here are the current model ids and prices, the four parameters that now return a 400 if you learned this API in 2025, the tool-result mistake that silently kills parallel calls, why streaming stops being optional above a certain max_tokens, and how to tell from usage alone whether your prompt cache is working.
RAG Chunking Strategies — Which One, and How to Tell Yours Is Wrong
A chunk is the smallest unit your system can retrieve, so a boundary in the wrong place puts the answer permanently out of reach — and it shows up as a generation problem. Here are the six strategies and when each is right, what the size and overlap numbers actually trade, why structure beats token counts, the metadata that turns out to be a security control, and the recall@k harness that tells you whether chunking is your problem at all.
Build an MCP Server in Python — From Fifteen Lines to Deployable
The Python SDK turns a decorated function into a working MCP server, and nobody gets stuck there. They get stuck on which operations deserve to be tools, what the descriptions say, and what a failed call leaves the agent holding. Here is the runnable minimum, the three primitives and who invokes each, why the docstring is the product, the five things between works and deployable, and the selection test that catches a reworded description breaking everything.
What Is a Forward Deployed Engineer? The Role, the Week, and the Trap
It is real engineering — but the week splits differently, and that is what makes it a different job. Here is where an FDE's hours actually go against a platform engineer's, the three roles people confuse it with and the one row that separates them, what the first six weeks of a deployment look like, the instinct that sinks engagements, and how to move in from platform or from consulting.
60 LLM & AI Agent Interview Questions (With Answers)
Sixty questions from real AI engineer and FDE loops, each with the two-or-three-sentence answer you would actually give before the interviewer decides whether to go deeper — grouped by the eight areas they move through, from fundamentals and tradeoffs to memory, tool design, cost, evaluation, security, and the thirteen curveballs that decide most offers.
LangChain vs LlamaIndex vs LangGraph — Which One, and When
LangChain is an integration layer, LlamaIndex is a data layer, LangGraph is a control layer — comparing them head-to-head is like asking whether requests, SQLAlchemy or Celery is the best Python library. Here is the stack drawn out, a three-question decision tree, a side-by-side on the nine dimensions that decide, three worked scenarios including one that needs no framework at all, and the production shape where two of them combine.
RAG vs Fine-Tuning — A Decision Framework (and When to Do Neither)
RAG and fine-tuning are not competing implementations of one idea — they fix different defects. Retrieval changes what is in the context window; fine-tuning changes what the model does with it. Here is the ten-second test that tells them apart, a side-by-side on the eleven dimensions that decide, four worked scenarios, the cost shape nobody models up front, and the four-arm eval that ends the argument in an afternoon.
Graph Engineering — Designing What the Agent Is Allowed to Do Next
Nodes do the work, edges decide what runs next, and one state object threads through both. Here's the mental model behind graph engineering, the four patterns that cover almost every production graph — router, fan-out/join, guarded cycle, human checkpoint — with runnable LangGraph examples, a worked refund agent, an honest list of when a graph is the wrong tool, and where the same shape shows up outside LangGraph.
The Forward Deployed Engineer Interview — What They Actually Ask
OpenAI, Anthropic and Google are all hiring FDEs, and a lot of strong engineers are now interviewing for a role whose loop they have never seen. Here is the quadrant the job actually occupies, the three stages, a minute-by-minute map of the decomposition case, and why speed to architecture is the thing that sinks people.
Does the AI Engineer Interview Still Have a LeetCode Round?
One round, about 45 minutes, and almost always a data-structure design problem rather than hard dynamic programming — because the job is the plumbing around a model. Here are eight free LeetCode problems that are quietly the same problems you solve in production, and the twelve minutes of follow-up that actually decide the round.
Context Engineering — Managing the Model's Working Memory
Prompt engineering tuned one message; context engineering designs the whole information flow around a finite window. It's the discipline behind agents that survive step 30 — and the sibling to harness and loop engineering. Here's the mental model, the failure mode (context rot), and the four moves that keep the window lean.
Loop Engineering — Stop Prompting, Start Designing the System
Direct prompting still works. But the practitioners closest to the frontier — Boris Cherny at Anthropic, Peter Steinberger, Addy Osmani — have all landed on the same shift: their job is to write loops that prompt the model, not to prompt the model themselves. Loop Engineering is the layer above harness engineering, and it changes what ‘shipping with an agent’ looks like.
Skills vs Subagents vs MCP — A Decision Tree
The three extension points look interchangeable at first glance. They aren't. Skills package knowledge, subagents package isolated reasoning, MCP servers package external access. Pick the wrong one and you'll end up rebuilding it — here's a decision tree and three worked scenarios.
Harness Engineering — Configuring the Shell Around the Model
The model does the work. The harness decides what it can touch, what runs before and after, and how permissions cascade from your home directory into a specific repo. Here's the layer stack, the failure modes of the wildcard permission, and a starter config that scales past one project.
Hooks in Claude Code — The Automation Layer Most People Skip
Hooks let Claude Code fire a shell command on tool events — before an edit, after an edit, when a run finishes. Used well, they replace half the ‘please format this’ / ‘don't touch .env’ nudges you type today. Here's the mental model, the JSON contract, and three worked examples.