1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Productionising LangGraph

Deploy a LangGraph agent: API worker vs queue consumer vs scheduled job on one Postgres, stream modes per consumer, bounding steps and tokens, and testing routers without a model.

Last updated

After this section you can

  • Choose a deployment shape for a graph and stream the modes each consumer needs
  • Bound steps, node time, run time and retries, and run the graph async behind a server
  • Test routers and gates without a model, and ship graph changes without stranding paused threads
40

Productionising LangGraph

The compiled graph is a library object. What makes it production is everything around it: where it runs, what the user sees while it runs, what stops it, how you test it, and how you change it while runs are still in flight.

Key idea

Nothing about a graph is a server. The same compiled object can sit behind an HTTP handler, a queue consumer and a cron job, and because the thread lives in the checkpointer, a run started by a request can be finished by a worker days later. Your work is the streams, the bounds, the tests and the migrations.

One compiled graph: three ways to run it yourself, or one managed service
YOU RUN IT · ONE COMPILED GRAPH, THREE SHAPES API worker FastAPI → graph.astream() runs that fit a request queue consumer one message = one run or one resume scheduled job sweeps unanswered gates retries, expiry Postgres checkpointer + store threads, checkpoints, pending interrupts OR MANAGED LangSmith Deployment runs Agent Server: HTTP API, task queue, persistence, Studio The thread lives in the database, not in a worker, so a run started by an API call can be resumed by a consumer the next day. Everything else in this section is what you add around the graph: streams, bounds, async, tests and safe changes.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium