1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Messages API

Claude Messages API deep dive: request/response format, system prompts, multi-turn conversations, and best practices.

Last updated

After this section you can

  • Build a Messages API request and say which fields are required and which were removed on current models
  • Read a response by block type, act on every stop_reason and cost a call from its usage
  • Handle errors with the SDK's typed exceptions and built-in retries, and choose a model for a workload
20

Messages API: Your First Claude Call

Every Claude integration, from a chatbot to a long-running agent, is built from one HTTP call. Learn what goes into it, what comes back, how it fails, and which model to put in the model field.

Key idea

You send the whole conversation, Claude sends back one message. That message is a list of typed blocks plus a stop_reason telling your code what to do next, and a usage object telling you what it cost.

One request in, one response out. The server keeps nothing between calls.
Request · POST /v1/messages model required max_tokens required messages[] required system · tools output_config · thinking stream · cache_control Claude API one model call remembers nothing Response · one Message content[] text · thinking · tool_use · … stop_reason why it stopped = your next move usage tokens in, out, cached 1 · send 2 · reply 3 · append response.content to messages[], then send the whole array again Every turn is a fresh, complete request. Where that growing array should live is Section 03.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium