1:1 mentoring with Big Tech AI engineers
LLM & Agentic

Streaming with Claude

Stream Claude API responses for real-time UX: server-sent events, token-by-token rendering, and production streaming patterns.

Last updated

Core9 min readFirst readMessages APITool Use with Claude

After this section you can

  • Trace the SSE event order and say which event carries text, tool arguments, stop_reason and usage
  • Stream a tool-using turn safely, including eager input streaming and validation
  • Recover a broken stream on current models and decide when not to stream
23

Streaming: Real-Time Token Delivery

A long answer can take many seconds to generate. Streaming sends it as server-sent events while it is written, so users see text at once and long requests stay clear of HTTP timeouts.

Key idea

Streaming changes when you see the output, not how long it takes. The reply arrives as a fixed sequence of events; text and tool arguments exist only in the deltas, and stop_reason arrives near the end, in message_delta.

A streamed reply is a fixed sequence of events. Text lives only in the deltas.
ONE STREAMED RESPONSE, IN ORDER message_start Message, content empty usage.input_tokens REPEATS PER CONTENT BLOCK · index 0, 1, 2 … content_block_start announces the type no content yet content_block_delta × N, one fragment each the only place text is content_block_stop block complete parse tool JSON now message_delta stop_reason · final usage message_stop the stream ends THE FOUR DELTA TYPES, INSIDE content_block_delta text_delta text of a text block input_json_delta tool arguments as partial JSON strings thinking_delta thinking text; empty under the default display signature_delta sent just before a thinking block closes ping events can arrive anywhere, and an error event can end a stream that already returned HTTP 200.

Related

More in LLM & Agentic

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium