Integration Guides

Anthropic Python SDK Tutorial: Setup, Streaming, Tools

Kenji Watanabe

The Anthropic Python SDK is the official anthropic package on PyPI. Install it with pip, create a client that reads ANTHROPIC_API_KEY from the environment, and call client.messages.create() with a model, a token limit and a message list. The same client also covers streaming, async, tool use, typed errors, retries, timeouts and a base_url option for gateways or relays.

Installing the Anthropic Python SDK

The package is published on PyPI and developed in the open on GitHub. Install it into a virtual environment so that upgrades stay isolated from the rest of your system:

python -m venv .venv
source .venv/bin/activate
pip install -U anthropic
python -c "import anthropic; print(anthropic.__version__)"

The SDK requires a reasonably current Python 3. If your project pins an older interpreter, check the version table in the repository README before upgrading, because major SDK releases have raised the minimum Python version in the past.

Creating the client and supplying the API key

The client constructor is anthropic.Anthropic(). Called with no arguments, it looks for the key in the ANTHROPIC_API_KEY environment variable. That is the recommended setup: the key never appears in source code, and the same code runs unchanged on a laptop, in CI and in production.

export ANTHROPIC_API_KEY="sk-ant-..."
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

You can also pass the key explicitly with anthropic.Anthropic(api_key=...), which is useful when the key comes from a secrets manager rather than the process environment. Avoid hard-coding it as a literal. If the variable is missing or malformed you will see an authentication error on the first request rather than at construction time; the article on fixing an Anthropic API key that is not working walks through the common causes.

Your first messages.create call

Every request goes through the Messages API. The three required parameters are model, max_tokens and messages. A system prompt is optional and sits outside the message list.

import anthropic

client = anthropic.Anthropic()
MODEL = "claude-sonnet-4-5"  # replace with the model id you intend to use

response = client.messages.create(
    model=MODEL,
    max_tokens=1024,
    system="You are a concise assistant for Python developers.",
    messages=[
        {"role": "user", "content": "Explain what a context manager is in two sentences."}
    ],
)

print(response.content[0].text)

The messages list alternates between user and assistant roles and must begin with a user turn. The API is stateless, so for a multi-turn conversation you append the assistant's reply and the next user message to the list and send the whole history again. For a broader walkthrough of the request shape, see how to use the Claude API.

Reading the response content blocks

response.content is a list of typed content blocks, not a single string. For a plain text request it usually contains one TextBlock, but when tools or extended thinking are involved it can contain several blocks of different types. Code that checks block.type before reading block.text keeps working as you add features:

for block in response.content:
    if block.type == "text":
        print(block.text)

print(response.stop_reason)         # "end_turn", "max_tokens", "tool_use", ...
print(response.usage.input_tokens, response.usage.output_tokens)

Two fields are worth logging from day one. stop_reason tells you why generation ended; a value of max_tokens means the answer was cut off and you should raise the limit. usage reports token counts, which is what billing and rate limits are based on.

Streaming with the stream helper

For anything longer than a short answer, stream the response so the user sees output immediately and the HTTP connection is not left idle. The SDK provides a context-manager helper, client.messages.stream(), that accumulates the message for you and exposes a text_stream iterator:

with client.messages.stream(
    model=MODEL,
    max_tokens=4096,
    messages=[{"role": "user", "content": "Write a short guide to Python dataclasses."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

    final = stream.get_final_message()

print()
print(final.usage.output_tokens)

get_final_message() returns the same Message object you would have received from a non-streaming call, so downstream code does not need two paths. If you need the raw event stream instead, pass stream=True to messages.create() and iterate the events yourself. The dedicated post on Claude API streaming covers event types and server-sent events in more depth.

The async client

Web frameworks and concurrent pipelines should use anthropic.AsyncAnthropic. It has the same methods as the synchronous client, awaited:

import asyncio
import anthropic

async def main() -> None:
    client = anthropic.AsyncAnthropic()

    response = await client.messages.create(
        model=MODEL,
        max_tokens=512,
        messages=[{"role": "user", "content": "Give me one tip for writing async Python."}],
    )
    print(response.content[0].text)

    async with client.messages.stream(
        model=MODEL,
        max_tokens=1024,
        messages=[{"role": "user", "content": "Now expand that tip into a paragraph."}],
    ) as stream:
        async for text in stream.text_stream:
            print(text, end="", flush=True)

asyncio.run(main())

Create one async client per process and reuse it; the underlying HTTP connection pool is what makes concurrent calls efficient. Mixing the sync client into an async event loop blocks the loop, so pick one style per code path.

Basic tool use

Tool use lets the model ask your program to run a function. You describe each tool with a name, a description and a JSON Schema for its input. When the model decides to call one, the response has stop_reason == "tool_use" and contains a ToolUseBlock with the arguments. You run the function, then send the result back in a tool_result block.

import json

tools = [
    {
        "name": "get_weather",
        "description": "Get the current weather for a city.",
        "input_schema": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    }
]

def get_weather(city: str) -> str:
    return json.dumps({"city": city, "condition": "clear", "temp_c": 21})

messages = [{"role": "user", "content": "What is the weather in Lisbon right now?"}]

response = client.messages.create(
    model=MODEL, max_tokens=1024, tools=tools, messages=messages
)

while response.stop_reason == "tool_use":
    messages.append({"role": "assistant", "content": response.content})
    results = []
    for block in response.content:
        if block.type == "tool_use":
            output = get_weather(**block.input)
            results.append(
                {"type": "tool_result", "tool_use_id": block.id, "content": output}
            )
    messages.append({"role": "user", "content": results})
    response = client.messages.create(
        model=MODEL, max_tokens=1024, tools=tools, messages=messages
    )

for block in response.content:
    if block.type == "text":
        print(block.text)

Three details matter. Append the assistant's full response.content to the history, not just the text, because the tool_use block must be present for the API to match the tool_result by tool_use_id. If the model requests several tools in one turn, return all the results in a single user message. If your function fails, return the error text with "is_error": True in the tool result instead of dropping it, so the model can recover.

Error classes, retries and timeouts

The SDK raises typed exceptions, all subclasses of anthropic.APIError. Catch the specific ones you can act on first and the general ones last:

import anthropic

try:
    response = client.messages.create(
        model=MODEL, max_tokens=256,
        messages=[{"role": "user", "content": "ping"}],
    )
except anthropic.AuthenticationError:
    print("API key missing or rejected")
except anthropic.RateLimitError as e:
    print("rate limited; retry-after:", e.response.headers.get("retry-after"))
except anthropic.APIStatusError as e:
    print("HTTP", e.status_code, e.message)
except anthropic.APIConnectionError:
    print("network problem reaching the endpoint")

Retries are built in. By default the client retries connection errors and HTTP 408, 409, 429 and 5xx responses twice with exponential backoff. The default request timeout is ten minutes. Both are configurable on the client and can be overridden per call:

client = anthropic.Anthropic(max_retries=3, timeout=60.0)

quick = client.with_options(max_retries=0, timeout=10.0).messages.create(
    model=MODEL, max_tokens=64,  # overrides apply to this request only
    messages=[{"role": "user", "content": "ok?"}],
)

Because timeouts are retried, the wall-clock worst case is roughly timeout multiplied by max_retries + 1. For latency-sensitive endpoints set both lower; for long generations switch to streaming instead of raising the timeout.

Pointing the SDK at a gateway or relay with base_url

By default the client sends requests to Anthropic's API host. The base_url parameter changes that host while keeping every other part of the request identical, which is how teams route traffic through an internal gateway, a logging proxy or a third-party relay that speaks the Anthropic protocol:

client = anthropic.Anthropic(
    base_url="https://your-endpoint.example.com",
    api_key="key-issued-by-that-endpoint",
)

The same setting is available as the ANTHROPIC_BASE_URL environment variable, so you can switch environments without touching code:

export ANTHROPIC_BASE_URL="https://your-endpoint.example.com"
export ANTHROPIC_API_KEY="key-issued-by-that-endpoint"

Two things to verify when you do this. The key must be the one issued by the endpoint you are pointing at, not your Anthropic key, unless the relay forwards it unchanged. And the endpoint must accept the /v1/messages path that the SDK appends, so pass the origin (scheme and host, optionally a path prefix) rather than a full URL ending in /messages. If you prefer the OpenAI client shape, some endpoints expose both protocols; that is a separate setup with the OpenAI client and is outside the scope of this tutorial.

ROIBest AI is an API relay endpoint compatible with the Anthropic and OpenAI protocols. If you use it, the base_url setting shown above is where it is configured, together with the key issued by the service; everything else in this tutorial stays the same. Details are at ai.roibest.com.

FAQ

Does the Anthropic Python SDK read the API key automatically?

Yes. anthropic.Anthropic() with no arguments reads ANTHROPIC_API_KEY from the environment. Pass api_key= explicitly only when the key comes from another source such as a secrets manager.

What is the difference between messages.create and messages.stream?

messages.create() returns the complete message once generation finishes. messages.stream() is a context manager that yields text as it is produced and still gives you the full message via get_final_message().

How do I handle a tool_use response?

Run the function named in the ToolUseBlock, then append the assistant message and a user message containing a tool_result block with the matching tool_use_id, and call the API again. Loop until stop_reason is no longer tool_use.

Can I use the SDK with a non-Anthropic endpoint?

Yes, as long as the endpoint implements the Anthropic Messages API. Set base_url on the client or export ANTHROPIC_BASE_URL, and use the key issued by that endpoint.