Short answer
An agentic loop calls the model, runs any tools it asks for, sends the results back and repeats until the model gives a final answer or a step limit stops it. It fits on one screen of TypeScript or Python. Test it with a fake client first, so it costs nothing until you run it for real.
Before you start
- Node 24 and npm, or Python 3.13
- The tool calling guide, or the same idea from elsewhere
- An Anthropic API key, only for the live run
What you'll build: A reusable loop that looks jobs up and drafts team updates until the task is done.
Time: about 30 minutes
Tested with: node 24, @anthropic-ai/sdk 0.96.0, python 3.13, anthropic 1.9.0
Step 1: set up the project
An AI agent, stripped back, is a loop. You give the model a job and some tools. It replies asking for a tool. Your code runs the tool and sends back what happened. The model reads that and either asks for another tool or gives its answer. This guide builds that loop from scratch, in TypeScript and Python, and points it at a small job: a tile delivery is running late and someone has to tell the team.
It builds on the tool calling guide. That guide makes one round trip: ask, run the tool, send the result back, get the answer. One trip is enough while a question needs a single lookup. This loop lets the model ask again, and again, until it is done. It reuses that guide's fake client and its way of handling errors, so read it first if tool calling is new to you.
Where a step shows code, the TypeScript comes first, then the same code in Python. Both are tested with no API key. Use the language you work in: the install commands and folder layouts below are alternatives, so follow one set, not both.
Start by installing the SDK and a test runner. You do not need an API key until the step that runs it for real.
npm install @anthropic-ai/sdk
npm install --save-dev vitestpip install anthropic
pip install pytestThen lay the files out like this. The TypeScript files keep the folder names that the commands later in this guide use. The Python files all sit in one folder, so each can import the others by name.
examples/learn/
shared/
model.ts
client.ts
fakeClient.ts
build-an-agentic-loop/
agent-loop.ts
agent-loop.test.ts
demo.ts
agent-loop.live.test.tsbuild-an-agentic-loop-py/
model.py
fake_client.py
build_an_agentic_loop.py
test_build_an_agentic_loop.py
test_live_build_an_agentic_loop.pyThe shared files are the ones from the tool calling guide. If you built that guide's examples, model.ts, client.ts and fakeClient.ts are already in place, and you only need the new files. Python readers should copy model.py and fake_client.py from that guide's folder into this one. If you did not build that guide, the fake client is shown whole, in both languages, in the testing step of that guide. The model and client files are shown below: model.ts and client.ts in TypeScript, and model.py in Python.
Every example takes the model name from one constant, so moving to a newer model is a one-line edit in one place, and no model name appears anywhere else. Python has its own copy of the constant.
/**
* The one model every Learn example uses. To move to a newer model, change this
* line (and its Python twin, model.py).
*/
export const MODEL = 'claude-opus-5-5';"""The one model every Learn example uses. See model.ts for the TypeScript twin."""
MODEL = "claude-opus-5-5"The loop also takes its client as a parameter instead of creating one. This type describes the only part of the client the loop calls. A real Anthropic client fits it, and so does the fake one used for testing. That is what lets the whole loop run in a test with no key and no cost. Python needs no such type: any object with a messages.create method will do.
import type Anthropic from '@anthropic-ai/sdk';
/**
* The slice of the Anthropic client the examples call. A real
* `new Anthropic()` satisfies it, and the tests pass a fake, so every example
* runs in CI with no API key and no cost.
*/
export type MessagesClient = {
messages: {
create(params: Anthropic.MessageCreateParamsNonStreaming): Promise<Anthropic.Message>;
};
};Step 2: describe the tools and the result
The loop needs three things from you: the tools the model may ask for, the code that runs each one, and a limit on how many times it may go round. The first block of the loop file holds the types for them, and the type of what comes back. This block starts the main file, agent-loop.ts in TypeScript and build_an_agentic_loop.py in Python. Every later block from the same file goes below it, in the order the page shows them.
import type Anthropic from '@anthropic-ai/sdk';
import type { MessagesClient } from '../shared/client';
import { MODEL } from '../shared/model';
/** Your code for one tool: takes the model's input, returns text for the model. */
export type ToolHandler = (input: Record<string, unknown>) => Promise<string> | string;
export type AgentOptions = {
tools: Anthropic.Tool[];
handlers: Record<string, ToolHandler>;
/** The most model calls one task may make. */
maxSteps: number;
system?: string;
};
export type AgentResult = { text: string; steps: number; stopped: 'done' | 'step_limit' | 'refused' };from dataclasses import dataclass
from typing import Any, Callable
from model import MODEL
# Your code for one tool: takes the model's input, returns text for the model.
ToolHandler = Callable[[dict[str, Any]], str]
@dataclass
class AgentResult:
text: str
steps: int
stopped: str # "done", "step_limit" or "refused"A handler is your code for one tool. It takes the input the model chose and returns text for the model to read. The handlers object maps each tool's name to its handler, and that is how the loop finds what to run. A tool with no handler is a tool the loop cannot run, whatever the model asks for. The model asks. Your code decides.
The options also carry the tool descriptions the model sees, the most model calls one task may make, and an optional system prompt. That limit is the loop's only safety net, and it is required, not optional. A loop with no limit is a loop that can run away.
The result says how the run ended, and there are three ways. The model gave a final answer, and the run is done. The loop reached its limit while the model was still asking for tools, and it stopped at the step limit. Or the model declined, and the run was refused. The result also carries the final text, which is empty when the run never finished, and the number of model calls it took. Returning the reason matters, because the code that called the loop should treat a finished job, a step limit and a refusal differently.
Step 3: write the loop
This is the . The function starts a conversation with the task as the first user message, then goes round at most as many times as the limit allows. Each time round, it sends the whole conversation and the tools to the model and looks at what came back.
export async function runAgent(client: MessagesClient, task: string, options: AgentOptions): Promise<AgentResult> {
const messages: Anthropic.MessageParam[] = [{ role: 'user', content: task }];
for (let step = 1; step <= options.maxSteps; step++) {
const reply = await client.messages.create({
model: MODEL,
max_tokens: 16000,
...(options.system ? { system: options.system } : {}),
tools: options.tools,
messages,
});
messages.push({ role: 'assistant', content: reply.content });
// A tool call cut off at max_tokens can look complete. Never run it.
if (reply.stop_reason === 'max_tokens') throw new Error('The reply hit max_tokens; raise max_tokens and retry.');
if (reply.stop_reason === 'refusal') return { text: textOf(reply), steps: step, stopped: 'refused' };
if (reply.stop_reason !== 'tool_use') return { text: textOf(reply), steps: step, stopped: 'done' };
const results: Anthropic.ToolResultBlockParam[] = [];
for (const block of reply.content) {
if (block.type === 'tool_use') results.push(await runTool(block, options.handlers));
}
messages.push({ role: 'user', content: results });
}
return { text: '', steps: options.maxSteps, stopped: 'step_limit' };
}def run_agent(client, task: str, *, tools: list[dict], handlers: dict[str, ToolHandler], max_steps: int, system: str | None = None) -> AgentResult:
messages: list[dict[str, Any]] = [{"role": "user", "content": task}]
extra = {"system": system} if system else {}
for step in range(1, max_steps + 1):
reply = client.messages.create(model=MODEL, max_tokens=16000, tools=tools, messages=messages, **extra)
messages.append({"role": "assistant", "content": reply.content})
# A tool call cut off at max_tokens can look complete. Never run it.
if reply.stop_reason == "max_tokens":
raise RuntimeError("The reply hit max_tokens; raise max_tokens and retry.")
if reply.stop_reason == "refusal":
return AgentResult(_text_of(reply), step, "refused")
if reply.stop_reason != "tool_use":
return AgentResult(_text_of(reply), step, "done")
results = [_run_tool(block, handlers) for block in reply.content if block.type == "tool_use"]
messages.append({"role": "user", "content": results})
return AgentResult("", max_steps, "step_limit")The model's reply goes straight back into the conversation, exactly as it arrived, tool requests and all. Do not trim it or rebuild it. The API is strict about the order (handle tool calls): the results must come straight after the assistant turn that asked for them, and each result carries the id of the request it answers.
Then the loop checks the stop reason, in this order:
- max_tokens. The reply ran out of room. A tool request cut off part-way can still look complete, so the loop never runs it. It raises an error instead, and the fix is a larger max_tokens.
- refusal. The model declined. That is a reply, not a failure of your code, and it is not a finished job either, so the loop returns it as its own outcome, refused, with whatever text came back.
- Anything other than tool_use. The model has given its answer. The loop returns the text and reports done.
- tool_use. The model wants tools run. The loop runs them and goes round again.
One reply can ask for several tools. The loop answers every request in the reply, in the order they arrived, and sends all the results back in a single user message. Answering only the first is an easy way to get an error back from the API.
If the loop goes round as many times as the limit allows and the model is still asking for tools, it returns an empty text and the step-limit outcome. It does not throw, because reaching the limit is a result your code should plan for, not a crash.
Step 4: run each tool safely
The loop hands each tool request to one small function. It takes the request and returns a tool result: a block holding the id of the request it answers and the content. Tools fail in two ways the loop can see. The model asks for a tool you never gave a handler, or your handler throws. Neither lets the loop fall over. Each becomes a tool result with the error flag, is_error, set to true and a message in plain words.
async function runTool(call: Anthropic.ToolUseBlock, handlers: Record<string, ToolHandler>): Promise<Anthropic.ToolResultBlockParam> {
// An own-property check: a plain object also answers to names like toString.
if (!Object.hasOwn(handlers, call.name)) {
return { type: 'tool_result', tool_use_id: call.id, content: `Unknown tool: ${call.name}`, is_error: true };
}
try {
return { type: 'tool_result', tool_use_id: call.id, content: await handlers[call.name](call.input as Record<string, unknown>) };
} catch (error) {
// Tell the model what went wrong, so it can try another way or explain.
const message = error instanceof Error ? error.message : String(error);
return { type: 'tool_result', tool_use_id: call.id, content: `Tool failed: ${message}`, is_error: true };
}
}
function textOf(reply: Anthropic.Message): string {
return reply.content.filter((b): b is Anthropic.TextBlock => b.type === 'text').map((b) => b.text).join('');
}def _run_tool(call, handlers: dict[str, ToolHandler]) -> dict[str, Any]:
handler = handlers.get(call.name)
if handler is None:
return {"type": "tool_result", "tool_use_id": call.id, "content": f"Unknown tool: {call.name}", "is_error": True}
try:
return {"type": "tool_result", "tool_use_id": call.id, "content": handler(dict(call.input))}
except Exception as error: # Tell the model what went wrong, so it can try another way or explain.
return {"type": "tool_result", "tool_use_id": call.id, "content": f"Tool failed: {error}", "is_error": True}
def _text_of(reply) -> str:
return "".join(block.text for block in reply.content if block.type == "text")The flag tells the model the tool did not work, and the message tells it why. From there the model can do something sensible: try another way, or tell the person what went wrong. Anthropic's documentation suggests writing messages that say what went wrong and what to try next (handle tool calls). Compare that with letting the error escape. The whole run falls over, and the person who asked gets nothing. The model reads whatever you put in the message, so keep it to what someone on your team could act on.
In TypeScript the lookup checks Object.hasOwn before it reads the handlers object. A plain object also answers to names like toString. Without the check, a model that asked for a tool called toString would find something that is not your handler, and the loop would report success. Python's dictionary get has no such trap. The tool calling guide explains the same trap for a job lookup.
The small function at the bottom of the block joins the text blocks of a reply into one string, which is what the loop returns as the final text.
Step 5: give it a real job
The loop is general. It does nothing useful until you give it tools. The job desk has two: a lookup in the job book, and a drafter for team updates. The task is a tiling job running late. To handle it, the model has to look the job up, then draft an update, then stop. That is two tool calls and a final answer.
The TypeScript block is the whole of demo.ts, imports included. The Python block goes at the bottom of build_an_agentic_loop.py, below the run-tool block.
import type Anthropic from '@anthropic-ai/sdk';
import type { MessagesClient } from '../shared/client';
import { runAgent, type AgentResult, type ToolHandler } from './agent-loop';
/** What the model sees for the first tool: a name, a plain description and a JSON Schema for the input. */
export const lookupJobTool: Anthropic.Tool = {
name: 'lookup_job',
description:
'Look up one job in the job book by its job number. Returns the suburb and booked date. Use it whenever a task names a job number.',
input_schema: {
type: 'object',
properties: {
job_number: { type: 'string', description: 'The job number, for example J-1042.' },
},
required: ['job_number'],
},
};
export const draftUpdateTool: Anthropic.Tool = {
name: 'draft_update',
description:
'Write a short update message for the team. This only writes a draft: nothing is sent. Use it after you have looked the job up.',
input_schema: {
type: 'object',
properties: {
text: { type: 'string', description: 'The message, in plain words.' },
},
required: ['text'],
},
};
/** A stand-in for your job-management system. */
const JOBS: Record<string, { suburb: string; booked: string }> = {
'J-1042': { suburb: 'Footscray', booked: '2026-10-14' },
};
const handlers: Record<string, ToolHandler> = {
lookup_job: ({ job_number }) => {
if (typeof job_number !== 'string') return 'Give a job_number, for example J-1042.';
// An own-property check: a plain object also answers to names like toString.
if (!Object.hasOwn(JOBS, job_number)) return `No job ${job_number} in the job book.`;
return JSON.stringify(JOBS[job_number]);
},
// This writes a draft and nothing else. It never sends anything: a person reads the draft and sends it.
draft_update: ({ text }) => {
if (typeof text !== 'string') return 'Give the text of the update to draft.';
return `Drafted: ${text}`;
},
};
export function runJobDesk(client: MessagesClient, task: string): Promise<AgentResult> {
return runAgent(client, task, {
tools: [lookupJobTool, draftUpdateTool],
handlers,
maxSteps: 6,
system: 'You help a small trade business keep jobs on track. Look jobs up before you draft anything.',
});
}import json
# What the model sees for the first tool: a name, a plain description and a JSON Schema for the input.
LOOKUP_JOB_TOOL = {
"name": "lookup_job",
"description": (
"Look up one job in the job book by its job number. Returns the suburb and booked date. "
"Use it whenever a task names a job number."
),
"input_schema": {
"type": "object",
"properties": {"job_number": {"type": "string", "description": "The job number, for example J-1042."}},
"required": ["job_number"],
},
}
DRAFT_UPDATE_TOOL = {
"name": "draft_update",
"description": (
"Write a short update message for the team. This only writes a draft: nothing is sent. "
"Use it after you have looked the job up."
),
"input_schema": {
"type": "object",
"properties": {"text": {"type": "string", "description": "The message, in plain words."}},
"required": ["text"],
},
}
# A stand-in for your job-management system.
JOBS = {"J-1042": {"suburb": "Footscray", "booked": "2026-10-14"}}
def _lookup_job(input: dict[str, Any]) -> str:
job_number = input.get("job_number")
if not isinstance(job_number, str):
return "Give a job_number, for example J-1042."
job = JOBS.get(job_number)
if job is None:
return f"No job {job_number} in the job book."
return json.dumps(job)
def _draft_update(input: dict[str, Any]) -> str:
# This writes a draft and nothing else. It never sends anything: a person reads the draft and sends it.
text = input.get("text")
if not isinstance(text, str):
return "Give the text of the update to draft."
return f"Drafted: {text}"
def run_job_desk(client, task: str) -> AgentResult:
return run_agent(
client,
task,
tools=[LOOKUP_JOB_TOOL, DRAFT_UPDATE_TOOL],
handlers={"lookup_job": _lookup_job, "draft_update": _draft_update},
max_steps=6,
system="You help a small trade business keep jobs on track. Look jobs up before you draft anything.",
)The two tool definitions have the same shape as in the tool calling guide: a name, a plain description that says when to use it, and a schema for the input. The job book is a small object with one job in it, standing in for your job system. The lookup handler checks that the job number is a string and, in TypeScript, uses the same own-property check as the loop. When the job is missing, it returns a readable sentence that says what was missing, or what to send, and the model can act on it. It returns plain text, so the result goes back without the error flag. If you want the flag on a lookup miss, have the handler throw, and the loop turns that into an error result.
The system prompt is two short sentences, and it does real work: it tells the model to look jobs up before it drafts anything. The last function calls the loop with the two tools, their handlers, a limit of six steps and that prompt.
Step 6: test it without spending anything
Because the loop never builds its own client, a test can hand it a fake one. The fake plays back replies you script in advance, one for each request, and records what your code sent. That makes every path testable, including the awkward ones such as a refusal or a reply cut off at max_tokens, on your own machine or in CI. It is the fake client from the tool calling guide, so it is not shown again here. If you do not have it yet, it is shown whole, in both languages, in the testing step of that guide.
Each test file starts with its imports. A short setup follows, which the tests share: two tool definitions and a function that builds the options for a run. The test sits under them. In the TypeScript file the test lives inside a describe block, which is left out here, because a test at the top level of the file runs the same way.
import { describe, it, expect, vi } from 'vitest';
import type Anthropic from '@anthropic-ai/sdk';
import { createFakeClient, multiToolUseReply, textReply, toolUseReply, withStopReason } from '../shared/fakeClient';
import { runAgent, type AgentOptions } from './agent-loop';const lookup: Anthropic.Tool = {
name: 'lookup_job',
description: 'Look up a job by number.',
input_schema: { type: 'object', properties: { job_number: { type: 'string' } }, required: ['job_number'] },
};
const update: Anthropic.Tool = {
name: 'draft_update',
description: 'Draft a message to the team.',
input_schema: { type: 'object', properties: { text: { type: 'string' } }, required: ['text'] },
};
const options = (over: Partial<AgentOptions> = {}): AgentOptions => ({
tools: [lookup, update],
handlers: { lookup_job: () => '{"suburb":"Footscray"}', draft_update: ({ text }) => `Drafted: ${String(text)}` },
maxSteps: 5,
...over,
});it('keeps calling tools until the model gives a final answer', async () => {
const client = createFakeClient([
toolUseReply('lookup_job', { job_number: 'J-1042' }, 'toolu_1'),
toolUseReply('draft_update', { text: 'J-1042 moves to Thursday' }, 'toolu_2'),
textReply('Drafted the update for J-1042.'),
]);
const result = await runAgent(client, 'The tiles for J-1042 are late. Tell the team.', options());
expect(result).toEqual({ text: 'Drafted the update for J-1042.', steps: 3, stopped: 'done' });
expect(client.calls[2].messages.at(-1)).toMatchObject({
role: 'user',
content: [{ type: 'tool_result', tool_use_id: 'toolu_2', content: 'Drafted: J-1042 moves to Thursday' }],
});
// The model's own turn goes back as it arrived, so the result has a request to answer.
expect(client.calls[2].messages.at(-2)).toMatchObject({
role: 'assistant',
content: [{ type: 'tool_use', id: 'toolu_2', name: 'draft_update' }],
});
});import json
from unittest.mock import Mock
import pytest
from fake_client import FakeClient, multi_tool_use_reply, text_reply, tool_use_reply, with_stop_reason
from build_an_agentic_loop import AgentResult, run_agentLOOKUP = {
"name": "lookup_job",
"description": "Look up a job by number.",
"input_schema": {"type": "object", "properties": {"job_number": {"type": "string"}}, "required": ["job_number"]},
}
UPDATE = {
"name": "draft_update",
"description": "Draft a message to the team.",
"input_schema": {"type": "object", "properties": {"text": {"type": "string"}}, "required": ["text"]},
}
def options(**over):
return {
"tools": [LOOKUP, UPDATE],
"handlers": {
"lookup_job": lambda input: '{"suburb":"Footscray"}',
"draft_update": lambda input: f"Drafted: {input['text']}",
},
"max_steps": 5,
**over,
}def test_keeps_calling_tools_until_the_model_gives_a_final_answer():
client = FakeClient([
tool_use_reply("lookup_job", {"job_number": "J-1042"}, "toolu_1"),
tool_use_reply("draft_update", {"text": "J-1042 moves to Thursday"}, "toolu_2"),
text_reply("Drafted the update for J-1042."),
])
result = run_agent(client, "The tiles for J-1042 are late. Tell the team.", **options())
assert result == AgentResult("Drafted the update for J-1042.", 3, "done")
sent_back = client.calls[2]["messages"][-1]
assert sent_back["role"] == "user"
assert sent_back["content"] == [
{"type": "tool_result", "tool_use_id": "toolu_2", "content": "Drafted: J-1042 moves to Thursday"}
]
# The model's own turn goes back as it arrived, so the result has a request to answer.
assistant_turn = client.calls[2]["messages"][-2]
assert assistant_turn["role"] == "assistant"
assert assistant_turn["content"][0].type == "tool_use"
assert assistant_turn["content"][0].id == "toolu_2"Read the test from the top. The first scripted reply asks for the lookup, under an id the test makes up. The second asks for a draft. The third is the final answer. The test checks that the loop returns that answer, that it took three steps, and that it stopped because the model was done. It also checks the last two messages the loop sent: the model's own turn, and then the tool result for the draft, under the same id the model used.
The rest of each test file covers the other paths. A reply that starts with a thinking block, which must go back unchanged, ahead of its tool request. A run that hits the step limit. An unknown tool, and a tool that throws, including tool names like toString. Two tool requests in one reply, which must come back together, in one user message, in order. A reply cut off at max_tokens, which must never run its tool. A refusal, which must not be reported as done. And a system prompt that goes out only when one is given. Test the job desk the same way. Check that both tools are offered, that the system prompt is the one above, that a missing job comes back as a readable message, and that the limit of six really stops a run after six model calls.
Now run them. The TypeScript command runs from the top of your project:
npx vitest run examples/learn/build-an-agentic-loopThe Python command runs from the folder that holds the Python files. Run the one for your language, not both:
python -m pytestThese tests need no key and cost nothing, so run them on every change. Be clear about what they prove. They show your plumbing is right: the loop keeps its turns straight, answers every request and stops when it should. They cannot show that the model will choose those tools in that order, because the model is not in the room. That is what the live run is for.
Step 7: run it for real
Set ANTHROPIC_API_KEY in your environment. The SDK reads it when you create the client, and the live checks do exactly that. The Python one also gives the client a two-minute timeout for each request. Here they are, whole:
import Anthropic from '@anthropic-ai/sdk';
import { describe, it, expect } from 'vitest';
import { runJobDesk } from './demo';
// Skipped unless LEARN_LIVE=1, so an ordinary test run never makes a paid call.
// Run it yourself, with your own key, to watch the real model work through the job.
describe.skipIf(process.env.LEARN_LIVE !== '1')('agentic loop: live', () => {
it('finishes the job desk task through the real API', async () => {
const result = await runJobDesk(new Anthropic(), 'The tiles for job J-1042 will be two days late. Draft an update for the team.');
console.log(result.text);
expect(result.stopped).toBe('done');
expect(result.text.length).toBeGreaterThan(0);
}, 120_000);
});import os
import pytest
from build_an_agentic_loop import run_job_desk
@pytest.mark.skipif(os.environ.get("LEARN_LIVE") != "1", reason="live API check; run by a person with LEARN_LIVE=1")
def test_live_finishes_the_job_desk_task():
import anthropic
# Give up on any one request after 120 seconds.
client = anthropic.Anthropic(timeout=120.0)
result = run_job_desk(client, "The tiles for job J-1042 will be two days late. Draft an update for the team.")
print(result.text)
assert result.stopped == "done"
assert result.textThen run one. Use the command for your language. The TypeScript command runs from the top of your project:
LEARN_LIVE=1 npx vitest run examples/learn/build-an-agentic-loop/agent-loop.live.test.tsThe Python command runs from the folder that holds the Python files:
LEARN_LIVE=1 python -m pytest test_live_build_an_agentic_loop.pyBoth live checks skip themselves unless the LEARN_LIVE variable is set to one, so an ordinary test run never calls the API. The commands above set it in front of the program, which works in bash and zsh. In Windows PowerShell, use the matching form below instead, and remove the variable afterwards so ordinary test runs stay off the API:
$env:LEARN_LIVE=1; npx vitest run examples/learn/build-an-agentic-loop/agent-loop.live.test.ts
Remove-Item Env:LEARN_LIVE$env:LEARN_LIVE=1; python -m pytest test_live_build_an_agentic_loop.py
Remove-Item Env:LEARN_LIVEThe check gives the job desk the late-tiles task and expects the run to finish, with some text in the answer. It prints that text, so you can read what the model said at the end. One run spends a little usage. It passing tells you the model finished this time. It does not tell you the model always will, and a live check is one sample, not a guarantee.
This is also the run that earns the guide its Last checked date. Passing against the fake client is not enough. Someone has to run it against the real API, and only then does the date move.
Before you use it for real work
The step limit is the only budget this loop has. It stops a run that never finishes, but it does not know what each step costs, how long the run has taken, or whether the model keeps asking for the same thing. Every step also sends the whole conversation again, so each one carries more than the last. The stop conditions and budgets guide is the next one to read. It adds the rest: token budgets, time, repeated requests and approval.
There is one more gap. The loop reports any stop reason it does not name as done. That includes pause_turn, which a reply can end with when you add tools that Anthropic runs for you, such as web search. It also includes a stop at the model's context window limit, which the API reports as model_context_window_exceeded. Anthropic's documentation says to treat that response as truncated (stop reasons), and this loop would call it done. The step limit here, and the token budget in the next guide, make that less likely.
The Anthropic SDKs also ship a beta tool runner that runs the loop for you (tool runner). In TypeScript it is client.beta.messages.toolRunner, and you can define tools for example with betaZodTool. In Python it is client.beta.messages.tool_runner, and you can define tools for example with the beta_tool decorator. It is the shortcut once you understand what it does underneath, and you have just built that. Both SDKs list it as beta, so check the current documentation before you depend on it.
Whichever way you go, the rule stays put. The model asks and your code decides. Start with tools that only read, keep the limit low, test every path with the fake client, and add the tools that change things last, with a person approving them.
Want this built for you instead? See how we build AI workflows, or book a free 30-minute consultation.

Peter McLean
Founder, Neurastruct
Australian small-business operator since 2001 and 16 years as a national account manager; AI certificates from Anthropic (2026) and Google (2025).
© Neurastruct Pty Ltd. Text licensed CC BY 4.0. Code samples licensed MIT. CC BY 4.0 · MIT