Short answer
Tool calling lets an AI model ask your software to do something. You describe each tool, the model replies with the tool it wants and the input, and your code runs it and sends back the result. The model never runs anything itself. With tools you define, your code decides what actually happens.
Before you start
- Node 24 and npm, or Python 3.13
- An Anthropic API key, only for the live run (the tests need none)
What you'll build: A function that answers questions about jobs by calling a job-lookup tool.
Time: about 20 minutes
Tested with: node 24, @anthropic-ai/sdk 0.96.0, python 3.13, anthropic 1.9.0
Step 1: describe the tool
A is an action your code offers the model: look up a customer, check stock, draft an invoice. is the exchange that lets the model use one. The model asks for a tool by name and says what to give it. Your code runs the tool and sends back the result. The model carries on from there. It never runs anything itself. With tools you define, as in this guide, everything that happens is something your code chose to do.
This guide builds the smallest useful version of that: a function that answers questions about jobs by looking them up in a job book. Think of the call from a customer who wants to know when the tilers are coming. The question arrives in plain English, the model works out that it needs a lookup, and your code does the lookup. Where a step shows code, the TypeScript comes first, then the same code in Python. Both are tested with no API key.
Start by installing the SDK and a test runner. You do not need an API key until the step that runs it for real.
npm install @anthropic-ai/sdk
npm install --save-dev vitestpip install anthropic
pip install pytestThen lay the files out like this. The TypeScript files keep the folder names that the commands later in this guide use. The Python files all sit in one folder, so each can import the others by name.
examples/learn/
shared/
model.ts
client.ts
fakeClient.ts
tool-calling/
tool-calling.ts
tool-calling.test.ts
tool-calling.live.test.tstool-calling-py/
model.py
fake_client.py
tool_calling.py
test_tool_calling.py
test_live_tool_calling.pyThe shared files come first. The examples take the model name from one constant, so moving to a newer model is a one-line edit in one place, and no model name appears anywhere else. The Python version has its own copy of the constant.
/**
* The one model every Learn example uses. To move to a newer model, change this
* line (and its Python twin, model.py).
*/
export const MODEL = 'claude-opus-5-5';"""The one model every Learn example uses. See model.ts for the TypeScript twin."""
MODEL = "claude-opus-5-5"The TypeScript examples also take the client as a parameter. This type describes the only part of the client they call. A real Anthropic client fits it, and so does the fake one used for testing later on. Python needs no such type: any object with a messages.create method will do. The fake client is the third TypeScript file and the second Python file, and it arrives in the testing step.
import type Anthropic from '@anthropic-ai/sdk';
/**
* The slice of the Anthropic client the examples call. A real
* `new Anthropic()` satisfies it, and the tests pass a fake, so every example
* runs in CI with no API key and no cost.
*/
export type MessagesClient = {
messages: {
create(params: Anthropic.MessageCreateParamsNonStreaming): Promise<Anthropic.Message>;
};
};Now the tool. The model never sees your code. It sees three things about a tool: a name, a description and a JSON Schema for the input. That makes the description the part that does the work. Anthropic's own documentation calls it the most important factor in how well a tool gets used, and asks for detail on what the tool does and when to use it (define tools). The description below is three short sentences: it looks up a job by its number, it returns the customer, suburb, status and booked date, and it should be used whenever a question names a job number.
The schema describes the input. Its required list names the inputs the model must supply, which here is only the job number. Each property gets its own short description with an example of a good value, because the model reads those too. Keep the required list short and make everything else optional.
This block starts the main example file, tool-calling.ts in TypeScript and tool_calling.py in Python. Every later block from the same file goes below it, in the order the page shows them.
import type Anthropic from '@anthropic-ai/sdk';
import type { MessagesClient } from '../shared/client';
import { MODEL } from '../shared/model';
/** What the model sees: a name, a plain description and a JSON Schema for the input. */
export const lookupJobTool: Anthropic.Tool = {
name: 'lookup_job',
description:
'Look up one job in the job book by its job number. Returns the customer, suburb, status and booked date. Use it whenever a question names a job number.',
input_schema: {
type: 'object',
properties: {
job_number: { type: 'string', description: 'The job number, for example J-1042.' },
},
required: ['job_number'],
},
};import json
from model import MODEL
LOOKUP_JOB_TOOL = {
"name": "lookup_job",
"description": (
"Look up one job in the job book by its job number. Returns the customer, suburb, "
"status and booked date. Use it whenever a question names a job number."
),
"input_schema": {
"type": "object",
"properties": {"job_number": {"type": "string", "description": "The job number, for example J-1042."}},
"required": ["job_number"],
},
}The model decides when to call a tool from these words alone, so vague words get you vague behaviour. A clear name, a plain description and a stated trigger cost nothing to write and save a lot of guessing later.
Step 2: send the question with the tool
The first request is an ordinary one with one addition: the list of tools the model may ask for. The model then decides. It either answers in plain text or replies asking for a tool.
/**
* One round trip: ask, run whatever tools the model requests, send the results
* back, and return its answer. If it asks for tools a second time, that needs a
* loop (the next guide builds one), so this stops rather than guessing.
*/
export async function answerWithTools(client: MessagesClient, question: string): Promise<string> {
const messages: Anthropic.MessageParam[] = [{ role: 'user', content: question }];
const first = await client.messages.create({ model: MODEL, max_tokens: 16000, tools: [lookupJobTool], messages });
if (first.stop_reason !== 'tool_use') return textOf(first);
const calls = first.content.filter((b): b is Anthropic.ToolUseBlock => b.type === 'tool_use');
messages.push({ role: 'assistant', content: first.content });
messages.push({ role: 'user', content: calls.map((c) => runTool(c)) });
const second = await client.messages.create({ model: MODEL, max_tokens: 16000, tools: [lookupJobTool], messages });
if (second.stop_reason === 'tool_use') throw new Error('The model asked for more tools: this needs a loop.');
return textOf(second);
}
function textOf(reply: Anthropic.Message): string {
return reply.content.filter((b): b is Anthropic.TextBlock => b.type === 'text').map((b) => b.text).join('');
}def answer_with_tools(client, question: str) -> str:
"""One round trip: ask, run the requested tools, send the results back, return the answer."""
messages = [{"role": "user", "content": question}]
first = client.messages.create(model=MODEL, max_tokens=16000, tools=[LOOKUP_JOB_TOOL], messages=messages)
if first.stop_reason != "tool_use":
return _text_of(first)
calls = [block for block in first.content if block.type == "tool_use"]
messages.append({"role": "assistant", "content": first.content})
messages.append({"role": "user", "content": [run_tool(call) for call in calls]})
second = client.messages.create(model=MODEL, max_tokens=16000, tools=[LOOKUP_JOB_TOOL], messages=messages)
if second.stop_reason == "tool_use":
raise RuntimeError("The model asked for more tools: this needs a loop.")
return _text_of(second)
def _text_of(reply) -> str:
return "".join(block.text for block in reply.content if block.type == "text")The reply says which through its stop reason, a field on the response called stop_reason. A value of tool_use means the model wants a tool run. The content then holds one or more tool_use blocks, each with an id, the name of the tool and the input to pass it. The code checks this first. If the stop reason is anything else, the reply is the answer, and the function returns its text.
That last part is deliberately blunt. A reply can also stop because it ran out of room, or because the model declined the request, and this example treats both as a plain answer. That is fine for a first pass. Before you rely on it for real work, look at those cases and decide what your code should do about each.
The limit a reply can run into is max_tokens, which covers the model's thinking as well as its reply, so the code sets it with room to spare.
Tool choice stays automatic. That is the default when you pass tools: the model decides whether a tool is needed at all. Say hello and it answers directly. Ask about a job and it asks for the lookup. You might be tempted to force a call. With the model these examples use, forcing a tool call returns an error. If you need the tool used, say so in the prompt, and check that the reply really does contain a tool request before your code relies on one.
Step 3: run the tool and send the result back
The tool needs something to look in. In your business that is a job system or a spreadsheet. Here it is a small object with two jobs in it, so the tests run anywhere.
export type Job = { customer: string; suburb: string; status: 'booked' | 'in progress' | 'invoiced'; booked: string };
/** A stand-in for your job-management system. */
export const JOB_BOOK: Record<string, Job> = {
'J-1042': { customer: 'Sam Lee', suburb: 'Footscray', status: 'booked', booked: '2026-10-14' },
'J-1043': { customer: 'Priya Rao', suburb: 'Geelong', status: 'in progress', booked: '2026-10-02' },
};JOB_BOOK = {
"J-1042": {"customer": "Sam Lee", "suburb": "Footscray", "status": "booked", "booked": "2026-10-14"},
"J-1043": {"customer": "Priya Rao", "suburb": "Geelong", "status": "in progress", "booked": "2026-10-02"},
}Now the function that does the work. It takes one tool request, runs it and returns a tool result: a small block holding the id of the request it answers and the content. Here the content is a string: the job as JSON, which the model reads without help.
/** Runs one tool call and wraps the answer, or the problem, for the model. */
export function runTool(call: Anthropic.ToolUseBlock): Anthropic.ToolResultBlockParam {
if (call.name !== lookupJobTool.name) {
return { type: 'tool_result', tool_use_id: call.id, content: `Unknown tool: ${call.name}`, is_error: true };
}
const { job_number } = call.input as { job_number?: unknown };
if (typeof job_number !== 'string') {
return { type: 'tool_result', tool_use_id: call.id, content: 'Give a job_number, for example J-1042.', is_error: true };
}
// An own-property check: a plain object also answers to names like toString.
if (!Object.hasOwn(JOB_BOOK, job_number)) {
return { type: 'tool_result', tool_use_id: call.id, content: `No job ${job_number} in the job book.`, is_error: true };
}
return { type: 'tool_result', tool_use_id: call.id, content: JSON.stringify(JOB_BOOK[job_number]) };
}def run_tool(call) -> dict:
"""Runs one tool call and wraps the answer, or the problem, for the model."""
if call.name != LOOKUP_JOB_TOOL["name"]:
return {"type": "tool_result", "tool_use_id": call.id, "content": f"Unknown tool: {call.name}", "is_error": True}
job_number = call.input.get("job_number") if isinstance(call.input, dict) else None
if not isinstance(job_number, str):
return {"type": "tool_result", "tool_use_id": call.id, "content": "Give a job_number, for example J-1042.", "is_error": True}
job = JOB_BOOK.get(job_number)
if job is None:
return {"type": "tool_result", "tool_use_id": call.id, "content": f"No job {job_number} in the job book.", "is_error": True}
return {"type": "tool_result", "tool_use_id": call.id, "content": json.dumps(job)}Look again at the middle of the function from the previous step. The assistant's reply goes back into the conversation exactly as it arrived, tool requests and all. Do not trim it or rebuild it. The tool results then go in the next user message, and each one carries the id of the request it answers, which is how the model matches answers to questions. The API is strict about that shape (handle tool calls): the results must come straight after the assistant turn that asked for them, and any text you add to that user message goes after the results, not before.
One reply can ask for several tools. Ask the model to compare two jobs and it can request both lookups at once. The code answers every request in the reply, in the order they arrived, and sends all the results back in a single user message. Answering only the first is an easy way to get an error back from the API.
Step 4: tell the model when a tool fails
Tools fail. The job number is not in the book, the model leaves the input out, or it asks for a tool you never defined. Go back to the function in the step above. It gives up in three places: an unknown tool, a job number that is missing or not a string, and a job number that is not in the book. Each returns a tool result with the error flag, is_error, set to true and a message in plain words.
The flag tells the model the tool did not work. The message tells it why. From there the model can do something sensible: try again with a corrected input, or tell the person it could not find the job. Anthropic's documentation suggests writing messages that say what went wrong and what to try next (handle tool calls). Compare that with letting the function throw. The whole request falls over, and the person asking gets nothing.
Validate the input yourself. The schema tells the model what to send. It does not stop a bad call reaching your code, and the model can still leave out a required field or send the wrong type. So the function checks that the job number is a string, then looks it up in a way that only finds real jobs. In TypeScript that means Object.hasOwn, because a plain object also answers to names like toString, which would look like a job that exists. Python's dictionary lookup has no such trap. Anthropic also offers a strict setting on the tool definition that guarantees the input matches the schema, and it is worth reading up on. Even then, matching the schema is not the same as being a real job. That check is yours.
One more rule, for tools that change things. A lookup is harmless. A tool that sends an email, edits a booking or issues a refund is not. When you add one, your code is where permission gets checked, and where a person gets asked first if it matters. The model can request. Only your code decides.
Step 5: test it without an API key
None of this needs an API key to test. The tests use a fake client: a stand-in that plays back replies you script in advance, one for each request, and records what your code sent. The test then checks what your code did with those replies.
Here is the fake client, whole. Each scripted reply is typed as the SDK's own Message, and the Python version also validates every reply against the SDK's own model, so an SDK change that alters the shape fails there first.
import type Anthropic from '@anthropic-ai/sdk';
import type { MessagesClient } from './client';
import { MODEL } from './model';
type Reply = Anthropic.Message;
type Block = Anthropic.Message['content'][number];
type Params = Anthropic.MessageCreateParamsNonStreaming;
export type FakeClient = MessagesClient & {
/** Every request the code under test sent, in order, as it was at send time. */
readonly calls: Params[];
};
/** Plays back `script`, one reply per call, and fails loudly when the script runs out. */
export function createFakeClient(script: readonly Reply[]): FakeClient {
const calls: Params[] = [];
return {
calls,
messages: {
async create(params: Params): Promise<Reply> {
// A snapshot, so a caller that keeps appending to its messages array
// cannot rewrite what an earlier call sent.
calls.push(structuredClone(params));
const reply = script[calls.length - 1];
if (!reply) {
throw new Error(`fake client: no scripted reply for call ${calls.length} (script has ${script.length})`);
}
return reply;
},
},
};
}
let seq = 0;
// The SDK's Message type carries fields the examples never read, and new SDK
// versions add more. The fake fills what examples use and casts the rest.
function message(content: Block[], stopReason: Reply['stop_reason']): Reply {
seq += 1;
return {
id: `msg_fake_${seq}`,
type: 'message',
role: 'assistant',
model: MODEL,
content,
stop_reason: stopReason,
stop_sequence: null,
usage: { input_tokens: 0, output_tokens: 0 },
} as unknown as Reply;
}
export function textReply(text: string): Reply {
return message([{ type: 'text', text, citations: null } as unknown as Block], 'end_turn');
}
export function toolUseReply(name: string, input: Record<string, unknown>, id = `toolu_fake_${name}`): Reply {
return message([{ type: 'tool_use', id, name, input } as unknown as Block], 'tool_use');
}
/** Several tool calls in one reply, the way the model asks for parallel work. */
export function multiToolUseReply(calls: { name: string; input: Record<string, unknown>; id: string }[]): Reply {
if (calls.length === 0) {
throw new Error('multiToolUseReply: give at least one call');
}
return message(
calls.map(({ name, input, id }) => ({ type: 'tool_use', id, name, input }) as unknown as Block),
'tool_use',
);
}
/** A copy of `reply` that reports this token usage, for budget tests. */
export function withUsage(reply: Reply, inputTokens: number, outputTokens: number): Reply {
return {
...reply,
content: structuredClone(reply.content),
usage: { ...reply.usage, input_tokens: inputTokens, output_tokens: outputTokens },
} as Reply;
}
/** A copy of `reply` with a different stop reason, for max_tokens and refusal tests. */
export function withStopReason(reply: Reply, stopReason: Reply['stop_reason']): Reply {
return { ...reply, content: structuredClone(reply.content), stop_reason: stopReason };
}"""A stand-in for anthropic.Anthropic that plays back scripted replies.
The TypeScript twin is fakeClient.ts. Every Python example takes its client as a
parameter, so the tests pass this and a reader passes anthropic.Anthropic().
Replies are real anthropic.types.Message objects, validated by the SDK, so an
SDK upgrade that changes the shape fails here first.
"""
from __future__ import annotations
import copy
from typing import Any
from anthropic.types import Message, StopReason
from model import MODEL
_seq = 0
def _message(content: list[dict[str, Any]], stop_reason: str) -> Message:
global _seq
_seq += 1
return Message.model_validate(
{
"id": f"msg_fake_{_seq}",
"type": "message",
"role": "assistant",
"model": MODEL,
"content": content,
"stop_reason": stop_reason,
"stop_sequence": None,
"usage": {"input_tokens": 0, "output_tokens": 0},
}
)
def text_reply(text: str) -> Message:
return _message([{"type": "text", "text": text}], "end_turn")
def tool_use_reply(name: str, input: dict[str, Any], id: str | None = None) -> Message:
return _message([{"type": "tool_use", "id": id or f"toolu_fake_{name}", "name": name, "input": input}], "tool_use")
def multi_tool_use_reply(calls: list[dict[str, Any]]) -> Message:
"""Several tool calls in one reply, the way the model asks for parallel work."""
if not calls:
raise ValueError("multi_tool_use_reply: give at least one call")
return _message(
[{"type": "tool_use", "id": c["id"], "name": c["name"], "input": c["input"]} for c in calls],
"tool_use",
)
def with_usage(reply: Message, input_tokens: int, output_tokens: int) -> Message:
"""A copy of `reply` that reports this token usage, for budget tests."""
return Message.model_validate({
**reply.model_dump(),
"usage": {**reply.usage.model_dump(), "input_tokens": input_tokens, "output_tokens": output_tokens}
})
def with_stop_reason(reply: Message, stop_reason: StopReason | None) -> Message:
"""A copy of `reply` with a different stop reason, for max_tokens and refusal tests."""
return Message.model_validate({**reply.model_dump(), "stop_reason": stop_reason})
class _Messages:
def __init__(self, script: list[Message], calls: list[dict[str, Any]]):
self._script = script
self._calls = calls
def create(self, **params: Any) -> Message:
# A snapshot, so a caller that keeps appending to its messages list
# cannot rewrite what an earlier call sent.
self._calls.append(copy.deepcopy(params))
n = len(self._calls)
if n > len(self._script):
raise AssertionError(f"fake client: no scripted reply for call {n} (script has {len(self._script)})")
return self._script[n - 1]
class FakeClient:
def __init__(self, script: list[Message]):
self.calls: list[dict[str, Any]] = []
self.messages = _Messages(list(script), self.calls)Each test file starts with imports, and the test sits under them.
import { describe, it, expect } from 'vitest';
import type Anthropic from '@anthropic-ai/sdk';
import { createFakeClient, multiToolUseReply, textReply, toolUseReply } from '../shared/fakeClient';
import { answerWithTools, lookupJobTool, runTool } from './tool-calling';it('runs the tool the model asks for and sends the result back under the same id', async () => {
const client = createFakeClient([
toolUseReply('lookup_job', { job_number: 'J-1042' }, 'toolu_a'),
textReply('J-1042 is booked for 14 October in Footscray.'),
]);
await expect(answerWithTools(client, 'When is J-1042 booked?')).resolves.toBe('J-1042 is booked for 14 October in Footscray.');
expect(client.calls[0].tools).toEqual([lookupJobTool]);
expect(client.calls[1].tools).toEqual([lookupJobTool]);
expect(client.calls[1].messages[2]).toMatchObject({ role: 'user', content: [{ type: 'tool_result', tool_use_id: 'toolu_a' }] });
});import json
from types import SimpleNamespace
import pytest
from fake_client import FakeClient, multi_tool_use_reply, text_reply, tool_use_reply
from tool_calling import LOOKUP_JOB_TOOL, answer_with_tools, run_tooldef test_runs_the_tool_the_model_asks_for_and_sends_the_result_back_under_the_same_id():
client = FakeClient([
tool_use_reply("lookup_job", {"job_number": "J-1042"}, "toolu_a"),
text_reply("J-1042 is booked for 14 October in Footscray."),
])
assert answer_with_tools(client, "When is J-1042 booked?") == "J-1042 is booked for 14 October in Footscray."
assert client.calls[0]["tools"] == [LOOKUP_JOB_TOOL]
assert client.calls[1]["tools"] == [LOOKUP_JOB_TOOL]
sent_back = client.calls[1]["messages"][2]
assert sent_back["role"] == "user"
assert sent_back["content"][0]["type"] == "tool_result"
assert sent_back["content"][0]["tool_use_id"] == "toolu_a"Read the test from the top. The first scripted reply asks for the lookup, under an id the test makes up. The second is the final answer. The test checks three things: the function returns that final answer, the tool definition went out on both requests, and the second request carries a tool result under the same id the model used. The rest of the test file covers the other paths: a direct answer with no tool, two lookups in one reply, the error cases (including names like toString), and the stop when the model asks for tools a second time.
Now run them. The TypeScript command runs from the top of your project:
npx vitest run examples/learn/tool-callingThe Python command runs from the folder that holds the Python files. Run the one for your language, not both:
python -m pytestThese tests run on every change to this site's code, in both languages, with no key and no cost. Be clear about what they prove. They show your plumbing is right. They cannot show the model will choose well, because the model is not in the room. That is what the live run is for.
Step 6: run it for real
Set ANTHROPIC_API_KEY in your environment. The SDK reads it when you create the client with no arguments, and the live checks do exactly that. Here they are, whole:
import Anthropic from '@anthropic-ai/sdk';
import { describe, it, expect } from 'vitest';
import { answerWithTools } from './tool-calling';
// Skipped unless LEARN_LIVE=1, so an ordinary test run never makes a paid call.
// Run it yourself, with your own key, to watch the real model use the tool.
describe.skipIf(process.env.LEARN_LIVE !== '1')('tool calling: live', () => {
it('answers a job question through the real API', async () => {
const answer = await answerWithTools(new Anthropic(), 'When is job J-1042 booked, and in which suburb?');
console.log(answer);
expect(answer).toMatch(/Footscray/);
}, 120_000);
});import os
import pytest
from tool_calling import answer_with_tools
@pytest.mark.skipif(os.environ.get("LEARN_LIVE") != "1", reason="live API check; run by a person with LEARN_LIVE=1")
def test_live_answers_a_job_question():
import anthropic
answer = answer_with_tools(anthropic.Anthropic(), "When is job J-1042 booked, and in which suburb?")
print(answer)
assert "Footscray" in answerThen run one. Use the command for your language. The TypeScript command runs from the top of your project:
LEARN_LIVE=1 npx vitest run examples/learn/tool-calling/tool-calling.live.test.tsThe Python command runs from the folder that holds the Python files:
LEARN_LIVE=1 python -m pytest test_live_tool_calling.pyBoth live checks skip themselves unless the LEARN_LIVE variable is set to one, so an ordinary test run never calls the API. The commands above set it in front of the program, which works in bash and zsh. In Windows PowerShell, use the matching form below instead, and remove the variable afterwards so ordinary test runs stay off the API:
$env:LEARN_LIVE=1; npx vitest run examples/learn/tool-calling/tool-calling.live.test.ts
Remove-Item Env:LEARN_LIVE$env:LEARN_LIVE=1; python -m pytest test_live_tool_calling.py
Remove-Item Env:LEARN_LIVEThe check asks the real model when one job is booked and in which suburb, then looks for the suburb in the answer. The function makes at most two requests, so a run uses a small amount of usage.
This is also the run that earns the guide its Last checked date. Passing against the fake client is not enough. Someone has to run it against the real API, and only then does the date move.
When one round trip is not enough
The function makes at most two requests to the model, which is enough for a question with a single lookup. It is not enough when the first answer raises another question: find the job, then find who is booked on it, then check whether they are free. The model asks for tools a second time, and the function throws an error rather than guess. That is the check near the end of the function in the send-the-question step, and it has its own test.
Handling that needs a loop: call the model, run whatever tools it asks for, send the results back, and go round again until it gives an answer or you decide it has had enough. A separate guide, How to build an agentic loop from scratch, builds one from the pieces in this guide.
The Anthropic SDKs also ship a beta tool runner that runs the loop for you (tool runner). In TypeScript it is client.beta.messages.toolRunner, and you can define tools for example with betaZodTool. In Python it is client.beta.messages.tool_runner, and you can define tools for example with the beta_tool decorator. It is worth knowing about, and it makes more sense once you have seen what it does underneath. Both SDKs list it as beta, so check the current documentation before you depend on it.
Whichever way you go, the rule from the first step stays put. The model asks and your code decides. Start with tools that only read, test every path with the fake client, and add the ones that change things last.
Want this built for you instead? See how we build AI workflows, or book a free 30-minute consultation.

Peter McLean
Founder, Neurastruct
Australian small-business operator since 2001 and 16 years as a national account manager; AI certificates from Anthropic (2026) and Google (2025).
© Neurastruct Pty Ltd. Text licensed CC BY 4.0. Code samples licensed MIT. CC BY 4.0 · MIT