Short answer
A loop without limits can run up cost or keep acting after it should have stopped. Give every loop a step limit, a token budget and a time limit, stop it when it repeats the same tool call, and pause for a person before any risky action. Check the budget before each model call, not after.
Before you start
- The agentic loop guide, or a loop of your own
- Node 24 and npm, or Python 3.13
What you'll build: A loop that stops on steps, tokens, time or repetition, and asks a person before acting.
Time: about 30 minutes
Tested with: node 24, @anthropic-ai/sdk 0.96.0, python 3.13, anthropic 1.9.0
Step 1: decide the limits
An agent decides for itself how many times to go round its loop. That is the point of it, and it is also the risk. A loop with no limits can run up a bill, run for hours, or keep acting long after it should have stopped. This guide takes the loop from the agentic loop guide and puts around it: a cap on steps, a budget for spend, a time limit, a check for a model that keeps asking for the same thing, and a pause for a person before a risky action.
It builds on the tool calling guide and the agentic loop guide. The loop guide goes round as often as the model asks, up to a step limit, and that limit is the only limit it has. This guide re-states the loop with the other limits and a pause for approval added, so its code stands on its own. If the loop is new to you, read that guide first.
Where a section shows code, the TypeScript comes first, then the same code in Python. Both are tested with no API key. Use the language you work in: the install commands and folder layouts below are alternatives, so follow one set, not both.
Start by installing the SDK and a test runner. You do not need an API key until the section that runs it for real.
npm install @anthropic-ai/sdk
npm install --save-dev vitestpip install anthropic
pip install pytestThen lay the files out like this. The TypeScript files keep the folder names that the commands later in this guide use. The Python files all sit in one folder, so each can import the others by name.
examples/learn/
shared/
model.ts
client.ts
fakeClient.ts
stop-conditions-and-budgets/
guarded-loop.ts
guarded-loop.test.ts
guarded-loop.live.test.tsstop-conditions-and-budgets-py/
model.py
fake_client.py
stop_conditions_and_budgets.py
test_stop_conditions_and_budgets.py
test_live_stop_conditions_and_budgets.pyThe shared files are the ones from the tool calling guide. If you built that guide's examples, model.ts, client.ts and fakeClient.ts are already in place, and you only need the new files. Python readers should copy model.py and fake_client.py from that guide's folder into this one. If you did not build that guide, the fake client is shown whole, in both languages, in the testing section of that guide. The model and client files are shown below: model.ts and client.ts in TypeScript, and model.py in Python.
Every example takes the model name from one constant, so moving to a newer model is a one-line edit in one place, and no model name appears anywhere else. Python has its own copy of the constant.
/**
* The one model every Learn example uses. To move to a newer model, change this
* line (and its Python twin, model.py).
*/
export const MODEL = 'claude-opus-5-5';"""The one model every Learn example uses. See model.ts for the TypeScript twin."""
MODEL = "claude-opus-5-5"The loop takes its client as a parameter instead of creating one. This type describes the only part of the client it calls. A real Anthropic client fits it, and so does the fake one used for testing, so the whole loop runs in a test with no key and no cost. Python needs no such type: any object with a messages.create method will do.
import type Anthropic from '@anthropic-ai/sdk';
/**
* The slice of the Anthropic client the examples call. A real
* `new Anthropic()` satisfies it, and the tests pass a fake, so every example
* runs in CI with no API key and no cost.
*/
export type MessagesClient = {
messages: {
create(params: Anthropic.MessageCreateParamsNonStreaming): Promise<Anthropic.Message>;
};
};Now the limits. Each one protects against something different. There are four numbers and one pause.
- Steps. The most model calls one task may make. It stops a run that never finishes because the model keeps asking for another tool.
- Spend. A budget for the whole run, added up from what the API reports for each call. It stops a run that does finish, but costs more than the job is worth. The next section covers how it is counted.
- Time. The most elapsed time a whole task may take. It stops a job that drags on, because every call is slow or the tools are.
- Repeats. How many times in a row the model may ask for the exact same tool calls. It stops a stuck loop, where each call looks fine but the run is going nowhere.
- Approval. Not a number. A person is asked before a tool runs, and a no stops it. It protects you from an action you cannot take back.
The four numbers are the first block of the loop file, along with its imports. It starts the main file, guarded-loop.ts in TypeScript and stop_conditions_and_budgets.py in Python. Every later block from the same file goes below it, in the order the page shows them.
import type Anthropic from '@anthropic-ai/sdk';
import type { MessagesClient } from '../shared/client';
import { MODEL } from '../shared/model';
export type Limits = {
/** The most model calls one task may make. */
maxSteps: number;
/** Input plus output tokens across every call. Each call re-sends the conversation, so this grows quickly. */
maxTokens: number;
/** Elapsed time for the whole task, in milliseconds. */
maxMillis: number;
/** How many times in a row the model may ask for the exact same tool calls. */
maxRepeats: number;
};
export type StopReason = 'done' | 'refused' | 'step_limit' | 'token_budget' | 'time_limit' | 'repeating';
type Spend = { steps: number; tokens: number; startedAt: number };
/** Why to stop before the next model call, or null to carry on. */
export function overLimit(spend: Spend, limits: Limits, now: number): StopReason | null {
if (spend.steps >= limits.maxSteps) return 'step_limit';
if (spend.tokens >= limits.maxTokens) return 'token_budget';
if (now - spend.startedAt >= limits.maxMillis) return 'time_limit';
return null;
}import json
import time
from dataclasses import dataclass
from typing import Any, Callable
from model import MODEL
@dataclass
class Limits:
max_steps: int # The most model calls one task may make.
max_tokens: int # Input plus output tokens across every call. Each call re-sends the conversation, so this grows quickly.
max_seconds: float # Elapsed time for the whole task, in seconds. The loop measures it with time.monotonic.
max_repeats: int # How many times in a row the model may ask for the exact same tool calls.
# The reason a run stopped is one of these strings:
# "done", "refused", "step_limit", "token_budget", "time_limit" or "repeating".
def over_limit(spend: dict, limits: Limits, now: float) -> str | None:
"""Why to stop before the next model call, or None to carry on."""
if spend["steps"] >= limits.max_steps:
return "step_limit"
if spend["tokens"] >= limits.max_tokens:
return "token_budget"
if now - spend["started_at"] >= limits.max_seconds:
return "time_limit"
return NoneThe four numbers travel together in one object. The function overLimit takes what the run has spent so far, the limits, and the time now, and returns the name of the first limit reached, or nothing. It looks at steps first, then spend, then time. The time now is passed in, so a test can hand it a fake clock and never wait. Both loops read a clock that only moves forward, so a change to the computer's own clock cannot upset them. TypeScript uses performance.now() and counts milliseconds. Python uses time.monotonic and counts seconds.
Be clear about what the time limit does. The clock is checked before each model call, so it cannot cut short a call, or a tool, that is already running. A single request that never returns is bounded by the client's own timeout instead. The SDK sets a default, and you can set your own when you create the client: in milliseconds in TypeScript and in seconds in Python. The SDK also retries a request that times out, so the worst case is longer than one timeout. A tool that can hang needs a timeout of its own, inside its handler.
Step 2: check the budget before every call
This is the loop from the agentic loop guide with the limits added. The block opens with the options the loop takes and the result it returns, then the loop itself. The options are the tools, the handlers that run them, the limits, an optional approval function and an optional clock. The result holds the final text, why the run stopped, how many steps it took and how many it used.
export type ToolHandler = (input: Record<string, unknown>) => Promise<string> | string;
export type GuardedOptions = {
tools: Anthropic.Tool[];
handlers: Record<string, ToolHandler>;
limits: Limits;
/** Asked before a known tool runs. A name with no handler is reported without asking. */
approve?: (call: Anthropic.ToolUseBlock) => Promise<boolean> | boolean;
/** Milliseconds now, from a clock that only moves forward. Defaults to performance.now(). Tests pass a fake clock. */
now?: () => number;
};
export type GuardedResult = { text: string; stopped: StopReason; steps: number; tokens: number };
export async function runGuardedAgent(client: MessagesClient, task: string, options: GuardedOptions): Promise<GuardedResult> {
const now = options.now ?? (() => performance.now());
const spend: Spend = { steps: 0, tokens: 0, startedAt: now() };
const messages: Anthropic.MessageParam[] = [{ role: 'user', content: task }];
const stop = (stopped: StopReason, text = ''): GuardedResult => ({ text, stopped, steps: spend.steps, tokens: spend.tokens });
let previousCalls = '';
let repeats = 0;
while (true) {
const limit = overLimit(spend, options.limits, now());
if (limit) return stop(limit);
const reply = await client.messages.create({ model: MODEL, max_tokens: 16000, tools: options.tools, messages });
spend.steps += 1;
spend.tokens += reply.usage.input_tokens + reply.usage.output_tokens;
messages.push({ role: 'assistant', content: reply.content });
if (reply.stop_reason === 'max_tokens') throw new Error('The reply hit max_tokens; raise max_tokens and retry.');
if (reply.stop_reason === 'refusal') return stop('refused', textOf(reply));
if (reply.stop_reason !== 'tool_use') return stop('done', textOf(reply));
const calls = reply.content.filter((b): b is Anthropic.ToolUseBlock => b.type === 'tool_use');
const signature = JSON.stringify(calls.map((c) => [c.name, c.input]));
repeats = signature === previousCalls ? repeats + 1 : 0;
previousCalls = signature;
if (repeats >= options.limits.maxRepeats) return stop('repeating');
const results: Anthropic.ToolResultBlockParam[] = [];
for (const call of calls) results.push(await runApproved(call, options));
messages.push({ role: 'user', content: results });
}
}# Your code for one tool: takes the model's input, returns text for the model.
ToolHandler = Callable[[dict[str, Any]], str]
@dataclass
class GuardedResult:
text: str
stopped: str # One of the reasons listed above.
steps: int
tokens: int
def run_guarded_agent(
client,
task: str,
*,
tools: list[dict],
handlers: dict[str, ToolHandler],
limits: Limits,
approve: Callable[[Any], bool] | None = None, # Asked before a known tool runs. A name with no handler is reported without asking.
now: Callable[[], float] = time.monotonic, # Seconds now. Tests pass a fake clock.
) -> GuardedResult:
spend = {"steps": 0, "tokens": 0, "started_at": now()}
messages: list[dict[str, Any]] = [{"role": "user", "content": task}]
previous_calls = ""
repeats = 0
def stop(stopped: str, text: str = "") -> GuardedResult:
return GuardedResult(text, stopped, spend["steps"], spend["tokens"])
while True:
limit = over_limit(spend, limits, now())
if limit:
return stop(limit)
reply = client.messages.create(model=MODEL, max_tokens=16000, tools=tools, messages=messages)
spend["steps"] += 1
spend["tokens"] += reply.usage.input_tokens + reply.usage.output_tokens
messages.append({"role": "assistant", "content": reply.content})
if reply.stop_reason == "max_tokens":
raise RuntimeError("The reply hit max_tokens; raise max_tokens and retry.")
if reply.stop_reason == "refusal":
return stop("refused", _text_of(reply))
if reply.stop_reason != "tool_use":
return stop("done", _text_of(reply))
calls = [block for block in reply.content if block.type == "tool_use"]
signature = json.dumps([[call.name, call.input] for call in calls])
repeats = repeats + 1 if signature == previous_calls else 0
previous_calls = signature
if repeats >= limits.max_repeats:
return stop("repeating")
results = [_run_approved(call, handlers, approve) for call in calls]
messages.append({"role": "user", "content": results})Read the top of the loop first. Before every model call, it asks overLimit whether any limit has been reached. If one has, the function returns at once with that reason and makes no call. If none has, it makes the call, counts a step, and adds the call's input and output tokens to a running total. Both come straight from the reply: input_tokens is what the model read, and output_tokens is what it wrote.
The check comes before the call, not after it, for a plain reason: a call you have not made cannot cost anything. A test proves it. With no steps allowed, the fake client has nothing scripted, and no call goes out.
That has a consequence, and you should plan for it. The check runs before each call, so one call can still take the run over budget. The check then stops the next call. The token test in the testing section shows it: two scripted calls together end nine hundred tokens past the budget, and the run stops before the third. The budget is not a hard ceiling. Set it a call's worth below the most you can afford to spend.
Input is the part that grows. Each call sends the whole conversation again: the task, every earlier reply and every tool result, so a long run costs more than its step count suggests. It is also how a long run edges towards the limit of the , the most the model can take in at once.
Two cautions about the count. It adds only input_tokens and output_tokens. With prompt caching on, the API reports cached input in separate fields, and you would add those too. And the budget is not the max_tokens setting on the request, which caps the length of one reply.
After the call, the loop checks the API's own stop reason, in the same order as the loop guide. There are two things called stopping here. The stop reason says why the model stopped writing one reply. The stopped field in the result says why the whole run ended. A reply cut off at max_tokens raises an error, and its tool request never runs. A refusal ends the run as refused, not as done. Anything other than tool_use ends the run as done.
That last rule is blunter than it sounds. The loop reports any stop reason it does not name as done. That includes pause_turn, which a reply can end with when you add tools that Anthropic runs for you, such as web search. It also includes a stop at the model's context window limit, which the API reports as model_context_window_exceeded. Anthropic's documentation says to treat that response as truncated (stop reasons), and this loop would call it done. The token budget makes a long run less likely to get there. It does not rule it out: one very large tool result can fill the window in a single step.
If the model asked for tools, the loop checks for a repeat, which is the next section, then runs each request and sends all the results back in one user message. One reply can ask for several tools, and answering only the first is an easy way to get an error back from the API, so the loop answers every request, in order. A test checks that too.
Step 3: catch a loop that repeats itself
A loop can be stuck without having reached any limit yet. The model asks for the same lookup, gets the same answer and asks again. Each call looks fine, but the run is going nowhere and spending as it goes. The step limit would end it eventually. The repeat check ends it sooner.
Go back to the loop block, to the lines after the stop reason checks. The loop writes down the calls the model just asked for: each tool's name and its input, as one piece of text. That is the signature of the reply. If it matches the signature of the reply before, a counter goes up by one. If it differs, the counter goes back to zero. When the counter reaches the repeat limit, the loop stops and reports repeating.
Where the check sits matters. It runs before any tool in the reply runs. That is on purpose, because running a repeat is often the harm: a second email, a second booking, a second refund. The loop stops on the ask, not after the action.
Take a repeat limit of two. The model asks for the same call, and it runs. It asks again, and it runs a second time. It asks a third time, and the run stops before that third call runs. The test scripts three identical replies and checks that the run stops as repeating after three model calls, with the handler called twice.
Know what the check cannot see. It compares each reply with the one before it, and only for an exact match. A model that changes one word of the input each time, or that swings between two different calls, never trips it. The step limit and the budget are what end those runs. And some jobs repeat a call on purpose, such as checking a status until it changes. If yours does, raise the limit to fit.
Step 4: pause for a person before risky actions
Some actions should not run just because a model asked. Sending an email, changing a booking, refunding a customer: you want a person to say yes first. The last block of the loop file does that. Every tool request goes through one function, and the lookup comes first. If the name has no handler, the request goes straight back to the model as an error and nobody is asked. For a known tool, if an approval function was passed in, it is asked next, and it gets the whole tool request: the name and the input.
async function runApproved(call: Anthropic.ToolUseBlock, options: GuardedOptions): Promise<Anthropic.ToolResultBlockParam> {
// An own-property check: a plain object also answers to names like toString.
if (!Object.hasOwn(options.handlers, call.name)) {
return { type: 'tool_result', tool_use_id: call.id, content: `Unknown tool: ${call.name}`, is_error: true };
}
if (options.approve && !(await options.approve(call))) {
return {
type: 'tool_result',
tool_use_id: call.id,
content: 'A person declined this action. Do not try it again; say what you would have done instead.',
is_error: true,
};
}
try {
return { type: 'tool_result', tool_use_id: call.id, content: await options.handlers[call.name](call.input as Record<string, unknown>) };
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
return { type: 'tool_result', tool_use_id: call.id, content: `Tool failed: ${message}`, is_error: true };
}
}
function textOf(reply: Anthropic.Message): string {
return reply.content.filter((b): b is Anthropic.TextBlock => b.type === 'text').map((b) => b.text).join('');
}def _run_approved(call, handlers: dict[str, ToolHandler], approve: Callable[[Any], bool] | None) -> dict[str, Any]:
handler = handlers.get(call.name)
if handler is None:
return {"type": "tool_result", "tool_use_id": call.id, "content": f"Unknown tool: {call.name}", "is_error": True}
if approve is not None and not approve(call):
return {
"type": "tool_result",
"tool_use_id": call.id,
"content": "A person declined this action. Do not try it again; say what you would have done instead.",
"is_error": True,
}
try:
return {"type": "tool_result", "tool_use_id": call.id, "content": handler(dict(call.input))}
except Exception as error: # Tell the model what went wrong, so it can try another way or explain.
return {"type": "tool_result", "tool_use_id": call.id, "content": f"Tool failed: {error}", "is_error": True}
def _text_of(reply) -> str:
return "".join(block.text for block in reply.content if block.type == "text")If the answer is no, the handler never runs. The model gets a tool result with the error flag, is_error, set to true, and a message: a person declined this action, do not try it again, say what you would have done instead. The wording is deliberate. Without it, the model may retry the same call, or carry on as if it had worked. With it, the model can tell the person what it wanted to do and why. If the model ignores the message and asks again, the repeat check from the previous section is the backstop.
The approval function is yours. It can say yes straight away to tools that only read, and stop to ask a person only for the ones that change something. In TypeScript it can wait on a promise, so it can wait for a click. In Python it is an ordinary function that returns true or false. The clock keeps running while a person thinks, so a run that stops for approval needs a time limit that allows for that.
The lookup guard from the earlier guides is here too, and it runs before the approval, so a person is never asked to approve a tool that does not exist. In TypeScript the guard is Object.hasOwn, because a plain object also answers to names like toString, and a model that asked for a tool of that name would otherwise find something that is not your handler. A test covers it. Python's dictionary get has no such trap. The agentic loop guide explains that trap, and how an unknown tool and a tool that throws are reported to the model. This block does both the same way.
The small function at the bottom joins a reply's text blocks into the final text.
Step 5: test every way it can stop
Every way this loop can stop should have a test, and none of them needs an API key. The tests use the fake client from the tool calling guide, so it is not shown again here. If you do not have it yet, it is shown whole, in both languages, in the testing section of that guide.
Each test file starts with its imports, then a short setup the tests share: a tool definition, a set of limits and a function that builds the options for a run. In the TypeScript file the test lives inside a describe block, which is left out here, because a test at the top level of the file runs the same way.
import { describe, it, expect, vi } from 'vitest';
import type Anthropic from '@anthropic-ai/sdk';
import { createFakeClient, multiToolUseReply, textReply, toolUseReply, withStopReason, withUsage } from '../shared/fakeClient';
import { overLimit, runGuardedAgent, type GuardedOptions, type Limits } from './guarded-loop';const tool: Anthropic.Tool = {
name: 'lookup_job',
description: 'Look up a job by number.',
input_schema: { type: 'object', properties: { job_number: { type: 'string' } }, required: ['job_number'] },
};
const limits: Limits = { maxSteps: 10, maxTokens: 50_000, maxMillis: 120_000, maxRepeats: 2 };
const opts = (over: Partial<GuardedOptions> = {}): GuardedOptions => ({
tools: [tool],
handlers: { lookup_job: () => '{"suburb":"Footscray"}' },
limits,
now: () => 0,
...over,
});it('stops on the token budget before the next call, counting input and output', async () => {
const client = createFakeClient([
withUsage(toolUseReply('lookup_job', { job_number: 'A' }), 30_000, 500),
withUsage(toolUseReply('lookup_job', { job_number: 'B' }), 20_000, 400),
textReply('never reached'),
]);
const result = await runGuardedAgent(client, 'x', opts());
expect(result).toEqual({ text: '', stopped: 'token_budget', steps: 2, tokens: 50_900 });
expect(client.calls).toHaveLength(2);
});from dataclasses import replace
from unittest.mock import Mock
import pytest
from fake_client import FakeClient, multi_tool_use_reply, text_reply, tool_use_reply, with_stop_reason, with_usage
from stop_conditions_and_budgets import GuardedResult, Limits, over_limit, run_guarded_agentTOOL = {
"name": "lookup_job",
"description": "Look up a job by number.",
"input_schema": {"type": "object", "properties": {"job_number": {"type": "string"}}, "required": ["job_number"]},
}
# The time limit is in seconds here, and the fake clock counts seconds too.
LIMITS = Limits(max_steps=10, max_tokens=50_000, max_seconds=120, max_repeats=2)
def opts(**over):
return {
"tools": [TOOL],
"handlers": {"lookup_job": lambda input: '{"suburb":"Footscray"}'},
"limits": LIMITS,
"now": lambda: 0,
**over,
}def test_stops_on_the_token_budget_before_the_next_call_counting_input_and_output():
client = FakeClient([
with_usage(tool_use_reply("lookup_job", {"job_number": "A"}), 30_000, 500),
with_usage(tool_use_reply("lookup_job", {"job_number": "B"}), 20_000, 400),
text_reply("never reached"),
])
result = run_guarded_agent(client, "x", **opts())
assert result == GuardedResult("", "token_budget", 2, 50_900)
assert len(client.calls) == 2Read the test from the top. The first two scripted replies each ask for a tool and report their own usage, and between them they pass the budget. The third reply is never used. The run stops for the token budget, it counted input and output together, and only two calls went out. Scripted usage is what lets a test check the budget without a real run that spends it.
The rest of each test file covers the other ways out, one test for each stop reason:
- Step limit. A script with exactly as many replies as steps allowed. The fake client fails loudly if the loop asks for another.
- Time limit. A fake clock. Each tool call moves it on by more than half the limit, so the second call ends past it. The clock counts milliseconds in TypeScript and seconds in Python, and the test says which.
- Repeating. Three identical replies. The run stops after the third reply, and the handler ran twice.
- Refused. A scripted refusal is reported as refused, not as done.
- Done. A run that finishes normally reports what it spent.
The tests also hold the loop to the rules above. A no from the approval function means the handler never runs, and a yes runs it after asking once. An unknown tool is reported without asking a person, and names like toString are not mistaken for tools. A reply cut off at max_tokens raises an error, and its tool never runs. Two tool requests in one reply come back together, in one user message, in order. A tool that throws comes back as an error result. With no steps allowed, no call is made. In TypeScript, two tests cover the default clock: one runs with no clock passed, and one checks that it reads performance.now().
Now run them. The TypeScript command runs from the top of your project:
npx vitest run examples/learn/stop-conditions-and-budgetsThe Python command runs from the folder that holds the Python files. Run the one for your language, not both:
python -m pytestThese tests need no key and cost nothing, so run them on every change. They show your limits fire when they should and that a no from a person stops a tool. They cannot show what the real model does when it meets a limit, because the model is not in the room. That is what the live run is for.
Step 6: run it for real
Set ANTHROPIC_API_KEY in your environment. The SDK reads it when you create the client, and the live checks do exactly that. The TypeScript one gives the whole test two minutes. The Python one gives the client a two-minute timeout for each request, the client-side bound from the first section. Each file defines its own small job book and lookup tool. Here they are, whole:
import Anthropic from '@anthropic-ai/sdk';
import { describe, it, expect } from 'vitest';
import { runGuardedAgent } from './guarded-loop';
// A stand-in for your job-management system.
const JOBS: Record<string, { suburb: string; booked: string }> = {
'J-1042': { suburb: 'Footscray', booked: '2026-10-14' },
};
const lookupJob: Anthropic.Tool = {
name: 'lookup_job',
description: 'Look up one job in the job book by its job number. Returns the suburb and booked date.',
input_schema: {
type: 'object',
properties: { job_number: { type: 'string', description: 'The job number, for example J-1042.' } },
required: ['job_number'],
},
};
// Skipped unless LEARN_LIVE=1, so an ordinary test run never makes a paid call.
// Run it yourself, with your own key, to watch the limits and the approval step work with the real model.
describe.skipIf(process.env.LEARN_LIVE !== '1')('stop conditions: live', () => {
it('finishes a job lookup inside small limits', async () => {
const result = await runGuardedAgent(new Anthropic(), 'Look up job J-1042 and tell me when it is booked.', {
tools: [lookupJob],
handlers: {
lookup_job: ({ job_number }) => {
// An own-property check: a plain object also answers to names like toString.
if (typeof job_number !== 'string' || !Object.hasOwn(JOBS, job_number)) return 'No such job in the job book.';
return JSON.stringify(JOBS[job_number]);
},
},
limits: { maxSteps: 4, maxTokens: 60_000, maxMillis: 60_000, maxRepeats: 1 },
approve: () => true,
});
console.log(result);
expect(result.stopped).toBe('done');
expect(result.tokens).toBeGreaterThan(0);
}, 120_000);
});import json
import os
import pytest
from stop_conditions_and_budgets import Limits, run_guarded_agent
# A stand-in for your job-management system.
JOBS = {"J-1042": {"suburb": "Footscray", "booked": "2026-10-14"}}
LOOKUP_JOB = {
"name": "lookup_job",
"description": "Look up one job in the job book by its job number. Returns the suburb and booked date.",
"input_schema": {
"type": "object",
"properties": {"job_number": {"type": "string", "description": "The job number, for example J-1042."}},
"required": ["job_number"],
},
}
def lookup_job(input):
job_number = input.get("job_number")
if not isinstance(job_number, str) or job_number not in JOBS:
return "No such job in the job book."
return json.dumps(JOBS[job_number])
# Skipped unless LEARN_LIVE=1, so an ordinary test run never makes a paid call.
# Run it yourself, with your own key, to watch the limits and the approval step work with the real model.
@pytest.mark.skipif(os.environ.get("LEARN_LIVE") != "1", reason="live API check; run by a person with LEARN_LIVE=1")
def test_live_finishes_a_job_lookup_inside_small_limits():
import anthropic
# Give up on any one request after 120 seconds.
client = anthropic.Anthropic(timeout=120.0)
result = run_guarded_agent(
client,
"Look up job J-1042 and tell me when it is booked.",
tools=[LOOKUP_JOB],
handlers={"lookup_job": lookup_job},
limits=Limits(max_steps=4, max_tokens=60000, max_seconds=60, max_repeats=1),
approve=lambda call: True,
)
print(result)
assert result.stopped == "done"
assert result.tokens > 0Then run one. Use the command for your language. The TypeScript command runs from the top of your project:
LEARN_LIVE=1 npx vitest run examples/learn/stop-conditions-and-budgets/guarded-loop.live.test.tsThe Python command runs from the folder that holds the Python files:
LEARN_LIVE=1 python -m pytest test_live_stop_conditions_and_budgets.pyBoth live checks skip themselves unless the LEARN_LIVE variable is set to one, so an ordinary test run never calls the API. The commands above set it in front of the program, which works in bash and zsh. In Windows PowerShell, use the matching form below instead, and remove the variable afterwards so ordinary test runs stay off the API:
$env:LEARN_LIVE=1; npx vitest run examples/learn/stop-conditions-and-budgets/guarded-loop.live.test.ts
Remove-Item Env:LEARN_LIVE$env:LEARN_LIVE=1; python -m pytest test_live_stop_conditions_and_budgets.py
Remove-Item Env:LEARN_LIVEThe check asks the loop to look up one job and say when it is booked. It expects the run to finish on its own, with some tokens counted, and it prints the whole result, so you can read the answer and see the steps and tokens used. The limits are small on purpose: a few steps, a modest budget, a minute, and a repeat limit of one, so the first repeated call stops the run. Keep them small the first time you run a new loop. If it misbehaves, it stops early and cheaply.
The approval function here says yes to everything, so this run does not test a no. The fake client tests do that. One run spends a little usage, and it passing tells you the model finished this time, not that it always will. A live check is one sample. The steps and tokens it prints are your first real numbers for the next section.
This is also the run that earns the guide its Last checked date. Passing against the fake client is not enough. Someone has to run it against the real API, and only then does the date move.
How to pick numbers that fit your job
No set of numbers fits every job. Start from the shape of the job you have, and work outwards.
- Steps. Count the steps a person would take to do the job, then allow a few more for a retry. A limit that is many times the real number is not doing its job.
- Budget. Decide what you would accept paying for one run, and work backwards. API usage is priced per million tokens, with separate rates for the tokens the model reads and the tokens it writes. The Claude API pricing page has the current rates for each model, and this guide states none, because they change. Turn your amount into a token count from that page, then set the budget a call's worth below it, because one call can still take a run over. Watch a few real runs first. The result reports the tokens each one used.
- Time. Keep the time limit well inside what your caller will wait. A person watching a screen will wait far less than a job that runs overnight. Count the wait for approvals too. Set the client's timeout, as in the first section, and give any tool that can hang a timeout of its own, so one stuck request cannot hold the run up for long.
- Repeats. Keep it low, one or two, unless the job really does repeat a call.
- Approval. Put the pause in front of whatever you cannot take back. Start with more approvals than you think you need, and drop one only once you have seen the tool behave.
When a run hits a limit, treat that as a result, not a crash, and log which limit it was. A limit that fires often is telling you something about the job, the prompt or the tools, and raising it is the last thing to try. Start with tools that only read, keep every limit low the first time, and add the tools that change things last, with a person approving them.
Want this built for you instead? See how we build AI workflows, or book a free 30-minute consultation.

Peter McLean
Founder, Neurastruct
Australian small-business operator since 2001 and 16 years as a national account manager; AI certificates from Anthropic (2026) and Google (2025).
© Neurastruct Pty Ltd. Text licensed CC BY 4.0. Code samples licensed MIT. CC BY 4.0 · MIT