Building
How to build an AI agent for your business: steps, costs and guardrails
How to build an AI agent in 10 steps: when you need one, which model to use, what it costs to run, and the guardrails that keep it from going wrong.
In short
An AI agent is software that uses a language model to decide its own next steps (what to look up, which tool to call, when it’s done) to finish a task for you. To build one, pick one job with a measurable outcome, collect real past cases as a test set, give it only the tools it needs, write its instructions from how your team already works, and launch with a person reviewing its work. Model fees are usually the small part: in the illustration below, 2,000 conversations a month cost between about $6 and $128 depending on the model. The real risks are wrong answers and actions that can’t be undone, so every agent needs sources, a handoff to a person and a spending cap. Often, a simple workflow or a single model call does the job better than an agent.
Key takeaways
- Start with the simplest thing that works: a workflow or a single model call often beats an agent.
- Give an agent one job, a test set of real past cases, and a pass bar before it meets a customer.
- Model fees are usually small; the real costs are the build, the upkeep and the mistakes.
- You are responsible for what your agent says, as rulings in Canada and Germany show.
- Keep a person in the loop for anything that moves money or can’t be undone.
What is an AI agent?
An AI agent is software that uses a language model to decide its own next steps (what to look up, which tool to call, when the job is done) to complete a task for you. OpenAI’s guide to building agents puts it in one line: “Agents are systems that independently accomplish tasks on your behalf.” Google Cloud’s definition is close: “software systems that use AI to pursue goals and complete tasks on behalf of users.”
The word that matters is decide. Anthropic draws the line between two kinds of system: “Workflows are systems where LLMs and tools are orchestrated through predefined code paths. Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage.” Both are useful, and most businesses need the first more often than the second.
How is an AI agent different from a chatbot or an automation?
| Question | Chatbot | Workflow automation | AI agent |
|---|---|---|---|
| Who decides the next step? | Nobody: it answers one message at a time | You do, in advance | The model, within limits you set |
| Can it act in your tools? | Rarely | Yes, the same way every time | Yes, choosing which tool and when |
| Best for | Common questions and simple routing | Fixed steps with known rules | Judgment calls on messy input |
| Example | “What are your opening hours?” | A form entry goes to the CRM, then a welcome email | Read a refund request, check the order, decide, draft the reply |
| Main risk | Wrong or made-up answers | Breaks when the input changes | Wrong actions, errors that compound, higher cost |
OpenAI’s guide is explicit that apps which use a model without letting it control the work (“simple chatbots, single-turn LLMs, or sentiment classifiers”) “are not agents.” A chatbot that answers from your help center can be very useful. It becomes an agent when it can look things up and act.
Do you actually need an AI agent?
Often not. Anthropic, which makes the Claude models, recommends “finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all.” It adds that for many applications, “optimizing single LLM calls with retrieval and in-context examples is usually enough.”
OpenAI’s guide names three signs that a job suits an agent: complex decisions full of exceptions and judgment, rules that have become hard to maintain, and heavy reliance on unstructured data such as emails, documents and conversations. If a job doesn’t clearly meet them, “a deterministic solution may suffice.”
| Your situation | What to build |
|---|---|
| Fixed steps and known rules, such as a form entry going to your CRM | A workflow automation, not an agent |
| Answering questions from your documents, with no actions | A retrieval assistant: one model call, given the right documents |
| Judgment on messy input, two or three tools, and a clear way to check the result | A single-job AI agent |
| Actions that move money or can’t be undone, legal exposure, or no past cases to test on | An agent that drafts while a person approves, or nothing yet |
| Low volume, or nobody to own it after launch | Nothing: a saved reply or a checklist is cheaper |
What are the types of AI agents?
Textbooks, and IBM’s overview, sort agents by how they decide:
- Simple reflex agents act on the current input with fixed rules.
- Model-based reflex agents also keep track of what they have seen so far.
- Goal-based agents plan a sequence of steps toward a goal.
- Utility-based agents weigh the options and pick the best outcome.
- Learning agents improve from feedback over time.
Agents built on today’s language models mix several of these. For a business, it is more useful to sort them by the job they do.
What are examples of AI agents in a business?
| Agent | What it does | Where a person stays in |
|---|---|---|
| Customer support agent | Answers from your help docs, looks up orders or accounts, drafts or sends the reply | Refunds, cancellations, upset customers, and anything it isn’t sure of |
| Lead qualification agent | Asks the questions your sales team would, scores the answers, books or routes the good leads | Pricing, custom deals, and any lead it can’t score clearly |
| Document and inbox triage | Reads incoming email, forms or files, extracts what matters, and files or routes it | A review queue for anything it can’t classify confidently |
| Internal knowledge assistant | Answers staff questions from your own documents, with sources, within each person’s access | Policy questions the documents don’t cover |
| Operations agent | Compares records between two systems and flags the mismatches | Fixing each mismatch, until it has proven reliable |
| Coding agent | Writes and changes code in a sandbox, runs the tests, and proposes the change | Reviewing and merging; it never touches production directly |
How do you build an AI agent, step by step?
- Pick one job and define done. “Answer order-status questions and hand off the rest with a summary” is a job; “handle support” is not. Write down how you will measure it: cases resolved without a person, answers checked as correct, time saved.
- Collect 50–200 real past cases as a test set. Real questions, emails or tickets, each with the right answer or action. This is how you will know whether the agent works, and whether a later change broke it.
- List the data and the fewest tools it needs, and rate each tool’s risk. Reading an order is low risk. Sending an email is medium. Issuing a refund or deleting data is high: those need a person’s approval, or stay out of scope.
- Write its instructions from how your team already works. Your support macros, policies and escalation rules are the first draft. Spell out what it must never do, and what it says when it doesn’t know.
- Prototype on a capable model, then try cheaper ones. OpenAI suggests building the prototype “with the most capable model for every task to establish a performance baseline,” then swapping in smaller models to see if they still pass.
- Run the test set and fix what fails. Score every case, fix the instructions, the data or the tools, and run it again. Set a pass bar before launch, and don’t lower it to hit a date.
- Design the handoff to a person. Decide when it escalates (after repeated failures, before high-risk actions, or whenever someone asks) and pass the conversation along, so nobody has to start over.
- Launch in shadow mode, or with review. Let it draft while a person sends, or let it answer a small share of traffic, before it acts on its own.
- Cap spending and turn on logging. A monthly cost ceiling, a limit on retries per conversation, a log of every tool call, and alerts when something spikes.
- Review weekly, and grow slowly. Read a sample of conversations, add every failure to the test set, and add a new tool only once the last one has proven itself. In OpenAI’s words: “Start small, validate with real users, and grow capabilities over time.”
Which AI model should you use for an agent?
The cheapest one that passes your test set. Prices change often, and the same text can cost different amounts on different models: Anthropic notes that its newer tokenizer “produces approximately 30% more tokens for the same text.” Compare models on your own cases and your own bills, not on price lists alone.
| Model | Input | Output | Worth knowing |
|---|---|---|---|
| Claude Sonnet 5.5 (Anthropic) | $2 | $10 | Cached input $0.20 |
| Claude Haiku 4.5 (Anthropic) | $1 | $5 | Cached input $0.10 |
| gpt-6.1-sol (OpenAI) | $2 | $10 | Cached input $0.10 |
| gpt-6-luna (OpenAI) | $0.10 | $0.50 | Cached input $0.01 |
| Gemini 3.8 Flash (Google) | $0.75 | $3.75 | Rises to $1.50 and $7.50 on January 1, 2027 |
A token is a piece of a word. In English, Anthropic estimates one token at about four characters, or three-quarters of a word.
What does an AI agent cost to run?
Model fees are arithmetic: conversations × calls × tokens × price. Here is an illustration, not a quote: a support agent handling 2,000 conversations a month, making four model calls per conversation, each reading about 6,000 tokens (its instructions, the documents it retrieved, the conversation so far) and writing about 400.
| Model | Per conversation | Per month |
|---|---|---|
| Claude Sonnet 5.5 or gpt-6.1-sol | $0.064 | $128 |
| Claude Sonnet 5.5, with 4,000 tokens of each call cached | About $0.035 | About $70 |
| Claude Haiku 4.5 | $0.032 | $64 |
| Gemini 3.8 Flash, at its 2026 price | $0.024 | $48 |
| Gemini 3.8 Flash, from January 2027 | $0.048 | $96 |
| gpt-6-luna | $0.0032 | $6.40 |
The cached row leaves out the cost of writing the cache the first time, and tokenizers differ, so run your own test set for real numbers. Anthropic’s own example lands in the same range: about 3,700 tokens per conversation on Haiku 4.5 comes to “~$37.00 per 10,000 tickets.”
What the table leaves out is often larger: hosting, search and other paid tools, monitoring, upkeep, and the people who handle the handoffs. Gartner predicts that “by 2030, cost per resolution for generative AI will exceed $3, higher than many B2C offshore human agents” (as reported by MarketScreener). The lesson isn’t that agents are expensive. It is that automating every case is. Let the agent take the routine cases and hand the rest to a person.
What does it cost to build an AI agent?
Published ranges are wide and mostly unsourced; one agency guide quotes “$6,000 to over $300,000.” The spread comes from scope:
- How many jobs it does. One job is a small project. Several agents coordinating with each other is a large one.
- How many systems it touches, and whether they have APIs.
- How clean the data is. Current help docs are quick to use; scattered PDFs and years of old tickets are not.
- Whether a test set exists, or has to be written from scratch.
- Compliance. HIPAA, SOC 2 or financial rules add review, paperwork and time.
No-code agent builders cost less to start and suit simple jobs. Code earns its cost when the agent needs your own systems, exact permissions, or tests that run on every change.
A single-job agent (one job, two or three tools, answers from your data with sources) is the kind we build in about a week, at a fixed price agreed in writing after a free scoping call. See AI agents with guardrails, built in a week.
How long does it take to build an AI agent?
A prototype takes a day. A reliable agent takes longer, because most of the work is the test set, the edge cases and the handoff. A single-job agent with a few tools and a written test set can be built, hardened and launched in about a week when the job is scoped first. Several agents working together, a model trained on your own data, or a formal compliance program take longer: plan in months, not weeks.
What can go wrong with an AI agent, and who is responsible?
Gartner predicted in June 2025 that more than 40% of agentic AI projects will be canceled by the end of 2027, “due to escalating costs, unclear business value or inadequate risk controls” (as reported by Outlook Business). The risk controls are the part you decide. Four public cases show why:
- Air Canada, 2024. Its website chatbot told a customer that a bereavement fare could be claimed after the flight, which Air Canada’s policy didn’t allow. Air Canada argued, in effect, that the chatbot was “a separate legal entity that is responsible for its own actions.” The tribunal disagreed: “It should be obvious to Air Canada that it is responsible for all the information on its website.” Air Canada was ordered to pay $812.02.
- Cursor, April 2025. The code editor’s AI support agent told users about a login policy that didn’t exist (The Register). A cofounder apologized on Hacker News, refunded the user, and said AI responses used for email support “are now clearly labeled as such.”
- Replit, July 2025. An AI coding agent deleted data from a production database during development. Replit’s CEO called it “Unacceptable and should never be possible.”
- Germany, May 2026. A clinic’s chatbot described its two doctors as specialists in plastic and aesthetic surgery, credentials they didn’t hold. The Higher Regional Court of Hamm held that the chatbot is not a “third party”: the clinic is responsible for what it says. An appeal to the Federal Court of Justice was allowed.
The pattern: what your agent says, you said. Customers expect a way out, too. In a 2026 Gartner survey of 3,566 customers, 87% said it is essential to have an option to reach a human agent when companies use generative AI (Insurance-Canada), and only 27% said they would try a chatbot again after a negative experience (IT-Online).
What guardrails does an AI agent need?
- Answers cite a source from your documents, or the agent says it doesn’t know.
- Tools are on an allowlist, and anything that writes, sends or pays needs a person’s approval.
- A retry limit, after which the case goes to a person.
- A visible way to reach a person, in every conversation.
- AI replies are labeled as AI.
- Answers about policies, prices and credentials come only from your source documents, never from the model’s memory.
- The agent is built and tested away from live data, and never gets access it doesn’t need in production.
- A monthly spending cap, a log of every tool call, and alerts.
- The test set must pass before any change goes live.
- The test set includes prompt-injection cases: messages and documents that try to give the agent new instructions.
- Personal data is kept out of logs, or kept only as long as needed.
What fits in a one-week AI agent build?
One agent, one job, with guardrails:
- Answers grounded in your documents, with sources.
- Two or three tools it can call, such as an order lookup.
- A handoff to a person when it isn’t sure, with the conversation attached.
- A monthly cost ceiling and usage logging.
- A test set of real questions it must pass, and the results it passed at launch.
- Chat on your site, or inside a tool you already use.
What doesn’t fit in a week: actions that move money without a person’s approval, several agents coordinating with each other, training or fine-tuning your own model, or formal compliance programs such as HIPAA or SOC 2. See how a one-week agent build works.
What next?
Write down the one job, the ten cases you see most often, and what the agent must never do. If the job passes the checks above, send us your requirements or see how a one-week agent build works. If a workflow or a single model call would do the job, we’ll tell you plainly.
Sources
- Anthropic — Building effective agents
- OpenAI — A practical guide to building agents (PDF)
- Google Cloud — What are AI agents?
- IBM — Types of AI agents
- Anthropic — Claude API pricing
- OpenAI — API pricing
- Google — Gemini API pricing
- Uptech — How to build an AI agent
- Outlook Business — Gartner on agentic AI project cancellations
- MarketScreener — Gartner on GenAI cost per resolution
- Insurance-Canada — Gartner survey on reaching a human agent
- IT-Online — Gartner survey on chatbots after a negative experience
- Civil Resolution Tribunal — Moffatt v. Air Canada, 2024 BCCRT 149
- The Register — Cursor’s AI support bot
- Hacker News — Cursor cofounder’s reply
- Amjad Masad (Replit CEO) on X
- Justiz NRW — OLG Hamm, 4 UKl 3/25 (German)