PutThrough Product engineering studio
Book a build call

Building

How to build an AI agent for your business: steps, costs and guardrails

How to build an AI agent in 10 steps: when you need one, which model to use, what it costs to run, and the guardrails that keep it from going wrong.

In short

An AI agent is software that uses a language model to decide its own next steps (what to look up, which tool to call, when it’s done) to finish a task for you. To build one, pick one job with a measurable outcome, collect real past cases as a test set, give it only the tools it needs, write its instructions from how your team already works, and launch with a person reviewing its work. Model fees are usually the small part: in the illustration below, 2,000 conversations a month cost between about $6 and $128 depending on the model. The real risks are wrong answers and actions that can’t be undone, so every agent needs sources, a handoff to a person and a spending cap. Often, a simple workflow or a single model call does the job better than an agent.

Key takeaways

  • Start with the simplest thing that works: a workflow or a single model call often beats an agent.
  • Give an agent one job, a test set of real past cases, and a pass bar before it meets a customer.
  • Model fees are usually small; the real costs are the build, the upkeep and the mistakes.
  • You are responsible for what your agent says, as rulings in Canada and Germany show.
  • Keep a person in the loop for anything that moves money or can’t be undone.

What is an AI agent?

An AI agent is software that uses a language model to decide its own next steps (what to look up, which tool to call, when the job is done) to complete a task for you. OpenAI’s guide to building agents puts it in one line: “Agents are systems that independently accomplish tasks on your behalf.” Google Cloud’s definition is close: “software systems that use AI to pursue goals and complete tasks on behalf of users.”

The word that matters is decide. Anthropic draws the line between two kinds of system: “Workflows are systems where LLMs and tools are orchestrated through predefined code paths. Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage.” Both are useful, and most businesses need the first more often than the second.

How is an AI agent different from a chatbot or an automation?

A chatbot, a workflow automation and an AI agent, compared
QuestionChatbotWorkflow automationAI agent
Who decides the next step?Nobody: it answers one message at a timeYou do, in advanceThe model, within limits you set
Can it act in your tools?RarelyYes, the same way every timeYes, choosing which tool and when
Best forCommon questions and simple routingFixed steps with known rulesJudgment calls on messy input
Example“What are your opening hours?”A form entry goes to the CRM, then a welcome emailRead a refund request, check the order, decide, draft the reply
Main riskWrong or made-up answersBreaks when the input changesWrong actions, errors that compound, higher cost

OpenAI’s guide is explicit that apps which use a model without letting it control the work (“simple chatbots, single-turn LLMs, or sentiment classifiers”) “are not agents.” A chatbot that answers from your help center can be very useful. It becomes an agent when it can look things up and act.

Do you actually need an AI agent?

Often not. Anthropic, which makes the Claude models, recommends “finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all.” It adds that for many applications, “optimizing single LLM calls with retrieval and in-context examples is usually enough.”

OpenAI’s guide names three signs that a job suits an agent: complex decisions full of exceptions and judgment, rules that have become hard to maintain, and heavy reliance on unstructured data such as emails, documents and conversations. If a job doesn’t clearly meet them, “a deterministic solution may suffice.”

What to build, by situation
Your situationWhat to build
Fixed steps and known rules, such as a form entry going to your CRMA workflow automation, not an agent
Answering questions from your documents, with no actionsA retrieval assistant: one model call, given the right documents
Judgment on messy input, two or three tools, and a clear way to check the resultA single-job AI agent
Actions that move money or can’t be undone, legal exposure, or no past cases to test onAn agent that drafts while a person approves, or nothing yet
Low volume, or nobody to own it after launchNothing: a saved reply or a checklist is cheaper

What are the types of AI agents?

Textbooks, and IBM’s overview, sort agents by how they decide:

  • Simple reflex agents act on the current input with fixed rules.
  • Model-based reflex agents also keep track of what they have seen so far.
  • Goal-based agents plan a sequence of steps toward a goal.
  • Utility-based agents weigh the options and pick the best outcome.
  • Learning agents improve from feedback over time.

Agents built on today’s language models mix several of these. For a business, it is more useful to sort them by the job they do.

What are examples of AI agents in a business?

Common AI agents, what they do, and where a person stays in
AgentWhat it doesWhere a person stays in
Customer support agentAnswers from your help docs, looks up orders or accounts, drafts or sends the replyRefunds, cancellations, upset customers, and anything it isn’t sure of
Lead qualification agentAsks the questions your sales team would, scores the answers, books or routes the good leadsPricing, custom deals, and any lead it can’t score clearly
Document and inbox triageReads incoming email, forms or files, extracts what matters, and files or routes itA review queue for anything it can’t classify confidently
Internal knowledge assistantAnswers staff questions from your own documents, with sources, within each person’s accessPolicy questions the documents don’t cover
Operations agentCompares records between two systems and flags the mismatchesFixing each mismatch, until it has proven reliable
Coding agentWrites and changes code in a sandbox, runs the tests, and proposes the changeReviewing and merging; it never touches production directly

How do you build an AI agent, step by step?

  1. Pick one job and define done. “Answer order-status questions and hand off the rest with a summary” is a job; “handle support” is not. Write down how you will measure it: cases resolved without a person, answers checked as correct, time saved.
  2. Collect 50–200 real past cases as a test set. Real questions, emails or tickets, each with the right answer or action. This is how you will know whether the agent works, and whether a later change broke it.
  3. List the data and the fewest tools it needs, and rate each tool’s risk. Reading an order is low risk. Sending an email is medium. Issuing a refund or deleting data is high: those need a person’s approval, or stay out of scope.
  4. Write its instructions from how your team already works. Your support macros, policies and escalation rules are the first draft. Spell out what it must never do, and what it says when it doesn’t know.
  5. Prototype on a capable model, then try cheaper ones. OpenAI suggests building the prototype “with the most capable model for every task to establish a performance baseline,” then swapping in smaller models to see if they still pass.
  6. Run the test set and fix what fails. Score every case, fix the instructions, the data or the tools, and run it again. Set a pass bar before launch, and don’t lower it to hit a date.
  7. Design the handoff to a person. Decide when it escalates (after repeated failures, before high-risk actions, or whenever someone asks) and pass the conversation along, so nobody has to start over.
  8. Launch in shadow mode, or with review. Let it draft while a person sends, or let it answer a small share of traffic, before it acts on its own.
  9. Cap spending and turn on logging. A monthly cost ceiling, a limit on retries per conversation, a log of every tool call, and alerts when something spikes.
  10. Review weekly, and grow slowly. Read a sample of conversations, add every failure to the test set, and add a new tool only once the last one has proven itself. In OpenAI’s words: “Start small, validate with real users, and grow capabilities over time.”

Which AI model should you use for an agent?

The cheapest one that passes your test set. Prices change often, and the same text can cost different amounts on different models: Anthropic notes that its newer tokenizer “produces approximately 30% more tokens for the same text.” Compare models on your own cases and your own bills, not on price lists alone.

API prices per million tokens, standard tier, checked October 1, 2026
ModelInputOutputWorth knowing
Claude Sonnet 5.5 (Anthropic)$2$10Cached input $0.20
Claude Haiku 4.5 (Anthropic)$1$5Cached input $0.10
gpt-6.1-sol (OpenAI)$2$10Cached input $0.10
gpt-6-luna (OpenAI)$0.10$0.50Cached input $0.01
Gemini 3.8 Flash (Google)$0.75$3.75Rises to $1.50 and $7.50 on January 1, 2027

A token is a piece of a word. In English, Anthropic estimates one token at about four characters, or three-quarters of a word.

What does an AI agent cost to run?

Model fees are arithmetic: conversations × calls × tokens × price. Here is an illustration, not a quote: a support agent handling 2,000 conversations a month, making four model calls per conversation, each reading about 6,000 tokens (its instructions, the documents it retrieved, the conversation so far) and writing about 400.

Model fees for 2,000 conversations a month (an illustration)
ModelPer conversationPer month
Claude Sonnet 5.5 or gpt-6.1-sol$0.064$128
Claude Sonnet 5.5, with 4,000 tokens of each call cachedAbout $0.035About $70
Claude Haiku 4.5$0.032$64
Gemini 3.8 Flash, at its 2026 price$0.024$48
Gemini 3.8 Flash, from January 2027$0.048$96
gpt-6-luna$0.0032$6.40

The cached row leaves out the cost of writing the cache the first time, and tokenizers differ, so run your own test set for real numbers. Anthropic’s own example lands in the same range: about 3,700 tokens per conversation on Haiku 4.5 comes to “~$37.00 per 10,000 tickets.”

What the table leaves out is often larger: hosting, search and other paid tools, monitoring, upkeep, and the people who handle the handoffs. Gartner predicts that “by 2030, cost per resolution for generative AI will exceed $3, higher than many B2C offshore human agents” (as reported by MarketScreener). The lesson isn’t that agents are expensive. It is that automating every case is. Let the agent take the routine cases and hand the rest to a person.

What does it cost to build an AI agent?

Published ranges are wide and mostly unsourced; one agency guide quotes “$6,000 to over $300,000.” The spread comes from scope:

  • How many jobs it does. One job is a small project. Several agents coordinating with each other is a large one.
  • How many systems it touches, and whether they have APIs.
  • How clean the data is. Current help docs are quick to use; scattered PDFs and years of old tickets are not.
  • Whether a test set exists, or has to be written from scratch.
  • Compliance. HIPAA, SOC 2 or financial rules add review, paperwork and time.

No-code agent builders cost less to start and suit simple jobs. Code earns its cost when the agent needs your own systems, exact permissions, or tests that run on every change.

A single-job agent (one job, two or three tools, answers from your data with sources) is the kind we build in about a week, at a fixed price agreed in writing after a free scoping call. See AI agents with guardrails, built in a week.

How long does it take to build an AI agent?

A prototype takes a day. A reliable agent takes longer, because most of the work is the test set, the edge cases and the handoff. A single-job agent with a few tools and a written test set can be built, hardened and launched in about a week when the job is scoped first. Several agents working together, a model trained on your own data, or a formal compliance program take longer: plan in months, not weeks.

What can go wrong with an AI agent, and who is responsible?

Gartner predicted in June 2025 that more than 40% of agentic AI projects will be canceled by the end of 2027, “due to escalating costs, unclear business value or inadequate risk controls” (as reported by Outlook Business). The risk controls are the part you decide. Four public cases show why:

  • Air Canada, 2024. Its website chatbot told a customer that a bereavement fare could be claimed after the flight, which Air Canada’s policy didn’t allow. Air Canada argued, in effect, that the chatbot was “a separate legal entity that is responsible for its own actions.” The tribunal disagreed: “It should be obvious to Air Canada that it is responsible for all the information on its website.” Air Canada was ordered to pay $812.02.
  • Cursor, April 2025. The code editor’s AI support agent told users about a login policy that didn’t exist (The Register). A cofounder apologized on Hacker News, refunded the user, and said AI responses used for email support “are now clearly labeled as such.”
  • Replit, July 2025. An AI coding agent deleted data from a production database during development. Replit’s CEO called it “Unacceptable and should never be possible.”
  • Germany, May 2026. A clinic’s chatbot described its two doctors as specialists in plastic and aesthetic surgery, credentials they didn’t hold. The Higher Regional Court of Hamm held that the chatbot is not a “third party”: the clinic is responsible for what it says. An appeal to the Federal Court of Justice was allowed.

The pattern: what your agent says, you said. Customers expect a way out, too. In a 2026 Gartner survey of 3,566 customers, 87% said it is essential to have an option to reach a human agent when companies use generative AI (Insurance-Canada), and only 27% said they would try a chatbot again after a negative experience (IT-Online).

What guardrails does an AI agent need?

  • Answers cite a source from your documents, or the agent says it doesn’t know.
  • Tools are on an allowlist, and anything that writes, sends or pays needs a person’s approval.
  • A retry limit, after which the case goes to a person.
  • A visible way to reach a person, in every conversation.
  • AI replies are labeled as AI.
  • Answers about policies, prices and credentials come only from your source documents, never from the model’s memory.
  • The agent is built and tested away from live data, and never gets access it doesn’t need in production.
  • A monthly spending cap, a log of every tool call, and alerts.
  • The test set must pass before any change goes live.
  • The test set includes prompt-injection cases: messages and documents that try to give the agent new instructions.
  • Personal data is kept out of logs, or kept only as long as needed.

What fits in a one-week AI agent build?

One agent, one job, with guardrails:

  • Answers grounded in your documents, with sources.
  • Two or three tools it can call, such as an order lookup.
  • A handoff to a person when it isn’t sure, with the conversation attached.
  • A monthly cost ceiling and usage logging.
  • A test set of real questions it must pass, and the results it passed at launch.
  • Chat on your site, or inside a tool you already use.

What doesn’t fit in a week: actions that move money without a person’s approval, several agents coordinating with each other, training or fine-tuning your own model, or formal compliance programs such as HIPAA or SOC 2. See how a one-week agent build works.

What next?

Write down the one job, the ten cases you see most often, and what the agent must never do. If the job passes the checks above, send us your requirements or see how a one-week agent build works. If a workflow or a single model call would do the job, we’ll tell you plainly.

Sources

  1. Anthropic — Building effective agents
  2. OpenAI — A practical guide to building agents (PDF)
  3. Google Cloud — What are AI agents?
  4. IBM — Types of AI agents
  5. Anthropic — Claude API pricing
  6. OpenAI — API pricing
  7. Google — Gemini API pricing
  8. Uptech — How to build an AI agent
  9. Outlook Business — Gartner on agentic AI project cancellations
  10. MarketScreener — Gartner on GenAI cost per resolution
  11. Insurance-Canada — Gartner survey on reaching a human agent
  12. IT-Online — Gartner survey on chatbots after a negative experience
  13. Civil Resolution Tribunal — Moffatt v. Air Canada, 2024 BCCRT 149
  14. The Register — Cursor’s AI support bot
  15. Hacker News — Cursor cofounder’s reply
  16. Amjad Masad (Replit CEO) on X
  17. Justiz NRW — OLG Hamm, 4 UKl 3/25 (German)

About the author

Anubhav Mehrotra

Founder and principal engineer at PutThrough. Six-plus years shipping software end to end — Android and iOS apps, backends, frontends, infrastructure and AI agents.

Questions

Questions this raises

Something we haven’t covered?

Ask on a call
01 How much does it cost to build an AI agent?

Published ranges run from about $6,000 to over $300,000, depending on scope. A single-job agent with a few tools and your own documents sits at the small end; several agents, many integrations and compliance programs sit at the large end. Running one is usually cheaper than building it: at current prices, model fees for 2,000 conversations a month come to roughly $6–$128, before hosting and review.

02 How long does it take to build an AI agent?

A prototype takes a day. A single-job agent with two or three tools, a test set and a handoff to a person takes about a week when the job is scoped first. Multi-agent systems and regulated work take longer.

03 What is the difference between an AI agent and a chatbot?

A chatbot answers messages. An AI agent decides its own next steps and acts in your tools to finish a task, such as looking up an order, updating a record or drafting a reply. OpenAI’s guide to building agents says simple chatbots are not agents.

04 Can I build an AI agent without coding?

Yes, for simple jobs: no-code agent builders let you connect documents and a few tools. The same rules apply either way: one job, a test set, a handoff to a person and a spending limit. Code becomes worth it when the agent needs your own systems, exact permissions or tests that run on every change.

05 What are the types of AI agents?

The classic five are simple reflex, model-based reflex, goal-based, utility-based and learning agents. For a business, it is more useful to sort them by job: customer support, lead qualification, document and inbox triage, and internal knowledge assistants.

06 Who is responsible when an AI agent makes a mistake?

The business that runs it. A Canadian tribunal ordered Air Canada to pay after its chatbot gave a wrong answer about fares, and a German court held a clinic responsible for its chatbot’s false claims about its doctors. Label AI replies, answer policy questions only from your source documents, and keep a person on anything that can’t be undone.

07 Which AI model is best for an AI agent?

The cheapest one that passes your test set. Prototype with a capable model to set a baseline, then try smaller, cheaper ones. Prices change often and tokenizers differ, so compare models on your own cases rather than on price lists.

08 Why do AI agent projects fail?

Gartner predicted in 2025 that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. The usual causes are no single job, no test set, too many tools, and no plan for handing off to a person.

Have something to build?

A free 30-minute call. We’ll tell you honestly what fits in a week, and send a fixed quote afterward.

Prefer to write? Send your requirements

What do you need?

What it should do, who will use it, the must-haves, and any links. A few sentences is plenty.

Budget, if you have one in mind
When would you like to start?