PutThrough Product engineering studio
Book a build call

Service 05

AI agent and LLM feature development

A demo that works once is easy. A feature that behaves on the thousandth call, at a cost you can predict, is the job.

Prefer to write? Send your requirements

Agents, tool-using features and MCP servers, built so that you can tell whether they are working. Most LLM features fail in ways ordinary tests do not catch — they get slower, more expensive, or quietly worse after a prompt change — so evaluation and cost control are part of the build rather than something added once the bill arrives.

Agents and tools

  • Tool and function design, including the failure and refusal paths
  • Retrieval that returns the right context rather than the most text
  • Multi-step flows with state, retries and a stopping condition
  • Streaming, cancellation and the interface around a slow response

MCP servers

  • Your product exposed as an MCP server, with real tool definitions
  • OAuth 2.1 with scoped permissions rather than a shared static token
  • Protected resource metadata and resource indicators, per the current spec
  • Audit logging of what an agent did on whose behalf

Evaluation

  • An eval set built from your real cases, not invented examples
  • Regression tests that fail the build when a prompt change degrades quality
  • Tracing, so a bad answer can be traced to the call that produced it

Cost and safety

  • Caching, batching and routing to cut spend without losing quality
  • Per-tenant budgets and rate limits on paid calls
  • Prompt injection and tool-permission controls on anything agentic
  • Disclosure where a user is talking to a machine

Built with

  • Claude
  • OpenAI
  • MCP
  • TypeScript
  • Python
  • Postgres / pgvector
  • Evals and tracing

What this is not

  • Not model training or fine-tuning. We build on hosted models, and we will say when a problem does not need a model at all.
  • Not a guarantee of model behavior. We can measure quality and hold it to a threshold; we cannot promise a model never gets something wrong.
  • Not a drop-in chatbot widget. If that is what you need, a product off the shelf will be cheaper than us.

01Self-assessment

Signs this is the work you need

Things that are either true of your situation or not. If the first two are false, you may not need a model at all — and we would rather say so.

  • The task involves judgement on unstructured input: language, documents, images, messy data.
  • A competent person could do it given the same information, but the volume makes that impractical.
  • You already have an LLM feature in production and cannot tell whether a prompt change made it better or worse.
  • Your model spend is growing faster than your usage, or you cannot attribute it to customers.
  • A user has reported a wrong answer and you could not reconstruct what the model was actually sent.
  • Customers are asking whether your product works with their AI assistant, and you have no integration story.
  • You have an agent that can take actions and no audit trail of what it did on whose behalf.
  • Your agent loop has no step limit and no cost ceiling.
  • You are sending whole tables or document sets into every prompt because it was the quickest thing that worked.

02Approach

How we approach it

Most "agent" problems are not agent problems

An agent is the most expensive, least predictable and hardest-to-test shape a feature can take. It earns that cost when the sequence of steps genuinely cannot be known in advance — when the work depends on what earlier steps found. That is a real category, and it is smaller than the current enthusiasm suggests.

A great many features described as agents are a fixed sequence: classify this, extract those fields, look up the record, draft a reply. Written as a workflow with model calls at the specific points that need judgement, that is cheaper, faster, testable, and far easier to debug when it goes wrong. Written as an agent with six tools and a loop, it is none of those.

The first thing we do on this work is push in the other direction: what can be decided by ordinary code, what needs one model call, and what genuinely needs the model to choose its own path. Complexity you did not need is complexity you will still be paying for in a year.

Evals before features, because the failure is invisible

Ordinary software fails loudly. An LLM feature fails by getting slightly worse — a little less accurate, a little more verbose, a little more likely to refuse something reasonable — and nothing in your test suite or your error tracker registers any of it.

That makes prompt changes uniquely dangerous. Everyone has shipped an improvement to one case that quietly degraded five others, because there was no way to see the five. An eval set is the instrument that makes the degradation visible: a collection of your real cases with expected outcomes, scored automatically, run in CI, failing the build when the number drops.

Building it is unglamorous and it is the difference between a feature you can improve and a feature you can only fiddle with. We build it first, from your actual data, before writing the thing it will measure. Fifty real cases beat five hundred invented ones.

Cost is a design parameter, not a line on a bill

Per-call pricing means architecture decides your unit economics. A retrieval step that stuffs the whole document set into every prompt, or an agent loop with no step limit, will not look different in testing and will look extremely different at the end of the month.

Agentic patterns are the amplifier: one user action can become ten or twenty model calls once the model is choosing its own steps. Multiply that by the number of users and a feature that seemed cheap in a demo becomes the largest line in your infrastructure spend.

The controls are well understood and mostly architectural. Prompt caching, so repeated context is not paid for at full price on every call. Batching for anything not interactive. Routing, so a cheap model handles the easy majority and an expensive one handles the rest. Trimming context to what is actually needed. Hard step limits and per-tenant budgets so a runaway loop stops rather than bills. We put a number on cost per action during the build, and treat it as a requirement like latency.

Security is about permissions, not about filtering

The moment a model can call tools, you have a system that takes instructions from whatever text it reads. A support ticket, a web page, a PDF, a row in your own database can all carry instructions aimed at the model. This is prompt injection, and the important thing to understand is that it cannot be solved by detecting bad text — a model cannot reliably tell instructions from data, because to it they are the same thing.

What contains it is the permission model. Tools scoped to the minimum they need. Read separated from write. Consequential actions requiring explicit confirmation from a human rather than from the model. Every tool call logged with who it was on behalf of, so you can answer afterwards what happened.

This is the same discipline as the rest of our work, which is why we do it: the question is never "can we stop the model being tricked" but "what is the worst thing it is permitted to do if it is". The 2026-07-28 MCP specification pushes in this direction too — servers are now OAuth 2.1 resource servers, with resource indicators so a token issued for one server cannot be replayed against another.

Four shapes an AI feature can take From left to right: ordinary code, a single model call, a fixed workflow with model steps, and a true agent with a loop over tools. Cost and unpredictability increase to the right. { } Ordinary code Deterministic AI One model call One judgement AI AI Fixed workflow Known order, testable AI A true agent Chooses its own path Cheap, predictable, easy to test Expensive, variable, hard to test
Prefer the leftmost shape that does the job. The cost of the rightmost is paid every day it runs.

03Glossary

Terms, defined plainly

What is an AI agent?

A system where a language model decides which actions to take, in what order, using tools you give it, until some stopping condition is met. The distinguishing feature is that the control flow is decided by the model rather than written by you. A chatbot that only produces text is not an agent; a fixed sequence of model calls you wrote is a workflow, not an agent.

What is a tool, or function call?

A function you expose to the model with a name, a description and a typed schema for its arguments. The model does not execute it — it asks for it to be called with particular arguments, your code decides whether to comply, runs it, and returns the result. Every security decision therefore stays on your side of that boundary.

What is MCP?

The Model Context Protocol is an open standard for connecting AI applications to external tools and data, so that a tool you expose once can be used by any client that speaks the protocol instead of being wired separately into each one. The specification was donated to the Linux Foundation’s Agentic AI Foundation, and its 2026-07-28 revision was the largest since launch — most significantly, it defines MCP servers as OAuth 2.1 resource servers.

What is an eval?

A test suite for probabilistic output. Instead of asserting an exact result, it runs a set of real cases and scores the responses — by exact match where possible, by a rubric applied by another model, or by a human review sample — and gives you a number you can compare across changes. Without one you cannot tell whether a prompt edit helped, and you will find out from customers.

What is RAG?

Retrieval-augmented generation: searching your own data for relevant material and putting it into the prompt so the model answers from it rather than from memory. It is the standard way to ground answers in current, private information. It is also over-applied — when the whole corpus fits comfortably in context, retrieval adds a failure mode without adding much.

What is prompt injection?

Content the model reads — a web page, a document, an email, a database row — containing instructions aimed at the model rather than at the user. Because a model cannot reliably distinguish instructions from data, this is not fixable by filtering text. It is contained by limiting what the model is permitted to do: scoped tool permissions, confirmation on consequential actions, and audit logs.

04Choices

Decisions you will have to make

Do you need an agent at all?

Ordinary code
The task is deterministic. Parsing, routing by rules, calculation. A model here adds cost and variance for nothing.
A single model call
One judgement is needed: classify, extract, summarise, rewrite. Cheap, fast, easy to evaluate.
A fixed workflow with model steps
Several judgements in a known order. Testable, debuggable, and covers most real products.
A true agent
The sequence genuinely depends on what earlier steps discover, and the task space is too large to enumerate.

Where we start The simplest of these that does the job. We will argue you down the list, not up it — the cost of the agent shape is paid every day it runs.

RAG, long context, or fine-tuning?

Retrieval (RAG)
The corpus is large, changes often, or is per-customer. The standard answer for grounding in private data.
Long context
The relevant material comfortably fits in the window. Simpler, and with caching often cheaper than it looks.
Fine-tuning
You need a consistent format or style at scale and have the examples. Rarely the answer to a knowledge problem.

Where we start Start with long context plus caching if the material fits, and move to retrieval when it genuinely does not. Retrieval adds a whole failure mode — the right answer not being retrieved — and that failure is silent.

How should tools be shaped?

Many narrow tools
Each action is distinct and you want tight permissions per tool. Easier to audit and constrain.
Fewer coarse tools
The model gets confused choosing between near-identical options, or the call overhead dominates.

Where we start Narrow enough to scope permissions properly, broad enough that the model is not choosing between six similar names. Tool descriptions are prompt engineering and deserve the same care as the prompt.

MCP server, or a bespoke integration?

MCP server
You want your product usable from many AI clients without wiring each one, or customers are asking for it as an integration.
Direct integration
One client, one use case, and no expectation of others. Less machinery.

Where we start MCP when the audience is "whatever assistant the customer uses", which is increasingly the question being asked. Build it as a proper OAuth 2.1 resource server from the start — retrofitting auth onto a server that shipped with a static token is a breaking change for everyone using it.

Which model, and how tied to it should you be?

One model, deeply used
You need a specific capability and the switching cost is acceptable.
Swappable behind an interface
You expect to route by cost, or to move as prices and capabilities change — which they do, quickly.

Where we start Swappable, with the eval set as the thing that tells you whether a swap is safe. Model choice should be a measured decision you can revisit, not an architectural commitment.

05Price

What this costs to build, and what it costs to run

There are two numbers here, and the second is the one people forget. Build cost is quoted after scoping. Running cost is a design output we commit to during the build.

How it is priced Quoted after a scoping call. Where the work is cost reduction on an existing feature, part of it can be priced against the savings we actually measure rather than against the hours.

Agent versus workflow

The largest factor in both numbers. A true agent is harder to build, far harder to test, and multiplies calls per user action.

Number of tools

Each is an interface, a permission decision, a failure path and a set of eval cases.

Whether an eval set exists

Building one from your real cases is real work and the highest-leverage part of the project.

Retrieval

Ingestion, chunking, embedding, indexing and re-ranking, plus keeping it current as your data changes.

Latency requirements

Interactive and streaming is more work than a background job, and constrains which models you can use.

Tenancy and budgets

Per-customer limits, attribution and spend caps are a billing subsystem, not a config value.

Context size per call

The main running-cost lever. Caching and trimming are where most savings are found.

Traffic shape

Bursty interactive traffic prices very differently from work that can be batched at a discount.

06Process

How a project runs

  1. Scoping call

    60 minutes, free

    What the feature should do, what "good" means for it, what data it needs, and — candidly — whether it needs a model at all.

    What we need
    Examples of the task being done well by a person.
    What you get
    A straight answer, including the simplest shape that would work.
  2. Evals first

    Three to five business days

    We build an eval set from your real cases with expected outcomes and a scoring method, and measure a baseline. Often this alone changes the plan.

    What we need
    Access to real examples, anonymised if necessary.
    What you get
    A number for current quality, and a written scope and price.
  3. Build

    Per milestone, agreed in the scope

    The feature, with tracing from the first commit and the eval suite running in CI. Cost per action is measured as we go, not at the end.

    What we need
    Review, and a decision-maker for the judgement calls that surface.
    What you get
    Working feature, eval scores, traces, and a cost-per-action figure.
  4. Handover

    Half a day

    A walkthrough of the evals, the traces, the cost dashboard and the failure modes we know about, plus how to add cases as new ones appear.

    What we need
    An hour of whoever will own it.
    What you get
    A system your team can change without guessing whether they made it worse.

07Deliverables

What you actually end up with

An eval set from your real cases

Your own examples with expected outcomes and a scoring method, plus a measured baseline. Frequently the most valuable artifact of the whole engagement, and it outlives any model you use.

The feature itself

Built in the simplest shape that does the job — ordinary code, one model call, a workflow, or a true agent, in that order of preference.

Regression tests in CI

The eval suite running on every change, failing the build when a prompt edit degrades quality that nothing else would have caught.

Tracing

Every request recoverable: the exact prompt, the context, the tool calls and the response, so a reported bad answer can be diagnosed rather than guessed at.

A cost-per-action figure

Measured during the build and treated as a requirement like latency, with the caching, routing and batching that produced it documented.

Permission and audit design

Scoped tools, read separated from write, human confirmation on consequential actions, and a log of every call with attribution.

An MCP server, where in scope

A proper OAuth 2.1 resource server with protected resource metadata, resource indicators and a published tool schema — not a static token.

08Stack

What we build it with

Hosted models behind an interface you can swap, because capabilities and prices move quickly enough that a hard commitment to one provider is a liability. Postgres with pgvector before a dedicated vector database, unless scale genuinely argues otherwise — one fewer system to operate is worth a lot.

AI Primary here

AI systems

Hosted models behind an interface you can swap, with the eval set deciding which one you use this quarter rather than a contract.

We default to

  • Claude
  • OpenAI
  • MCP
  • Evals
  • Tracing
  • Cost control

We also work in

  • Gemini
  • OpenRouter
  • LiteLLM
  • Vercel AI SDK
  • Anthropic SDK
  • Tool and function calling
  • Structured output
  • Prompt caching
  • Batch APIs
  • Model routing
  • pgvector
  • Qdrant
  • Pinecone
  • Chroma
  • Hybrid retrieval and re-ranking
  • Langfuse
  • LangSmith
  • Braintrust
  • Helicone
  • Ollama
  • vLLM
  • OAuth 2.1 for MCP
API Primary here

Backend and services

Boring, typed and synchronous until something measurably needs to be a queue. Most systems need fewer moving parts than they are given.

We default to

  • Node.js
  • Python
  • REST
  • tRPC
  • Background jobs
  • Queues

We also work in

  • Hono
  • Fastify
  • Express
  • NestJS
  • FastAPI
  • Django
  • GraphQL
  • OpenAPI
  • WebSockets
  • Server-sent events
  • BullMQ
  • Celery
  • Inngest
  • Trigger.dev
  • Cron and scheduled work
  • Stripe
  • Paddle
  • Razorpay
  • Resend
  • Postmark
  • Twilio
  • Webhooks in both directions
Data Primary here

Databases and storage

Postgres unless there is a specific reason otherwise, migrations in the repository, and a restore that has been executed rather than merely configured.

We default to

  • Postgres
  • Supabase
  • Redis
  • Migrations
  • Backups and restore

We also work in

  • Neon
  • PlanetScale
  • MySQL
  • SQLite
  • Turso
  • Firebase Firestore
  • Upstash
  • Drizzle ORM
  • Prisma
  • Kysely
  • pgvector
  • Row-level security
  • Point-in-time recovery
  • Meilisearch
  • Typesense
  • Elasticsearch
  • S3
  • Cloudflare R2
  • Read replicas
  • Connection pooling
Web

Frontend and full-stack

TypeScript everywhere, and a framework that renders on the server unless the product is purely an application behind a login.

We default to

  • TypeScript
  • React
  • Next.js
  • Astro
  • Tailwind CSS
  • Vite

We also work in

  • JavaScript
  • Remix
  • SvelteKit
  • Vue 3
  • Nuxt
  • shadcn/ui
  • Radix UI
  • TanStack Query
  • TanStack Table
  • React Hook Form
  • Zod
  • Zustand
  • Redux Toolkit
  • Framer Motion
  • MDX
  • Storybook
  • Vitest
  • Playwright
  • Testing Library
  • ESLint
  • Prettier
  • Recharts
  • D3
App

Mobile

Expo unless the product is the platform experience. One codebase is worth a great deal when the maintenance bill outlasts the project.

We default to

  • React Native
  • Expo
  • EAS Build
  • Swift
  • Kotlin

We also work in

  • Capacitor
  • Flutter
  • SwiftUI
  • Jetpack Compose
  • Expo Router
  • React Navigation
  • Reanimated
  • MMKV
  • WatermelonDB
  • StoreKit
  • Google Play Billing
  • RevenueCat
  • APNs
  • Firebase Cloud Messaging
  • Expo Notifications
  • Fastlane
  • App Store Connect
  • Google Play Console
  • Detox
  • Maestro
  • Privacy manifests
Ops

Infrastructure and observability

Managed platforms until scale or cost argues otherwise. The parts we insist on are a deploy you can roll back and an error that reaches a human.

We default to

  • Vercel
  • Cloudflare
  • Docker
  • GitHub Actions
  • Sentry

We also work in

  • AWS
  • Cloudflare Workers
  • Cloudflare R2
  • Cloudflare D1
  • Fly.io
  • Railway
  • Render
  • Terraform
  • Nginx
  • Caddy
  • GitLab CI
  • Grafana
  • Prometheus
  • Axiom
  • Better Stack
  • Checkly
  • Uptime checks
  • Structured logging
  • Staged rollouts
  • Tested rollback

Split into what we default to and what we also work in. If your system is outside this, tell us on the call — we would rather say no than learn a stack on your budget.

09Pitfalls

Mistakes we see most

Shipping without evals

Without a scored set of real cases you cannot tell whether a prompt change helped or hurt. You will find out from customers, weeks later, with no way to attribute it.

Putting everything into context

Passing whole tables or document sets into every prompt is the most common cause of a surprising bill, and it usually makes quality worse — relevant material gets buried in irrelevant material.

An agent loop with no stopping condition

No step limit and no cost ceiling means a single bad input can spin until it is noticed. Both are one line each and both are routinely missing.

Treating prompt injection as a text-filtering problem

A model cannot reliably separate instructions from data. Containment is a permissions question: scope the tools, separate read from write, confirm consequential actions with a human.

Static API tokens in an MCP server

A shared long-lived token gives every client identical, unscoped, unattributable access. The current specification defines servers as OAuth 2.1 resource servers precisely to end this, and retrofitting it is a breaking change.

No tracing

When someone reports a wrong answer you need the exact prompt, context, tool calls and response for that request. Reconstructing it later is not possible.

Not telling users they are talking to a machine

Disclosure is increasingly a legal requirement as well as a decent one, and it is trivial to add at build time.

Choosing a model first

Model capabilities and prices move fast. Build behind an interface and let the eval set decide which one you use this quarter.

10Honestly

When you should hire somebody else

  • You want a model trained or fine-tuned on your data as the primary deliverable. That is a different specialism.
  • You want a chat widget on a marketing site. An off-the-shelf product will be cheaper and better than anything we would build.
  • You need a guarantee that the model never gets something wrong. We can measure quality and hold it to a threshold; nobody can promise that.
  • The problem is deterministic. If rules can solve it, rules should, and we will say so on the call.

11Questions

Questions

Something we haven’t covered?

Ask on a call
01 What is an MCP server, and do I need one?

An MCP server exposes your product’s capabilities as tools any AI client can use, rather than you building a separate integration for each assistant. You need one when the answer to "which AI client will our customers use" is "we do not know" — which is increasingly the position. If you have exactly one integration target and no expectation of others, a direct integration is less machinery. Build it as a proper OAuth 2.1 resource server from the start: the 2026-07-28 specification defines them that way, and adding auth to a server that shipped with a shared static token breaks every client already using it.

02 How do I stop my AI feature costing more than it earns?

Treat cost per action as a requirement with a number, the same as latency, and measure it during the build rather than discovering it on a bill. The levers, roughly in order of effect: trim what goes into context; cache repeated context so it is not paid for at full price every call; route easy requests to a cheaper model and hard ones to an expensive one; batch anything that does not need to be interactive; and put hard step limits and per-tenant budgets on anything agentic, because agent loops turn one user action into ten or twenty model calls.

03 Do I actually need an agent?

Probably not, and that is the most useful thing we can tell you. An agent is right when the sequence of steps genuinely depends on what earlier steps discover. If the steps are known in advance — classify, extract, look up, draft — that is a workflow with model calls in it, and it will be cheaper, faster, testable and far easier to debug. We will argue you towards the simplest shape that does the job, because the cost of the agent shape is paid every day it runs.

04 What is an eval, and why do you build it first?

An eval is a test suite for probabilistic output: a set of your real cases with expected outcomes, scored automatically, producing a number you can compare across changes. We build it first because LLM features fail invisibly — slightly less accurate, slightly more verbose, slightly more likely to refuse — and nothing in a normal test suite or error tracker sees any of it. Without one, improving the feature is guesswork, and a prompt change that fixes one case can quietly break five others.

05 How do you handle prompt injection?

By constraining permissions rather than filtering text. A model cannot reliably distinguish instructions from data, so any content it reads — a ticket, a web page, a database row — can carry instructions aimed at it. What contains that: tools scoped to the minimum they need, read separated from write, human confirmation on consequential actions, and audit logging of every tool call and whose behalf it was on. The question we design against is not "can the model be tricked" but "what is the worst thing it is permitted to do if it is".

06 Which model should we use?

Whichever your eval set says is best for your task this quarter, accessed behind an interface you can swap. Capabilities and prices move quickly enough that a hard commitment to one provider is a liability, and routing by difficulty — cheap model for the easy majority, expensive one for the rest — is frequently the single largest cost saving available.

07 Can you reduce the cost of an AI feature we already have?

Yes, and it is a well-defined piece of work: measure where the spend actually goes, then apply context trimming, caching, routing and batching, with the eval set proving quality held. Where the savings are measurable we can price part of the work against them rather than against hours — so if we do not find savings, you do not pay for that part.

08 Do you fine-tune models?

Rarely, and we would usually talk you out of it. Fine-tuning is good for enforcing a consistent format or style at scale when you already have the examples. It is a poor answer to a knowledge problem — for that you want retrieval or a longer context, both of which stay current as your data changes, while a fine-tune is frozen at the moment you ran it.

09 How do you know the feature is working after launch?

Three things, all in place before launch. Evals running in CI so a change that degrades quality fails the build. Tracing, so any individual bad answer can be traced back to the exact prompt, context and tool calls that produced it. And cost-per-action monitoring, because the most common production surprise is not a quality collapse but a bill.

10 Is it safe to let an agent take actions in our system?

It can be, if the permissions are built for it rather than inherited from a human user. In practice: give the agent its own identity rather than a shared token, scope its tools to the narrowest set that does the job, separate read from write, require explicit human confirmation on anything consequential or irreversible, and log every call with attribution. Designed that way, the blast radius of a successful injection is small and, importantly, visible afterwards.

12Also

The other services

Rescue & Hardening

Your AI-built app made safe to launch: access control, payments, deploys and monitoring.

Mobile & Store Launch

Your app packaged, compliant and submitted — through review, not just uploaded.

Web development

New applications and new features, built production-ready from the first commit.

Tell us what you want built.

A scoping call costs nothing and ends with a written scope and a fixed price. If what you need is outside what we do, we will say so on the call.

Prefer to write? Send your requirements

What do you need?

What it should do, who will use it, the must-haves, and any links. A few sentences is plenty.

Budget, if you have one in mind
When would you like to start?