01 What is an MCP server, and do I need one?
An MCP server exposes your product’s capabilities as tools any AI client can use, rather than you building a separate integration for each assistant. You need one when the answer to "which AI client will our customers use" is "we do not know" — which is increasingly the position. If you have exactly one integration target and no expectation of others, a direct integration is less machinery. Build it as a proper OAuth 2.1 resource server from the start: the 2026-07-28 specification defines them that way, and adding auth to a server that shipped with a shared static token breaks every client already using it.
02 How do I stop my AI feature costing more than it earns?
Treat cost per action as a requirement with a number, the same as latency, and measure it during the build rather than discovering it on a bill. The levers, roughly in order of effect: trim what goes into context; cache repeated context so it is not paid for at full price every call; route easy requests to a cheaper model and hard ones to an expensive one; batch anything that does not need to be interactive; and put hard step limits and per-tenant budgets on anything agentic, because agent loops turn one user action into ten or twenty model calls.
03 Do I actually need an agent?
Probably not, and that is the most useful thing we can tell you. An agent is right when the sequence of steps genuinely depends on what earlier steps discover. If the steps are known in advance — classify, extract, look up, draft — that is a workflow with model calls in it, and it will be cheaper, faster, testable and far easier to debug. We will argue you towards the simplest shape that does the job, because the cost of the agent shape is paid every day it runs.
04 What is an eval, and why do you build it first?
An eval is a test suite for probabilistic output: a set of your real cases with expected outcomes, scored automatically, producing a number you can compare across changes. We build it first because LLM features fail invisibly — slightly less accurate, slightly more verbose, slightly more likely to refuse — and nothing in a normal test suite or error tracker sees any of it. Without one, improving the feature is guesswork, and a prompt change that fixes one case can quietly break five others.
05 How do you handle prompt injection?
By constraining permissions rather than filtering text. A model cannot reliably distinguish instructions from data, so any content it reads — a ticket, a web page, a database row — can carry instructions aimed at it. What contains that: tools scoped to the minimum they need, read separated from write, human confirmation on consequential actions, and audit logging of every tool call and whose behalf it was on. The question we design against is not "can the model be tricked" but "what is the worst thing it is permitted to do if it is".
06 Which model should we use?
Whichever your eval set says is best for your task this quarter, accessed behind an interface you can swap. Capabilities and prices move quickly enough that a hard commitment to one provider is a liability, and routing by difficulty — cheap model for the easy majority, expensive one for the rest — is frequently the single largest cost saving available.
07 Can you reduce the cost of an AI feature we already have?
Yes, and it is a well-defined piece of work: measure where the spend actually goes, then apply context trimming, caching, routing and batching, with the eval set proving quality held. Where the savings are measurable we can price part of the work against them rather than against hours — so if we do not find savings, you do not pay for that part.
08 Do you fine-tune models?
Rarely, and we would usually talk you out of it. Fine-tuning is good for enforcing a consistent format or style at scale when you already have the examples. It is a poor answer to a knowledge problem — for that you want retrieval or a longer context, both of which stay current as your data changes, while a fine-tune is frozen at the moment you ran it.
09 How do you know the feature is working after launch?
Three things, all in place before launch. Evals running in CI so a change that degrades quality fails the build. Tracing, so any individual bad answer can be traced back to the exact prompt, context and tool calls that produced it. And cost-per-action monitoring, because the most common production surprise is not a quality collapse but a bill.
10 Is it safe to let an agent take actions in our system?
It can be, if the permissions are built for it rather than inherited from a human user. In practice: give the agent its own identity rather than a shared token, scope its tools to the narrowest set that does the job, separate read from write, require explicit human confirmation on anything consequential or irreversible, and log every call with attribution. Designed that way, the blast radius of a successful injection is small and, importantly, visible afterwards.