Adding AI to a product or workflow you already have is different from building an AI product from scratch. You have real users, real data and systems that already work. The goal isn't to bolt a chat box onto everything. It's to make one existing job faster, cheaper or better, without breaking what's there.
The steps below work whether the AI faces customers, like a smart search or a drafting feature, or sits inside the business, like sorting invoices or summarizing sales calls.
Pick one job, not "AI" in general
Projects that start with "we should use AI" tend to stall. Projects that start with a specific, repeated task tend to ship. Describe the job in one sentence, with an input and an output.
- "Turn each incoming supplier email into a structured order line."
- "Suggest three tags for every new support ticket."
- "Draft a first reply to a sales inquiry using our product sheet."
- "Let users search our docs by asking a question in plain words."
Good first candidates happen often, follow a pattern a person could explain, and have a clear way to tell a good result from a bad one. Avoid jobs where one mistake is very expensive until you've done a few lower-risk ones.
Check the data before you write code
A model can only work with what you give it. Before building, look at the actual inputs and the information the model would need to do the job well.
- Collect 30 to 50 real examples of the input, including messy and unusual ones.
- Write down what a correct output looks like for each, or have the person who does the job today do it.
- List the extra information a person would look up while doing the job: product data, policies, past records.
- Check where that information lives, whether it's current, and whether it can be accessed through an API or export.
- Flag anything sensitive, so you can decide how it's handled before it goes anywhere near a model.
This step often reveals the real problem. If the documents a person relies on are outdated or scattered, fixing that is part of the project. For how to handle sensitive data, see our guide to using LLMs with company data.
Prototype quickly, outside the product
Before touching your codebase, test the idea with a simple script or notebook that runs your examples through a model and saves the results. This takes days, not weeks, and answers the main question early: can a model do this job well enough?
Keep the prototype honest. Use your real examples, not ones written to make the demo look good. Try two or three prompts and at least two models, including a smaller, cheaper one. Many tasks don't need the largest model available.
Build an evaluation set
The examples you collected become your evaluation set: a fixed list of inputs with expected outputs that you run every time something changes. Without one, every tweak to a prompt or model is a guess.
input: "Hi, can you send 40 units of SKU-1182 to the north warehouse by Friday?"
expected: { sku: "SKU-1182", quantity: 40, location: "North", due: "Friday" }
check: exact match on sku and quantity, location matches a known warehouse- Score what you can automatically: exact fields, valid formats, correct categories.
- For open-ended outputs like drafts or summaries, have a person grade a sample against a short rubric.
- Add every real failure you find later to the set, so it can't quietly come back.
- Decide on a pass bar before launch, for example the share of examples that must be fully correct.
Add guardrails around the model
Treat the model as a component that's usually right and sometimes confidently wrong. The code around it should catch the wrong cases.
- Ask for structured output and validate it. Reject or retry responses that don't match the expected format.
- Check outputs against your own data, such as whether a product ID actually exists.
- Give the model the smallest set of permissions and tools it needs. Reading data is much safer than changing it.
- Treat any text from users or outside sources as untrusted, since it can contain instructions aimed at the model.
- Set timeouts and a fallback, so the product still works when the AI provider is slow or down.
- Log inputs, outputs and decisions, with sensitive fields handled according to your data policy.
Roll out in stages
Even a feature that passes the evaluation set will meet cases you didn't expect. Release it so that early mistakes are cheap.
- Shadow mode: the AI runs on real inputs, but nobody sees the results except your team.
- Suggestion mode: a person sees the AI's output and accepts, edits or rejects it.
- Limited release: automatic for a small group of users or a narrow set of cases, behind a feature flag you can switch off.
- Wider release, one segment at a time, as the numbers hold up.
Suggestion mode is useful data in itself. How often people accept the output unchanged is one of the clearest signals of quality you can get.
Measure cost and quality from day one
AI features have a running cost that grows with usage, and quality can drift when you change prompts, models or data. Track both from the start, not after the first surprising bill.
- Cost per task, not just total spend: tokens used, number of model calls and retries per job.
- Quality: evaluation set score, acceptance rate of suggestions and the rate of corrections or complaints.
- Speed: how long users wait, and whether it changes how they use the feature.
- Business effect: time saved per task, or the outcome the feature was meant to improve.
Costs usually come down through simple changes: shorter prompts, sending only the relevant data, caching repeated work and using a smaller model for easier cases. Re-run the evaluation set after each change to make sure quality held.
Deeraf integrates AI into existing products and workflows as a Build Sprint, from evaluation set to staged rollout, on infrastructure you control, and keeps cost and quality in check through On Call.