NNabeel Hassan

Blog · September 4, 2026 · 8 min read

The First Two Weeks After a Voice Agent Goes Live

By Nabeel HassanAI Engineer · ICPC World Finalist

TL;DR: Launch day is not the end of the build, it is the start of the only test that counts. Ramp instead of switching: give the agent overflow and after-hours calls first, not the main line. Spend the first days listening to whole calls rather than tuning after every bad one, because the first complaint is a sample of one. Triage in a fixed order, batch the changes, ship them into quiet traffic, and hold the numbers until week two when the denominators mean something. Most of what breaks in week one is not the model. It is hours, transfer destinations and assumptions about who calls.

Everything before go-live is a hypothesis. You wrote refusals for the questions you imagined, you tested with people who already knew what the agent was for, and you agreed a definition of working with a client describing their business from memory. Then the forwarding rule flips and real people start ringing.

I have taken agents live for UK hospitals, vet clinics, estate agents and car garages at Fortell AI, and for a US product at Tested Media where calls land in GoHighLevel and the client watches them arrive in a dashboard. The shape of the first two weeks is consistent, and so is the way it goes wrong when it is treated as a victory lap.

Ramp, do not switch

The single biggest decision is how many calls the agent gets on day one, and the right answer is almost never all of them.

Start with the calls that are currently going nowhere. After-hours is the safest slice, because nobody is losing anything that was previously being handled: those callers heard a voicemail greeting or nothing at all. Overflow is next, where the agent picks up only when the line is engaged or unanswered after a set number of rings. Both give you real callers with real intent while the business keeps its existing front door.

This is a routing decision, not an agent decision, and it belongs in the phone number layer. Keep the main number where it is and point a forwarding rule at the agent. A ramp you can reverse in thirty seconds by editing one forwarding rule is worth more than any amount of pre-launch confidence.

The pressure to skip this is real: a client who has paid for something wants it answering everything on Monday. The argument that works is commercial rather than technical. The calls the agent takes in the first week decide whether the client trusts it in the second, and it is better that it earns that trust on the calls the business was already missing.

Day one is for listening, not tuning

The first bad call arrives faster than you expect and the instinct is to fix it immediately. Resist it for a couple of days.

A single call is a sample of one, and the change it provokes is almost always too specific. Someone asks about parking, so a line about parking goes into the prompt, and now the agent volunteers parking information to people who did not ask. Do that five times in a week and the prompt becomes a pile of patches aimed at individual callers, each one making the agent slightly worse for everybody else.

Collect instead. Every call that ends badly goes into a list with a one line note. After twenty or thirty calls the list sorts itself: three notes saying the same thing is a real problem worth a change, and one note on its own is an anecdote. Anecdotes are how prompts get fat.

The exception is anything actively harmful. Wrong opening hours, a promise the business does not honour, or a missed emergency escalation gets fixed the hour you hear it. Everything else waits for the batch.

Listen to whole calls before you look at a dashboard

In week one the dashboard is misleading. The volumes are too small for a rate to mean anything, and one strange afternoon moves every percentage on the page.

So I listen to whole calls, start to finish, not clips and not transcripts. Transcripts hide the two things that decide whether a call went well: how long the caller waited for the agent to start speaking, and the moment their tone changed. You can read a transcript of a call that ended in a booking and miss entirely that the caller repeated themselves three times to get there. Every call for the first day or two, then a switch to reading the failures rather than a random sample.

Triage in a fixed order

When the list has enough on it to be worth acting on, I sort it into three buckets and work them in this order.

Calls that lost the caller. The agent could not handle something and the caller hung up, or it handled something it should have escalated. Lost revenue and lost trust, so these get fixed first. Usually the fix is an exit rather than a capability: add a fourth path alongside the three handoff exits it already has, instead of teaching the agent to handle the case itself.

Calls that lost the data. The conversation went fine, the caller was happy, and nothing usable came out the other end. A name spelled unrecognisably, a number captured with a digit missing, a booking that never appeared in the calendar. These are quieter and more dangerous, because everyone thought the call went well. Most are capture problems or failures in the layer behind the call, not conversation problems.

Calls that were merely awkward. Slightly wordy, an odd pause, a strange emphasis on a word. Real, worth fixing, and last. There is always more of this than there is time for, and none of it costs the client money.

What actually breaks in week one

After enough launches the same things keep appearing, and almost none of them are the model.

Hours. The agent knows the published opening hours. It does not know about the half day, the bank holiday, the vet who does not work Wednesdays, or the phones being answered from eight even though the door opens at nine. Real callers find every one of these in the first fortnight.

The greeting. Whatever the agent does with the business name should have been settled when you chose the voice, but hear it once on the live line anyway.

The transfer destination. Transfers pass in testing because you were the one on the other end. In production the destination rings out, reaches a mobile in a workshop, or lands in a voicemail box nobody empties. Test the destination, not just the transfer, at the times of day transfers will actually happen.

Callers who are not the caller you designed for. Scoping produces an imagined caller. Reality delivers the delivery driver, the supplier chasing an invoice, the recruiter, the wrong number and a trickle of automated spam calls. None of these need clever handling. They need a short, polite exit that does not spend a minute trying to book them an appointment.

Stale knowledge. Prices that changed, a service that was dropped, a policy the owner rewrote and told nobody. Week one is when the knowledge layer gets audited by strangers, and staleness surfaces as confidently wrong answers rather than as gaps.

Batch, then ship into quiet traffic

When the list has three or four real fixes on it, they go out together, and they go out when the phone is not ringing.

A live agent is global state written in prose, so every prompt change is a deployment to production with no staging environment underneath it. That is the whole argument for a frozen scenario set and a reversible change: run the scenarios before and after, keep the previous version somewhere you can restore it in a minute, and never ship into the busiest hour of the week.

Tell the client what changed, in one line of plain language. Not because they need the detail, but because it turns their next piece of feedback from a complaint into a comparison.

Week two: the numbers finally mean something

By the second week there is enough volume for rates to stop lying, and this is when I put numbers in front of a client for the first time. Held back deliberately, because a bad Tuesday in week one gets quoted back at you for a month.

The number that matters is not how many calls the agent answered. It is what happened to the calls that would otherwise have been missed, measured against a denominator both sides agreed before the build. That agreement should have come out of the scoping call, and if it did not, week two is when you find out you and the client have been measuring different things.

Expect a mixed picture: bookings captured that were previously going to voicemail, calls the agent should not have taken, and a handful that are genuinely embarrassing. Show all three. The clients who stay are the ones who saw the failures early and watched them get shorter.

The client is going live too

The part that gets forgotten is that the humans are new to this too. A handoff only works if somebody picks up, and a captured lead only counts if somebody knows where it lands and looks at it.

So the launch includes a short conversation about their side: where the leads appear, who checks them and how often, what to do when the agent transfers a call, and who to tell when something sounds wrong. In a build that routes everything into a CRM, this is where most of the value gets realised or lost, and no amount of prompt tuning compensates for a lead list nobody opens.

When to widen the ramp

I widen when two weeks have passed with nothing landing in the first triage bucket, when the fixes going out are cosmetic rather than structural, and when the client has stopped forwarding me clips of individual calls. That last one is the real signal. When a client stops auditing every call, the agent has started to feel like infrastructure rather than an experiment, and the main line can move.

The short version

  1. Ramp with after-hours and overflow first, through a forwarding rule you can reverse instantly.
  2. Listen to whole calls for the first days and collect, do not patch after every bad call.
  3. Fix anything actively harmful the same hour, batch everything else.
  4. Triage lost callers first, lost data second, awkwardness last.
  5. Ship batched changes into quiet traffic with a scenario set and a way back.
  6. Hold the numbers until week two, then show the failures alongside the wins.
  7. Launch the humans too: where leads land, who picks up a transfer, who reports problems.

I build production AI voice agents and the automation, CRM and telephony layer behind them for founders across the US, UK and Europe. More about my work here, or book a call.

FAQ

How should an AI voice agent be launched, all at once or gradually?

Gradually, through a forwarding rule rather than a number port. Give the agent after-hours calls first, because nobody loses anything that was previously being handled, then overflow, where it picks up only when the line is engaged or unanswered after a set number of rings. Both deliver real callers with real intent while the business keeps its existing front door, and both can be reversed in thirty seconds by editing one forwarding rule. Widen to the main line once a couple of weeks pass with no calls being lost and the client has stopped forwarding clips of individual calls.

Should I fix a voice agent prompt after every bad call in the first week?

No, apart from anything actively harmful such as wrong opening hours, a promise the business does not honour, or a missed emergency escalation. A single bad call is a sample of one, and the change it provokes is almost always too specific: a caller asks about parking, a line about parking goes in, and now the agent volunteers parking information to everyone. Collect failures in a list with a one line note each, wait until three notes say the same thing, then ship the batch into quiet traffic with a scenario set and a way back.

What usually goes wrong in the first week of an AI receptionist being live?

Rarely the model. It is opening hours the published schedule does not capture, such as half days, bank holidays and staff who do not work certain weekdays. It is transfer destinations that ring out or land in a voicemail box nobody empties, because in testing you were the person on the other end. It is callers nobody designed for, like delivery drivers, suppliers, recruiters and spam calls, which need a short polite exit rather than an appointment. And it is stale knowledge surfacing as confidently wrong answers.

Building something in this space?

I take on AI-agent, automation and product work directly — scoped fast, shipped fast.

Book a discovery call →

Keep reading