NNabeel Hassan

Blog · October 3, 2026 · 8 min read

Which LLM Should Power Your AI Voice Agent? How I Choose the Model Behind a Retell Agent

By Nabeel Hassan — AI Engineer · ICPC World Finalist

TL;DR: The language model behind an AI voice agent matters less than most people think and in different ways than they expect. On a phone call the model is judged on three things: how fast it produces the first words, how reliably it follows instructions and calls tools when a caller wanders, and how it behaves when it does not know. Raw intelligence comes a distant fourth. My process is to start with a fast mid-tier model, keep each node's job small enough that a fast model can do it, test candidates against the same recorded scenarios instead of a demo call, upgrade only the nodes that actually need more reasoning, and never let the model choice become the fix for a problem that lives in the prompt, the tools or the call flow.

Almost every voice agent discovery call I have had includes some version of "which AI does it use?" The client has usually read that one model is the smartest this month and assumes the agent will be better if we pick it. More often the smartest model makes the agent slower, more talkative and no more accurate on the five things the phone line actually needs to do.

I build voice and chat agents on Retell AI, and the platform lets you pick the model behind an agent, and on conversation flows, override it per node. At Tested Media on CallSetter AI and at Fortell AI on lines for hospitals, real estate agents, vet clinics and car garages, I have swapped models on live agents more times than I can count. This is what I have learned about making that choice deliberately.

What a phone call actually asks of a model

A chat interface forgives a lot. The user sees a typing indicator, reads at their own pace and can scroll back. A phone call forgives almost nothing. The caller hears silence as confusion, hears a long answer as a lecture and hangs up on both.

So the questions I ask about a candidate model are specific to speech:

None of those are leaderboard benchmarks. You only see them in your own call flow, on real calls.

Why I start with a fast mid-tier model

My default for a new agent is the fastest model that reliably follows instructions and calls tools, not the most capable one available. Reasoning ability is rarely the bottleneck on a receptionist line. Booking an appointment, answering a question from a knowledge base, capturing a caller's details and handing off to a human are all narrow tasks once the call flow is designed properly.

The trick is in that last clause. A fast model struggles with a giant prompt that tries to describe the whole business, every edge case and every tool at once. The same model does well when each node has one job, a short instruction and only the tools it needs. That is the main reason I lean on conversation flows with components for anything beyond a simple line: smaller nodes let a cheaper, faster model do work that would otherwise need a bigger one.

So the first lever is never the model dropdown. It is making the job small enough that the model choice stops mattering much.

When a bigger model earns its place

There are nodes where I do reach for a stronger model, and they share a pattern: the agent has to reason over several facts at once, or the cost of a subtle mistake is high.

Triage with real consequences

On the hospital and medical clinic work I wrote about in voice agents for hospitals, the node that decides whether a caller's description sounds urgent and needs a human now is not a place to save a few hundred milliseconds. It is a short, rule-heavy decision where a stronger model following the escalation rules more consistently is worth the extra latency. The answer is spoken once, so the delay is felt once.

Messy multi-part requests

A garage caller who wants to know if their car is ready, whether the quote included the tyres, and if they can move the collection to Saturday is three intents in one breath. A stronger model untangles that better. Even so, I would rather design the node to take one intent at a time than slow the whole agent for one type of call.

Summaries and extraction after the call

Post-call work is where model speed matters least, because no one is waiting on the line. Summarising the call, extracting structured fields for GoHighLevel and deciding which follow-up to trigger can run on a more capable model without any cost to the caller's experience. In post-call automation the risk is a wrong field written into the CRM, so accuracy beats speed there.

The pattern is per node, not per agent. One agent can run a fast model on the greeting and the booking steps and a stronger one on a single triage decision.

How I actually test candidate models

The wrong way to choose a model is to make three demo calls, pick the one that sounded nicest and ship it. Demo calls are polite, on topic and recorded in a quiet room. Real calls are none of those things.

What I do instead:

  1. Keep a fixed scenario set. The same list of test calls I use for regression testing: the happy path, the caller who interrupts, the one who asks something the knowledge base does not cover, the one who wants a human, the one who gives a wrong date and corrects it.
  2. Run every candidate model through the same scenarios. Same prompts, same tools, same voice. Only the model changes, so the difference is attributable.
  3. Score behaviours, not vibes. Did it call the tool? Did it invent a result when the tool returned nothing? Did it stay inside the node's rules? How many words did it use? How long was the slowest turn, not the average?
  4. Listen to the worst call, not the best. A model that is great on nine scenarios and makes something up on the tenth is worse for a clinic than one that is fine on all ten.

At Fortell I built a pipeline with Claude Code and Comet browser automation to build and test agents quickly, which I wrote up in building and testing voice agents with Claude Code and Comet. The same idea applies to model choice: once running a scenario set is cheap, comparing models is an afternoon of evidence instead of a debate.

Model swaps that should have been prompt fixes

The most common mistake I see, and have made, is reaching for a new model to fix a problem that the model did not cause.

A model swap feels productive because it is one click. But if the underlying problem is structural, the new model hides it for a week and then the same failure shows up in a different form. I now treat "let's try a smarter model" as the last resort in a debugging session, not the first.

Do not let the model choice lock you in

Models change every few months. The one that is best for your agent today will be superseded, repriced or deprecated. That is fine if the agent is built so the model is a setting, and painful if the prompts are tuned to one model's quirks so tightly that a swap breaks everything.

Two habits help. First, keep prompts plain and explicit rather than relying on tricks that happen to work with one model. Second, keep the scenario set alive, so that when a new model appears you can test it against the same calls in an afternoon and switch on evidence. This is the model-level version of the argument I made in voice agent platform lock-in: own the parts that carry your business logic, and treat the parts you rent as replaceable.

It is also the same lesson I took from the consumer side with LectureNotes AI, where I wrote about keeping the model layer swappable for cost reasons.

The short version

If you are choosing the model for a voice agent, here is the order I would work in:

The model is the part of a voice agent that gets the most attention and usually deserves the least. The agents that work well on real phone lines are not the ones on the newest model. They are the ones where the job was made small and the behaviour was tested on calls that look like the real thing.


I build production voice and chat agents on Retell, wired into n8n, GoHighLevel and Twilio, and I pick the model on evidence rather than headlines. If you want an agent that holds up on real calls, more about my background here, or book a call.

FAQ

What is the best LLM for an AI voice agent?

Usually the fastest model that reliably follows instructions and calls tools, not the most capable one. On a phone call, time to first word, staying within the node's rules, correct tool use and admitting uncertainty matter more than raw reasoning ability.

Does a smarter model make a voice agent better?

Not by default. Bigger models are often slower and more talkative, which hurts on a phone line. Most voice agent problems come from the prompt, the tool return values or the call flow, and a model swap tends to hide those problems rather than fix them.

Can different parts of one voice agent use different models?

Yes. On Retell conversation flows you can set the model per node, so a fast model can handle the greeting and booking while a stronger one handles a single high-stakes decision such as medical triage. Post-call summaries and data extraction can use a more capable model because no caller is waiting.

How should you test which model to use for a voice agent?

Run every candidate through the same fixed set of test calls with the same prompts, tools and voice, and score behaviours: whether it called the right tool, invented results, stayed in scope, how many words it used and how slow its worst turn was. Judge by the worst call, not the best.

Building something in this space?

I take on AI-agent, automation and product work directly — scoped fast, shipped fast.

Book a discovery call →

Keep reading