TL;DR: A voice agent should not greet a caller who rang yesterday like a stranger. The fix is not a smarter model, it is a lookup: when the call comes in, check the caller's number against the CRM, pass a handful of facts into the agent as variables before it speaks, and write a short summary back to the contact when the call ends so the next call starts with context. On Retell that means an inbound call webhook that returns dynamic variables, backed by a fast query to GoHighLevel or whatever CRM the business runs. Keep the lookup fast with a safe default for unknown or slow lookups, pass the minimum the agent needs, and treat caller ID as a hint rather than proof of identity, because shared phones and spoofed numbers are real. Memory that the agent uses carelessly is worse than no memory at all.
Most of the voice agents I have built, on CallSetter AI at Tested Media and across the hospital, real estate, vet clinic and car garage agents at Fortell AI, talk to the same people more than once. Patients call back about an appointment. A pet owner calls to check on a booking. A driver calls the garage to ask whether their car is ready. A lead who spoke to the agent last week calls again with one more question.
If the agent opens every one of those calls with "Can I get your full name?", the business loses something a good human receptionist gives for free: the feeling of being known. This post is how I give agents that memory without turning it into a privacy problem.
Where the memory actually lives
The first design decision is the one people get wrong most often. The memory does not live in the model, and it should not live in the voice platform's call history either. It lives in the CRM.
The business already has a system of record. On CallSetter that is GoHighLevel, which holds the contacts, the pipeline stages and the appointments. At Fortell the agents plugged into whatever each client used. That system already knows who the caller is, what they booked and what stage they are at. The agent's job is to read a small slice of it at the start of the call and write a small slice back at the end.
This keeps one source of truth. If a receptionist updates an appointment by hand, the agent sees it on the next call. If the agent books something, the staff see it in the tool they already use. I made the same argument about dashboards in the CallSetter build story: the agent is a front door, and the CRM is the building behind it.
The lookup: before the agent says a word
Retell lets you configure an inbound call webhook. When a call arrives on your number, Retell sends your endpoint the caller's number and the number they dialled, and your endpoint can respond with dynamic variables for that call. The agent's prompt references those variables, so the greeting and the first few turns can use them immediately.
The flow I use looks like this:
- Call arrives. Retell hits the webhook with the caller's number.
- Normalise the number. Strip formatting and convert to E.164 so
07700 900123and+447700900123match the same contact. Skipping this step is the most common reason a lookup "randomly" fails. - Query the CRM. Find contacts with that phone number. One query, indexed, no chains of calls.
- Build a small variable set. First name, a status flag, the next appointment if there is one, and a one-line note from the last call.
- Return within a tight budget. If anything is slow or fails, return the default variables for an unknown caller and let the call carry on.
That last point matters more than it looks. The lookup happens before the caller hears anything, so every millisecond of it is dead air or ringing. I wrote about where time goes on a call in the voice agent latency budget, and the pre-call lookup has to fit inside it. A greeting that comes two seconds late costs more than a greeting without a name.
I usually build this endpoint in n8n for the first version, since the rest of the pipeline already lives there (the setup is in connecting Retell AI to n8n). If the lookup gets complicated or n8n's overhead starts eating the budget, it moves into a small dedicated function. The contract with the agent does not change either way: number in, variables out.
What to pass in, and what to leave out
The temptation is to dump the whole contact record into the prompt. Do not. Every extra field is something the agent might say out loud, misread, or mention at the wrong moment.
For most businesses, this is enough:
caller_first_name, only if there is exactly one matching contact.caller_status:new,known, orambiguous. The prompt branches on this, not on whether the name happens to be empty.next_appointment, already formatted for speech. "Thursday at half past two", not2026-10-15T14:30:00Z. I covered why formatting for the ear matters in writing voice agent prompts for speech.last_call_note, one sentence, written by the previous call's summary step.
Vertical by vertical, one more field usually earns its place. At a vet clinic it is the pet's name, because "Is this about Bella?" lands far better than "Is this about your pet?", which I touched on in voice agents for vet clinics. At a garage it is whether there is an open job on the car, since "is my car ready" is the call they get most, as covered in voice agents for auto repair shops. In real estate it is the property they enquired about.
What stays out: payment details, clinical information, internal notes written by staff, and anything the business would be embarrassed to hear read back. If the agent does not need a fact to handle the call, it does not get it.
Caller ID is a hint, not an identity
This is the part that turns a nice feature into a risky one if you skip it.
A phone number is not a person. Families share a landline. A husband calls about his wife's appointment from his own phone. A business line is answered by whoever is at the desk. Caller ID can also be spoofed, which means anyone who knows someone's number can make a call look like it came from them. I went through why the caller should never be treated as trusted input in voice agent security and prompt injection, and caller ID is part of that.
So the rules I write into every prompt that uses caller memory are:
- Confirm before personalising. "Hi, am I speaking with Sarah?" If the answer is no, the agent drops the variables and carries on as if the caller is new.
- Never volunteer sensitive details. The agent can say "I can see you have an appointment with us later this week" after confirming the name. It does not read out what the appointment is for, and in a clinical setting it says even less. For hospital and medical clinic agents, anything beyond the name and a time waits for proper verification or a human.
- Changes need more than caller ID. Cancelling or moving an appointment asks for a second detail the caller would know, such as the date of birth on file or the booking date, checked by a function call on the backend rather than by the model comparing strings in its head.
The principle is that memory should make the call smoother for the right person and give nothing useful to the wrong one.
The messy cases
Real CRMs are not clean, and the lookup has to cope with that.
Duplicate contacts. The same number is on three contacts because a lead form created a new one each time. If the lookup picks one at random, the agent greets the caller by the wrong name. My rule: more than one match means caller_status = ambiguous, no name, and the agent asks. The post-call step can flag the duplicate for staff to merge.
Stale data. The appointment in the CRM was cancelled by text last night and nobody updated it. This is why the agent phrases memory softly ("I can see an appointment on Thursday, is that still right?") instead of stating it as fact, and why booking changes always go through a live function call rather than trusting the pre-call snapshot. The function calling post covers how those calls should be designed.
Withheld numbers. No number, no lookup. The agent behaves exactly as it would with a new caller, and nothing breaks.
Slow CRM. The API has a bad minute. The timeout fires, the defaults go back, and the caller never knows. The worst outcome of a failed lookup should be a slightly less personal call, never a silent one.
Closing the loop: write back after every call
A lookup only gives the next call context if the previous call left some. That is the job of the post-call step.
When Retell sends the call analysis, the workflow writes a short note to the contact: what the caller wanted, what the agent did, and anything left open. Not the transcript, a sentence or two that a human or the next call can use. I wrote about making those notes useful in call summaries staff actually read, and the same summary doubles as the last_call_note for the next lookup.
The write-back is also where a new caller becomes a known one. If the agent captured the caller's details on the first call, the workflow creates or updates the contact with the normalised number, so the second call finds them. If that step fails quietly, the memory never builds up, which is one of the failure patterns in post-call automation failures.
How I test it
Caller memory adds branches, and branches need tests. Before a client goes live I run the same small set every time, and keep it as part of regression testing after:
- A known caller with one contact and an upcoming appointment.
- A known number where the person says "no, this is her husband".
- A number matching two contacts.
- A brand new number, then a second call from it after the first call has written back.
- A withheld number.
- A deliberately slow lookup to prove the default kicks in and the greeting still arrives on time.
If all six behave, the feature is ready. If the second call from a new number still treats them as a stranger, the write-back is broken, and that is far better found in testing than by a client wondering why their "smart" receptionist forgets everyone.
I build production voice and chat agents on Retell, n8n and GoHighLevel for businesses across the US, UK and Europe. If you want an agent that remembers your callers without saying too much, more about my work is here, or book a call.