TL;DR: If you are building an AI app that personalises something, a plan, a routine, a programme, do not make the chat window the product. A chatbot gives every user a blank box and a fresh answer each time, which is hard to trust, hard to track and impossible to improve. The better shape is a structured plan the model generates into a fixed format your app owns, a tracking loop that records what the user actually did, and a narrow, explicit path for the model to adjust the plan from that data. Keep the model on one job (drafting and revising the plan), keep the app in charge of state, set hard boundaries on topics where advice can hurt, and measure whether people come back on day seven, not how impressive day one looked.
I co-founded Lifemaxxing AI with Benjamin Siegel. It is an AI self-improvement app: it helps people reimagine where their life could go and then gives them personalised plans and tracking to get there. Before that I co-founded LectureNotes AI, and these days most of my client work is AI voice and chat agents. Across all of it, the same lesson keeps coming back: the model is the easy part. The product is everything that turns a model's output into something a person can follow tomorrow morning.
This post is about that layer, for founders and engineers building any "AI that knows you" product.
Why "just add a chatbot" is the wrong default
When a founder says "it's like a personal coach, powered by AI", the first build almost always ends up as a chat screen. It demos well. Ask a question, get a thoughtful paragraph back.
Then real users arrive, and three problems show up quickly.
The blank box problem. Most people do not know what to ask. A self-improvement app exists precisely because people are unsure what to do next. Handing them an empty text field puts the hardest step, framing the problem, back on the user.
Nothing persists. A chat answer is a paragraph that scrolls away. There is no "plan" to come back to, no checklist, no notion of done. The user has to remember the advice, which is the thing they came to the app to stop doing.
Nothing is measurable. If the output is free text, you cannot tell whether the user followed it, whether it helped, or whether a prompt change made it better or worse. You are flying blind on the one thing your product promises.
Chat has its place. I build chat agents for business websites and they work well when the job is answering questions. But a personalisation product's job is not answering questions. It is producing a plan and helping someone stick to it.
The shape that works: plan, track, adjust
The architecture I trust for this kind of product has three parts, and the model only owns one of them.
1. The plan is structured data the app owns
The model's job is to draft a plan, but the plan itself should be data in a format your app defines: goals, the actions under each goal, how often each action happens, and how progress is measured. The model fills that structure. It does not invent the structure.
This matters for a few reasons:
- The app can render it properly. Checklists, schedules, streaks and reminders all need structured fields. You cannot build a daily view on top of a paragraph.
- You can validate it. If an action has no frequency, or a plan has fifteen goals for a person who said they have twenty minutes a day, your code can reject it and ask the model again before the user ever sees it.
- It survives model changes. If the plan is a schema you own, swapping the model behind it is a config change and a test run, not a rewrite. I made the same argument for voice agents in how I choose the LLM behind a voice agent.
This is the same discipline as function calling in a voice agent: let the model decide what to say, but make it hand over anything the system acts on in a shape the system can check.
2. Tracking is the real product
A plan nobody follows is a nice document. The value of a self-improvement product is in the loop: the user does something, records it in a second or two, and sees that it counted.
Two rules I hold to here:
- Tracking must be cheaper than skipping. One tap to mark something done. If logging feels like homework, people stop logging, and once they stop logging the app has nothing to personalise with.
- The tracking data belongs to the app, not the conversation. Completed actions, missed days and user notes go into your database as records, not into a growing chat history the model has to re-read every time. That keeps cost predictable and makes the data usable for everything else, from reminders to analytics.
3. Adjustment is a narrow, explicit step
Personalisation is not the first plan. Anyone can generate a decent first plan. Personalisation is what happens in week two, when the user has skipped the morning routine five days running and nailed the evening one.
So the adjustment step gets its own deliberate design: the app summarises what actually happened, hands that summary to the model along with the current plan, and asks for specific revisions in the same structured format. The user sees what changed and why, and accepts it.
What I avoid is letting the model silently rewrite the whole plan every time it is called. Plans that change under people feel random, and random destroys trust faster than a mediocre plan does.
Boundaries: what the model must never do
A self-improvement app touches health, fitness, sleep, money, relationships and mood. Some of those are areas where generic AI advice can do real harm. That needs to be designed in, not left to a polite line in the system prompt.
The approach I take on any AI product in sensitive territory:
- Decide the out-of-scope topics up front, in writing, the way I would scope a voice agent for a medical clinic. Diagnosing conditions, prescribing diets for medical issues and anything that sounds like crisis support are not jobs for a plan generator.
- Handle them with a fixed response path, not improvisation. When a user's input lands in one of those areas, the app should respond with a calm, pre-written message that points them to a real professional, rather than letting the model wing it.
- Test the boundaries like features. I keep a fixed set of tricky inputs and run every prompt change against them, the same habit I describe in voice agent regression testing. "It refused nicely once in a demo" is not a test.
- Treat user text as untrusted input. People will paste odd things into a goal field, sometimes on purpose. The lessons from prompt injection in voice agents apply directly: the model should never be able to change its own rules because a user typed something clever.
Cost and context: do not feed the model everything
There is a tempting design where every request sends the model the user's whole history so it "really knows them". It works for a week, then it gets slow and expensive, and the quality often drops because the important signal is buried in noise.
The better pattern is to keep a compact profile and a short summary of recent progress, maintained by your app, and send that. The model gets what it needs to revise a plan, not a diary. I went deep on why every model call has to earn its cost in the unit economics of an AI consumer app, and personalisation products are where that discipline pays off most, because your most engaged users are also the ones with the longest histories.
It also helps to split the jobs. Generating or revising a plan is occasional and worth a capable model. Small things, like phrasing a reminder or tagging a note, run constantly and should use something cheaper, or no model at all.
What to measure
Day one of an AI personalisation app always looks good. The first plan is impressive, people screenshot it, and that is not the number that matters.
The metrics I care about for this kind of product:
- Day-7 and day-30 return rate. Do people come back after the novelty of the first plan wears off?
- Tracking rate. Of the actions in active plans, how many get logged at all, done or skipped? Silence is worse than a logged miss.
- Adjustment acceptance. When the app proposes changes, do users accept them? If they keep rejecting revisions, the adjustment step is not listening.
- Plans abandoned in week one. Usually a sign the first plan was too ambitious, which is a prompt and validation problem you can fix.
None of these are about the model's eloquence. They are about whether the loop works.
What I would tell a founder starting one
If you are building an AI product that promises to know the user, here is the short version I give in discovery calls:
- Design the plan format before you write a single prompt.
- Make tracking one tap, and store it as data, not chat.
- Give adjustment its own step, show the user what changed, and let them say no.
- Write down the topics the app must never advise on, and test them every release.
- Keep the model's context small and the model itself swappable.
- Judge the product on week two, not on the first screenshot.
Building Lifemaxxing AI with a co-founder in the US while I work from Lahore also reinforced the founding-team side of this, which I wrote about in being the technical co-founder with a US partner. The short version: one person owns the product promise, one owns the system that keeps it, and both read the retention numbers every week.
I build AI products, voice and chat agents, and mobile and web apps for founders across the US, UK and Europe, including my own AI apps with real users. If you are building something that turns a model into a personalised experience, more about me is here, or book a call and we can sketch the plan, track and adjust loop for your product.