TL;DR: In a normal app, a free user costs you almost nothing. In an AI app, every free user who actually uses the product costs you real money, every time, because each core action is a paid call to a transcription or language model. That flips the usual consumer playbook. Before you scale reach, work out what one core action costs you, what your heaviest users cost you per month, and where the paywall sits relative to that cost. Design the free tier around the expensive step, not around a feature list, cap usage in units the user understands, cache and reuse anything you already paid for, and keep the model layer swappable so a price change is a config edit. Viral growth on a product whose unit economics are upside down just loses money faster.
I learned this from my own product, not a client's. LectureNotes AI, the note-taking app I co-founded with Benjamin Rhodes, reached around 15 million organic Instagram views in two months and grew to nearly 10,000 users at roughly $1k MRR. I wrote the growth side of that story separately. This post is the other side: what it means to run a product where the thing people love is also the thing that costs you money every time they do it.
Why AI apps break the "free users are free" assumption
The classic consumer software model rests on one quiet fact: the marginal cost of a user is close to zero. A to-do app or a photo filter runs on the user's phone. Your server cost per extra user rounds to nothing, so you can give the product away to millions, convert a few percent, and the maths works.
AI products do not get that fact for free. When a student records a lecture in LectureNotes AI, the app has to turn that audio into text and then turn the text into a summary, takeaways and an outline. Both steps run on models somebody charges for. That cost is not a fixed server bill you amortise across the user base. It is paid per action, and it scales with exactly the behaviour you want: more use, more cost.
So the variable that matters is not "how many users do we have". It is "how much do our users use the expensive thing, and how much of that usage do we get paid for".
Start with the cost of one core action
Before any pricing conversation, I want one number written down: what does the core action cost, end to end?
For a lecture app, the core action is "one recorded lecture becomes usable notes". The honest cost includes:
- Transcription, which scales with audio length. A lecture is long audio, often an hour or more, not a two-line chat prompt. I covered why lecture audio and phone audio behave so differently in what a lecture recorder and a phone agent taught me about speech-to-text.
- Summarisation and structuring, which scales with transcript length, plus any extra passes you run for outlines or key points.
- Retries and failures, because jobs that fail get re-run and you pay for the failed attempt too. Resumable steps and rejecting unusable audio (silence, a phone in a bag) before it reaches a model are pure margin.
The same exercise applies to voice agents, where I spend most of my working time now. A call has a per-minute cost made of telephony, speech recognition, the model and the synthesised voice. I broke that down for businesses buying one in how much an AI receptionist costs. The principle is identical: know the cost of the unit the customer experiences, not just your monthly invoice.
Then look at the heaviest users, not the average
Averages lie in AI apps. Most free users try the product a couple of times and drift off, and a small group uses it constantly: in a lecture app, the student who records every class through exam season.
That student is your best advocate and, on a free plan, your most expensive cost. If your cost model only looks at the average user, the heavy tail quietly eats the margin. So I model three people, not one:
- The tourist, who tries the core action once or twice. Their cost is your acquisition cost for a chance at conversion.
- The regular, who uses it a few times a week. This is the person your pricing should be built around.
- The power user, who uses it every day. Their monthly cost should be comfortably below what they pay you, or they should be hitting a limit.
If the power user on the free plan costs more per month than your paid plan charges, you do not have a pricing problem. You have a free tier problem.
Design the free tier around the expensive step
The mistake I see most often in AI products, and one I think about a lot after LectureNotes AI, is designing the free tier as a feature list. "Free users get summaries, paid users get outlines and export." That divides by features, but cost does not follow features. It follows usage of the expensive step.
A better question is: what is the minimum amount of the expensive step a new user needs to feel the value? For a lecture app, that is roughly one good set of notes from a real lecture. That moment is what makes someone want to pay. Give them that, generously and quickly, and then put the limit on volume, not on the magic.
A few rules I now hold to:
- Limit in units the user understands. "Three lectures a week" or "60 minutes of recording" is honest and predictable. "500 credits" makes people do arithmetic and feel tricked.
- Put the paywall where the value is felt, not before it. The growth retrospective made the same point from the conversion side: a paywall before the first good result kills conversion, and no paywall at all lets your heaviest free users run up the bill indefinitely.
- Make the limit visible before it bites. A student who discovers the cap mid-exam-week feels ambushed. One who saw "1 lecture left this week" two days earlier makes a decision.
Engineering choices that move the number
Pricing is half of it. The other half is engineering.
Never pay for the same thing twice
If a user re-opens a lecture, regenerating the summary is wasted money. Store the output, and only re-run the model when something actually changed. The same goes for any expensive intermediate result, like a transcript. It is the cheapest optimisation there is, and it makes the app feel faster.
Match the model to the step
Not every step needs the most capable model. Turning a transcript into a clean structure is a different job from answering an open-ended question about it. Use the cheaper model where it is good enough, and spend on the expensive one only where users would notice the difference. Then check that claim with real outputs, the way I check prompt changes against a fixed scenario set in voice agent regression testing. "The cheaper model is fine" is a hypothesis until you have compared the outputs.
Keep the model layer swappable
Model prices change, in both directions, and new models appear constantly. If your provider is wired through the whole codebase, every price change becomes a project. If it sits behind one thin layer with the model and settings in config, a change is an edit and a test run. I argued the same thing for voice platforms in what you would rebuild if your platform doubled its price, and it matters even more for a consumer app with thin margins.
Measure cost the way you measure growth
Most early teams watch views, signups and revenue. Fewer watch model cost per active user, and that is the number that tells you whether growth is helping or hurting.
The metrics I would put on the same dashboard as the growth numbers:
- Cost per core action, tracked over time, because model and provider changes move it under you.
- Cost per active user per month, split by free and paid.
- Gross margin per paying user, after model costs, not before.
- Share of total cost from free users, which tells you whether your free tier is a funnel or a subsidy.
This is the same lesson I keep repeating about voice agents in the metrics that actually matter: measure the thing that decides the outcome, not the thing that is easiest to count. Views are easy to count. Margin per user decides whether you can keep going.
What viral growth does to an AI product
This is the part I would tell any founder about to celebrate a spike. In a classic app, a viral moment is almost pure upside. In an AI app, a viral moment brings a wave of tourists who each trigger the expensive step at least once, and some of them become heavy free users.
That is still worth having; it validated demand for LectureNotes AI fast. But it means a few things have to be true before the spike, not after:
- Your free tier caps the expensive step, so the cost of a spike is bounded.
- You know your cost per core action, so you can estimate the bill from the signup chart.
- Your paywall is in place and tested, so the spike converts while attention is high.
- You have provider spend limits and alerts set, so a bug or an abusive account cannot run up an unbounded bill overnight.
Reach multiplies whatever unit economics you already have. If each user makes you money, reach is wonderful. If each user costs you money, reach is a faster way to find out.
The honest summary
AI products are real businesses with a cost of goods sold, and the cost sits inside the core action. That is not a reason to avoid building them. It is a reason to do the arithmetic first: price one action, model your heaviest user, put the limit on the expensive step, and engineer so you never pay for the same output twice.
When I work with founders now, whether I am the technical co-founder on a product or building something for a client, this is one of the first conversations we have, before a line of code, because it is far cheaper to get right on a whiteboard than on a monthly invoice.
I build AI products, voice agents and mobile and web apps for founders across the US, UK and Europe, and I have run the numbers on my own AI products with real users. If you are building something where every user triggers a model call, more about me is here, or book a call and we can work out your cost per user before you scale.