TL;DR: I use Claude Code and Cursor every working day, across Unity and C#, React Native, Next.js, n8n and Retell voice agents. They earn their place on work that is structured, checkable and cheap to regenerate: scaffolding from a written spec, config that lives in files, reading an unfamiliar codebase, and writing the tests I would otherwise skip. I keep them on a short leash, or out entirely, wherever the output is hard to verify: native platform code where they invent APIs with total confidence, Unity scene and prefab files, anything touching production phone numbers, money or patient data, and every decision that is expensive to reverse. The rule underneath all of it: the tool writes, I own. Clients pay for the judgement, and the speed is only worth having if the review keeps up.
People ask me this in almost every discovery call now, usually phrased as "do you use AI to build this?" The honest answer is yes, heavily, and the more useful answer is where and where not. That split is what this post is about.
Why my stack makes this question harder than usual
Most "how I use AI coding tools" posts come from someone working in one web framework all day. My week is less tidy. In the same few days I might touch a Unity project for the RAQTS sports-tech platform, a React/TypeScript/GraphQL dashboard for Bettershop, and an n8n workflow behind a Retell voice agent for Tested Media's CallSetter AI. Before that came a C# SDK running on Windows, Android and iOS at Geonode, and AR work at ARCortex.
That spread matters because AI coding tools are not equally good everywhere. They are excellent where there is a huge amount of public, consistent code to learn from and a fast way to check the result. They get noticeably worse as you move toward niche platform APIs, binary-ish editor formats and systems where the only real test is a device in someone's hand. A single rule like "always use it" or "never trust it" is wrong for a stack like mine. The answer changes by layer.
Where it earns its place
Turning a written spec into a wired-together first draft
The clearest win I have had is the pipeline I built at Fortell AI to build and test voice agents in minutes instead of days, which I wrote up in full in building and testing voice agents with Claude Code and Comet. The important part was not the prompt. It was the spec. Once the ownership line, the fields to capture, the tools the agent can call and the escalation rules were written down plainly, Claude Code could generate the function definitions, the n8n webhook skeletons and the test scenarios in one pass.
The pattern generalises: the better the written spec, the more of the boring structure the tool can produce correctly. A vague request gets a plausible-looking guess. A precise one gets something close to what I would have typed, faster.
Config that lives in files
Anything that is declarative and checked into git is a good fit. Voice agent config exported to files, workflow definitions, CI pipeline YAML, environment templates, type definitions generated from a GraphQL schema. These have a clear shape, a clear diff and usually a validator. When I argued in avoiding voice agent platform lock-in that agent config should live outside the vendor's dashboard, this is part of the reason: config in files is config an AI tool can read, change and explain, and config a reviewer can diff.
Reading someone else's codebase
As a fractional CTO I regularly walk into code I did not write, which I described in stepping in as CTO mid-build at RAQTS. The first days used to be spent tracing data flows by hand. Now I ask the tool to map them: where a value is written, which screens read it, what calls this endpoint. I still verify the map against the code, because a confident wrong answer about architecture is worse than no answer. But as a way to get oriented it is the single biggest time saver I have found, and it is the least risky use because nothing it says ships.
The tests I would otherwise skip
Unit tests on pure logic, fixtures, edge case inputs, regression scenarios for a voice agent. Tests are a great fit because a wrong test usually fails loudly, and the discipline in regression testing a live voice agent depends on having a frozen set of scenarios that someone actually wrote down. An AI tool makes writing that set cheap enough that it actually gets written.
Where I keep it on a short leash
Native platform code
This is where I have seen the most confident nonsense. On the React Native VPN side of the Geonode work, the real behaviour lives in Android's VpnService and iOS Network Extensions, with entitlements, separate processes and background rules that differ per OS and per version. AI tools will happily produce a method that does not exist, a permission flow that was deprecated two versions ago, or a lifecycle assumption that only fails on a real device after the app has been backgrounded for ten minutes.
I still use the tool here, but only as a fast reader of documentation I then check myself, never as the author of code I accept on trust. The same goes for the cross-platform networking core I wrote about in the Repocket C# SDK: the shared logic is fine to generate against, the per-platform edges are not.
Unity scenes, prefabs and anything the editor owns
Unity serialises scenes and prefabs as YAML, which looks like text an AI can edit. In practice those files are full of object IDs and references that the editor manages, and a hand edit that is syntactically fine can quietly break a reference you only notice at runtime. I let the tool write C# scripts, editor tooling and build scripts, which I covered in the Unity CI/CD pipeline post. I do not let it edit what the editor owns. That is not a limitation of any one model. It is a question of who owns the file format.
Visual and spatial work
In AR and XR, "does it work" often means "does it line up with the real world when you stand there holding the device". No coding tool can see that. The geospatial AR coordinate problem is a good example: the maths can be generated, but whether an anchor drifts half a metre on a real phone is a field test, not a code review.
Where I keep it out entirely
Some things I do not delegate, regardless of how good the tools get:
- Production phone numbers, SMS sending and anything that talks to real customers. A misconfigured forward or a test message to a real lead is not something you can take back.
- Money and patient data. Payment flows, pricing logic, and anything near the clinical work I did at Nystag or the hospital agents at Fortell. The review burden there is total, so the time saved by generating is close to zero.
- Secrets. Client API keys, keystores and credentials never go into a prompt. Tools read what is in the repo and the terminal, so secrets stay in environment variables and the CI secret store, not in files the tool will open.
- Decisions that are expensive to reverse. Choosing a voice platform, a data model, a CRM as the system of record, the escalation line on a medical agent. An AI tool can list options. The choice, and the responsibility for it, is mine. That is most of what a client is paying a CTO for, which I wrote about in holding the CTO title at three startups.
The working rules I actually follow
- Spec first, prompt second. If I cannot write down what done looks like, the tool cannot either.
- Every generated change is reviewed as if a new junior engineer wrote it. Read the diff, run it, check the edges. If I would not merge it from a person, I do not merge it from a model.
- Keep the loop small. One focused change, verified, then the next. Large unreviewed batches are where confident mistakes hide.
- Give the tool the project's rules in the repo. A short context file with the stack, conventions and things never to touch saves a lot of wrong guesses, and it doubles as onboarding for the next human.
- Verify on the real target. A device, a real test call, a staging deploy. "It compiles" is not the same as "it works", and the gap is widest exactly where the tools are weakest.
- Treat prompts and agent instructions as untrusted input when they come from outside. The same thinking as voice agent prompt injection: a coding agent that reads web pages or issue text can be told to do things you did not ask for, so permissions stay tight.
What clients actually get out of it
The honest version: AI coding tools make me faster on the parts of a project that were never the hard part. Wiring, boilerplate, config, test scaffolding, reading unfamiliar code. That time goes back into what decides whether the product works: the architecture, the edge cases, the platform-specific details and the conversations with the client about what they really need.
What it does not do is replace the person who knows why a Unity reference broke, why an iOS extension got killed in the background, or why a voice agent should transfer a call instead of answering it. If anything, the tools raise the value of that knowledge, because they produce a lot of plausible output and someone has to know which parts are wrong.
So yes, I use AI on your project. I also read every line it writes, and I am the one accountable for it.
I build AI voice agents, mobile and web products and Unity/XR software for teams across the US, UK and Europe, using AI tools where they genuinely speed things up and my own judgement everywhere else. If you want a product built fast without the confident mistakes, more about me is here, or book a call.