TL;DR: You cannot test an augmented reality app at your desk, and the Unity editor will happily tell you everything works. AR bugs live in the space between the code and the world: lighting, surfaces, sensor noise, heat, motion and the person holding the device. After building AR and XR systems at ARCortex (the ERIS platform for the fire service, AR Planes, AR Properties, MQ-9 Reaper tracking on AR maps) and a clinical eye-tracking app at Nystag, the testing setup I trust has four layers: pull every piece of logic that does not need a camera out into plain unit tests, replay recorded sessions so tracking bugs can be reproduced, run a small device lab for the hardware spread, and then do scripted field sessions in the real environment with an on-device debug overlay running. The field session is the one you cannot skip, and the other three exist so that the field session is spent finding new problems instead of rediscovering old ones.
When I moved from Unity games into AR, my testing habits came with me and most of them were wrong. In a game, the world is yours. Every wall, light and floor is an asset you placed, so a bug you see once you can see again. In AR, half the scene belongs to reality, and reality does not take pull requests.
Why the editor lies
Unity gives you ways to fake an AR session in the editor, and they are useful for layout and UI work. They are also the source of most false confidence on an AR project, because a simulated session gets the hard part for free.
In the editor, tracking never drifts. Planes are detected instantly and stay put. The compass points exactly where it should. Frame time is whatever your desktop GPU makes it. The device never gets warm.
On a real phone or headset, every one of those assumptions breaks in a different way:
- Tracking depends on the room. A blank wall, a glossy floor, low light or a crowd walking past and the session loses its sense of where it is. Content that sat perfectly still in the editor starts to slide.
- Sensors are noisy in structured ways. The biggest error in outdoor AR is usually heading, not position, which I went into in the coordinate problem behind geospatial AR. Near a vehicle, a steel structure or a laptop bag, the magnetometer quietly lies.
- Performance changes over time. A phone running camera, tracking and rendering at once gets hot, and a hot phone throttles. The app that held its frame rate for two minutes on your desk can be dropping frames twenty minutes into real use.
- The user is not you. You hold the phone steady at chest height because you know where the content is.
None of that shows up until the app is in a real place in a real hand. So the question is not whether to field test. It is how to make each field session count.
Layer one: get the logic away from the camera
The cheapest AR bug to fix is the one that was never an AR bug. A surprising share of what goes wrong in an AR app is plain logic that happens to live inside a MonoBehaviour: converting a latitude and longitude into a local position, choosing which label to show, interpolating a moving target, deciding when data is too old to display.
On the geospatial work at ARCortex, the conversion from world coordinates into the AR session's local frame is pure math. It takes numbers and returns numbers. If it lives in a plain C# class instead of inside a component that needs a running session, you can test it with known inputs in milliseconds, including the nasty cases: targets kilometres away, altitudes in different datums, a heading right at the 0 and 360 degree boundary.
The rule I follow now is simple. If a function does not need the camera, it should not be able to see the camera. That makes AR code look a bit more like the backend code I write for mobile apps that have to survive old versions, and that is a good thing. The parts that can be tested deterministically get tested deterministically, on CI, on every commit, which is the whole point of having a Unity build pipeline in the first place.
Layer two: record sessions so bugs can be replayed
The worst AR bug report is "the label jumped once, near the car park, around four." You were not there, you cannot get the same light again, and the tester cannot tell you what the tracking state was at that moment.
Both major mobile AR platforms let you record a session and play it back. ARCore has a recording and playback API that captures the camera feed and sensor data, and on iOS you can record ARKit sessions with Apple's tools and replay them in Xcode. But it turns "it happened once outside" into a file you can open on a desk as many times as you need. It will not reproduce thermal throttling or live network data, but it pins down tracking bugs.
Once you have a handful of recordings from real conditions, a dim room, a featureless corridor, an outdoor walk, they become a regression set. When you change anything about placement, smoothing or anchoring, you replay them and watch for content that drifts or jumps more than it used to. It is the same idea I use for regression testing voice agents: freeze a set of real, messy inputs and run every change against them, because the bug you fixed last month is the one most likely to come back.
Layer three: a small device lab, chosen on purpose
You do not need fifty devices. You need the right five or six, picked to cover the ways hardware actually differs for AR:
- The oldest device you officially support. It decides your frame budget, your texture sizes and how much you can render at once.
- A device with depth sensing and one without. Plane detection and occlusion behave very differently, and the fallback path tends to be the least tested code in the app.
- One mid-range Android phone. Android camera and sensor behaviour varies by manufacturer more than people expect, and this is where the odd ones show up.
- The headset, if there is one. At Nystag that was the Vive Focus 3, and a headset is its own platform with its own input, comfort and performance limits. I wrote about how the hardware choice gets made in choosing an XR headset for enterprise deployment.
The device lab is where you run the soak test, the one almost nobody runs: leave the app running its heaviest scene for thirty minutes and watch frame time, temperature and battery. Real users do not stop after two minutes, and a firefighter or a clinician certainly does not.
Layer four: field sessions with a script
Field testing is where the real bugs are, and it is also where time disappears. Walking around with a phone and "seeing how it feels" produces impressions, not findings. What worked for me was treating each field session like a small experiment.
Write the route before you go. Which places, which lighting, which moves: approach the target, walk past it, turn quickly, cover the camera, lock and unlock the phone, take a call in the middle. The interruptions matter as much as the happy path, because recovering a lost session is where AR apps most often fall apart.
Run a debug overlay on the device. A small, toggleable panel that shows tracking state, frame time, the current heading and its accuracy, the age of any live data, and how many anchors exist. Without it, a tester sees that something looks wrong. With it, they can see why: tracking was limited, the heading accuracy dropped, the data was forty seconds old. That overlay has saved me more debugging time than any other tool on an AR project.
Capture everything. Screen recording, the app's own logs with timestamps, and a session recording if the platform supports it. The bug you notice in the field is often not the one you can explain in the field.
Put it in the hands of the people who will use it
The last layer is not technical. Engineers are terrible AR testers because we already know where the content is and how the app expects to be held.
For the ERIS work this meant thinking about the conditions of a real incident: gloves, bright daylight, someone moving fast with their attention on something else entirely. I covered those design constraints in building AR for firefighters, and every one of them is also a testing condition. If the controls only work bare-handed and indoors, the test plan has to say so, or it is testing the wrong thing.
For the clinical app it meant the opposite kind of user. A clinician fitting a headset on a patient cares about setup time, comfort and whether the measurement is repeatable, which is why I wrote in the Nystag build story that clinical-grade means boring before it means impressive. A test session there is less about whether the content looks right and more about whether the same procedure gives the same result twice.
What I would set up on day one of any AR project
If I joined an AR project tomorrow, as I have done a few times in fractional CTO roles, this is the order I would put things in:
- Move coordinate, placement and data logic into plain classes with unit tests on CI.
- Build the on-device debug overlay before the second feature, not after the first field disaster.
- Record a starter set of real sessions in at least three different environments.
- Pick the device lab on purpose, including the oldest supported device, and run a thirty-minute soak test.
- Write a field route with interruptions in it, and repeat it at different times of day.
- Put the app in a real user's hands, in their real conditions, and do not help them.
The lesson that carried over
I build mostly AI agents now, and this habit came with me more than any line of Unity code. A voice agent is tested in the same uncomfortable way: the editor (or the platform's test chat) is clean, and real phone calls are full of background noise, interruptions and people who do not behave like the developer. The fix is the same structure. Test the logic without the messy input, record real sessions and replay them, and then go and watch it fail in the real world on purpose. I wrote about that arc in going from Unity game dev to AI engineer, and testing is a big part of why the move felt natural.
I have built AR and XR systems for public safety, defense and healthcare, and now build production AI agents and products for startups across the US, UK and Europe. If you have an AR or XR app that works on your desk and nowhere else, more about my background is here, or book a call.