← All field notes

Beta testing · September 2026 · 7 min

What do testers see that the builder cannot?

A conversation about the very different app you see once somebody else is holding the phone.

A phone held in one hand, ready for someone else to use
A phone ready for someone else to use. Photo by Jakub Żerdzicki on Unsplash

Why isn’t testing your own app enough?

I already know what everything is meant to do, and that changes how I use it. If a label is not very clear, I still understand it because I wrote it. If something takes longer than expected, I know what process is happening in the background. I also know which account has the right test data and which order makes the flow work.

Some of that happens without me realising it. I can use the same build several times and stop seeing a problem because my behaviour has adjusted around it. A tester has not learned those little workarounds. They use what is actually there, which is why their confusion can tell me more than another successful test on my own phone.

What was the first testing problem that surprised you?

Getting the invitation itself became a problem. I had a TestFlight build, an internal group and the tester listed, but the email was not arriving even after I sent the invitation again. I was the account holder and still could not make the basic route into the build feel straightforward.

It had nothing to do with the app code, but for the tester that distinction did not matter. Their experience was that they could not test Tilawa. It made me widen what I considered part of onboarding, because it really begins wherever the user first has to depend on something working.

What changed when you saw the app on different phones?

A lot of details I thought were settled came back. The font on iOS did not feel like the Android version. The splash logo felt too large. I changed the icon and then had to judge whether the artwork was sitting properly inside the rounded shape Apple applies on the home screen. These sound small, but when they all appear together the app can feel less finished than it did on the device I used every day.

There were functional things too. The Makkah and Madinah live streams were not loading in the app, and Sentry showed a font-loading exception. Each issue had its own cause, but the person using Tilawa is not separating them into design, native configuration and network handling. They are deciding whether the app feels dependable.

Did that make you focus more on visual issues?

It made me take them seriously without letting them take over the release. If the logo is badly sized, the user notices before they have done anything else. That affects confidence. At the same time, I still have to compare it with a payment problem, a privacy problem or a call that does not end correctly.

I went through several versions of the splash and icon because I wanted the first impression to feel intentional. I just tried to keep the order sensible. A visual issue deserves care, but something that can lose a user’s money or expose the wrong information cannot sit behind it because the visual fix is more enjoyable.

What is the difference between internal and external testers for you?

Internal testers are useful when I need a quick answer about a build. Did it install, did the fixed flow work and did the new version reach the device? They also know me and usually know something about what I am trying to build, so there is already shared context.

External testers are closer to the experience I eventually care about. They may have a different phone, different expectations and less patience for something that is not clear. They are not carrying the same background story, so they expose places where the product is relying on knowledge that never made it into the interface.

How do you get feedback that is more useful than “I like it”?

I give the tester something real to do. I might ask them to sign in, find a teacher, place a call, leave the app and return, then go through the minutes page. I want to know where they paused, what they expected to happen and what they thought the app had done after an action.

If I only ask what they think, I often get comments about colours or whether the idea is good. Those comments are not useless, but watching somebody hesitate in front of the correct button tells me that the screen is not communicating as well as I assumed. Their behaviour is usually more specific than their opinion.

Do you wait for several people to report the same issue?

It depends on the issue. If one person dislikes some spacing, I can note it and see whether it appears again. If one person finds a way to be charged twice or see something private, I do not need a second report before treating it seriously.

I also try not to fix only the exact example. If a livestream fails on iOS, I need to know whether that one URL is wrong or whether I have a broader problem with how that platform handles the media. The report is where I start investigating, not always the full description of what needs to change.

What has beta testing changed in the way you build?

I explain less and observe more. My first instinct used to be telling somebody what a screen was meant to do, especially when I knew how much work sat behind it. Now I try to let the confusion happen and understand it before I help.

When several people misunderstand the same thing, I have to accept that the product is unclear. The beta period is useful precisely because it gives me time to discover that while the group is still small and I can talk to people properly. I would rather hear the uncomfortable version now than protect my own explanation and meet the same problem after release.