Case study
Exam Readiness
Summary
As sole product designer, I helped shape the strategy and led the design for Heylama’s pivot to B1 exam prep. The product tracked practice, but couldn’t answer what mattered most: am I ready to pass?
Team
Status
Product strategy.
Unshipped.
The stakes
You moved to Berlin. Now everything depends on one exam.
You did the hard part. New country, new language, a life you are building here. Then a letter arrives: your residency needs a B1 certificate. A date, a fee, a result.
Pass, and your life here continues. Fail, and you wait. Months, maybe. You pay again. You keep waiting for permission to stay where you already live.
It was never about the price of the test. It was about the price of living where you chose to live.
The problem
We couldn't beat the giants. So we stopped playing their game.
Heylama was a small team with a small budget in the most crowded category in software: general language learning. Against Duolingo money, we would lose.
So we narrowed. Not to a smaller version of the same thing, but to the one place the giants weren't built for: people with a real exam, a real deadline, and a real consequence riding on it.
This wasn't a hypothesis. Our CEO had watched friends in Berlin live it, visas and jobs waiting on a language certificate.
What we found
These learners didn't need more material. They were drowning in it.
They had apps, textbooks, YouTube, tutors, mock exams, forums. The problem was never access. It was that none of it connected, and none of it answered the question underneath all the studying.
They didn't lack ways to practice. They lacked a way to know if practicing was working.
The one question
Every worry we heard collapsed into the same one.
Different weak skills. Different deadlines. The same fear underneath.
The product could measure how much someone practiced. It couldn't tell them if it was enough.

Architecture
Readiness couldn't be a number on top. It had to be the whole system.
The easy version was a score bolted onto the existing product. We had one already, an old "level-up" estimate, and it was decoration. It couldn't carry a decision this heavy.
So we rebuilt readiness as a loop, not a feature. Daily practice generates evidence. A short mini exam reads that evidence across all four skills. The result updates the readiness score and reshapes what comes next. Then the loop runs again.
I designed how each stage worked, and how each one honestly earned the next.
Onboarding
The first screen couldn't ask learners what level they were.
Most language apps open by asking you to rate yourself. But self-report was the exact thing the product existed to replace. A learner who could judge their own level wouldn't need a readiness score.
So the first screen measured instead of asked. A short adaptive check estimated the starting level from real answers, before anything else.
We didn't ask how ready they were. We found out.
Feedback
A test result that stops at a number leaves you stuck.
A score tells you where you landed. It doesn't tell you why, or what to do next. For someone anxious and short on time, that gap is where they give up.
So the result never stopped at one number. It broke into four skills, because a learner can look ready overall while speaking sits below the line. Each skill came back in plain language, and its weakness became the next session's tasks.
We didn't grade the learner. We told them what to fix.
The score
A percentage looks certain, even when it's built on almost nothing.
After one short exam, the product could show a number. It couldn't yet honestly claim the number meant anything. Nerves, luck, or a familiar topic could move it.
So the score showed its own certainty. Not enough evidence, then an early estimate, then confident readiness. It only spoke plainly once the evidence backed it: repeated, recent, and across all four skills.
The real question wasn't how to score readiness. It was when the product had earned the right to.
More of the system
The rest worked the same way: measure, then act.
Practice that mirrors the exam
Same task types, same timing as the real B1. Practice that doesn't match the exam only trains you for the practice.
One connected path
Courses, practice, and the score ran off one shared source, so the plan always matched your most recent result.
The loop, in the app
Finish a mini exam and the next session rebuilds itself, targeting whichever skill came back weakest.
Reflection
The hard part wasn't making the product smart. It was making it honest.
The system never shipped. Heylama closed before development finished, so the scoring model and its thresholds were never validated. Only the reading exercises reached real learners. That limit is real, and worth saying plainly.
The work still changed how I design AI products. I used to start with the obvious question: how does the system produce a good answer? The harder questions came after. What backs that answer. How recent the evidence is. What's still missing. And how the interface should behave when the honest answer is "not sure yet."
The goal was never a more confident product. It was one honest about what it knew.


