Case study

Exam Readiness

Summary

As the sole product designer, I helped shape the strategy and led the design for Heylama’s pivot to B1 exam prep. The product could track practice, but it could not answer the question that mattered most:
“Am I ready to pass?”

Team

Me
FE
FE
BE
QA
CEO
CTO

Status

Readiness system
unshipped.

9:41📶 🔋

System scope

4

Exam skills

2

Exam formats

1

Readiness model

01Stakes

You moved to Berlin. Now one exam stands between you and what comes next.

New country, new language, and a life you are building anew. Then comes a deadline. You need a B1 certificate for residency, work, or the next step in the process.

Failing can mean another fee, another exam date, and months of waiting. Readiness is not abstract when the result affects what you can do next.

It was never just about the price of the test. It was about the price of living where you chose to live.

02Pivot

We narrowed from general language learning to one measurable outcome.

Heylama was a small team competing in a broad language-learning market where larger products had more content, reach, and resources.

B1 exam prep gave us a more focused problem. There was a defined exam, four tested skills, a deadline, and a clear outcome.

That changed the question from “How do we help people learn a language?” to “How do we help someone know when they are ready to pass?”

What we found

These learners did not need more material. They already had too much of it.

The problem was not access. Learners could find almost anything they wanted to study.

The problem was that all of that work stayed disconnected. One tool tracked vocabulary, another gave mock exams, another helped with speaking. Learners still had to piece together what those signals meant.

No single place turned their preparation into a clear picture of readiness.

One question

Different worries kept leading back to the same decision.

The wording changed, but the decision underneath it did not: book the exam now, or keep preparing?

That gave us a much clearer product target than progress, streaks, or activity. Readiness had to bring several skills together into one useful answer without hiding what that answer was based on.

Scattered learner quotes converging on one question: am I actually ready?
03Architecture

Readiness could not be a number on top. It had to be the whole system.

The easiest move was to add another score to the existing product. We already had a level-up estimate, and it showed the problem with that approach. A number that sits outside the learning experience quickly becomes decoration.

I treated readiness as a feedback loop instead. An assessment had to do more than update a score. It had to create evidence that changed what the learner did next. That made the score an output of the system, not the system itself.

04Onboarding

The first screen could not ask learners what level they were.

If the product existed to reduce guesswork, starting with self-report made little sense.

I used a short adaptive check to establish the starting point from performance instead. That asked for more effort upfront, but it gave the rest of the experience evidence it could actually use.

We were no longer asking learners to diagnose themselves before we helped them.

05Feedback

A test result that stops at a number leaves you stuck.

Two learners can land on the same overall result for completely different reasons. One may struggle with speaking, another with writing.

So the result needed to explain what was driving the score and change what happened next. The learner could see where the weakness came from, while the next plan responded to it.

The result became a decision about what to practice next, not just a record of what happened.

06Readiness score

A percentage looks certain, even when it is built on almost nothing.

The challenge was not calculating a percentage. It was deciding how much confidence that percentage deserved.

One strong attempt could be luck, nerves could distort another, and a familiar topic could make performance look stronger than it really was. I wanted confidence to grow only when the evidence became repeated, recent, and spread across all four skills.

So confidence became part of the interface itself. The product could be cautious when the evidence was thin and become more direct as that evidence became stronger.

The real question was when the product had earned the right to sound certain.

More of the system

The rest worked the same way: measure, then act.

Practice that mirrors the exam

The tasks followed the real B1 format and timing. If practice behaved differently from the exam, the evidence it produced would be less useful.

One connected path

Courses, practice, and readiness worked from the same evidence. New results could change the plan instead of sitting in a separate progress view.

The loop, in the app

Finishing a Mini Exam did more than update the score. The weakest skill became the focus of the next session, turning assessment directly into action.

07Reflection

The hard part was making the product honest about what it knew.

The readiness system never shipped. Heylama closed before development finished, so the scoring model, confidence thresholds, and adaptive loop were never validated with real learners. Only parts of the exam-preparation experience, including reading exercises, reached production.

That leaves the hardest assumptions unresolved. We still needed to learn how much evidence was enough, how quickly confidence should respond to new performance, and whether learners would trust a score that openly admitted uncertainty.

The project changed what I look for when designing AI products. A useful answer is only part of the job. I now ask what evidence supports it, how recent that evidence is, what is still missing, and how the product should behave when certainty is low.

The product should be no more certain than the evidence behind it.

While you’re here