Engineering Insights

LLM Evals Before Launch: A Studio Checklist That Caught Real Bugs

September 02, 2026 İdris Yıldız

An LLM eval is a fixed set of inputs with a judged expected behaviour — exact JSON, a citation, a refusal, a CEFR band — that you run before a prompt or model change reaches users. If you cannot fail a build, you do not have an eval. You have a vibe.

Our first useful eval at Gidysoft was ugly: 40 Polylingo sentences, a spreadsheet, and a script that asked the model for a level and a corrected line. It caught a prompt change that made every A2 sentence come back as C1. That would have shipped on a Friday. The eval is why it did not.

Three layers, not one score

Contract tests: can it still emit the schema? Grounding tests: did it only use the retrieved docs? Behaviour tests: did it refuse the thing it must refuse? We score them separately. A model that writes prettier prose and starts inventing sources has not “improved”. It has regressed on grounding.

Judges can be code (regex, JSON schema, contains-citation) or a second model with a tight rubric. We prefer code. When we use a model judge, we sample disagreements and have a human break the tie weekly. Otherwise the judge and the candidate drift together.

Production is the next eval

Offline sets go stale. We sample live traces — stripped of personal data — into a “this week” set, and we alert when the live distribution leaves the offline set behind. That is how we noticed EduPick’s assistant starting to answer policy questions after a content import. The import was correct. The retrieval filter was not.

What we do not do

We do not chase public leaderboards. They do not contain our worksheets, our Turkish–English mix, or our support macros. We do not let a founder’s favourite anecdote override a failed suite. And we do not A/B a broken schema “to gather signal”. Schema breaks are 500s with extra steps.

Questions we keep getting

How many cases is enough? Start with 30 that represent the jobs you sell. Add one every time production surprises you. Hundreds without ownership is theatre.

Who owns the suite? The product engineer who owns the feature, same as unit tests. Not a floating “AI person”.

Can we eval tone? Lightly. We eval harm, schema, and grounding first. Tone is a review comment until it becomes a support pattern.