AN AI WORKSHOP · HALF DAY · FOR TEAMS BUILDING WITH LLMS
No Surprises
How to tell whether your AI features are any good, with human reviewers and automated evals, worked through on your own product.
Book a session01 · THE PROBLEM
You can’t improve what you don’t measure.
Most teams ship LLM features on gut feel. That works for a demo, but not once real users, new models and constant prompt changes are involved.
Spot checks, not evidence.
A handful of prompts tried by hand and a demo that looked good. Nobody can say how often the feature gets it right.
Every change is a gamble.
A new model, a prompt tweak or a retrieval change fixes one case and quietly breaks three others, and nobody finds out until a user does.
Nobody agrees on “good”.
Product, engineering and domain experts each judge outputs differently, so quality arguments go round in circles.
02 · WHAT YOU’LL LEARN
Start with people, then automate.
Automated evals are only as good as the human judgement behind them. We teach both, in that order, the way we run evaluation on our own client work.
PART ONE
Human evaluation.
Agree on what good looks like, then measure it with people who know the domain.
- Writing a rubric your reviewers can apply consistently
- Annotation guidelines and worked examples
- Sampling real outputs so the results mean something
- Measuring agreement between reviewers, and fixing the rubric when they disagree
PART TWO
Automated evaluation.
Turn those human judgements into evals that run on every change.
- Building a test set from real usage and your human labels
- Code-based checks for the things you can assert
- LLM-as-judge for the things you can’t, calibrated against your human labels
- Running evals in CI so regressions are caught before release
03 · THE SESSION
Bring a real AI feature. Leave with evals for it.
Bring one AI feature you’re building or running, and some real examples of what it produces. We work on it together.
20 min
Why evals: how we evaluate the AI features we build for clients
70 min
Human evaluation, hands on: write a rubric for one of your features and label real outputs together
80 min
Automated evaluation, hands on: build a test set, write checks and an LLM judge, and compare it with your human labels
30 min
Putting it in your pipeline: running evals in CI, and next steps for your team
Book your team’s session
We’ll run it at a time and place that suits you. Get in touch with a line on the AI feature you want to evaluate.