AN AI WORKSHOP · HALF DAY · FOR TEAMS BUILDING WITH LLMS

No Surprises

How to tell whether your AI features are any good, with human reviewers and automated evals, worked through on your own product.

Book a session

01 · THE PROBLEM

You can’t improve what you don’t measure.

Most teams ship LLM features on gut feel. That works for a demo, but not once real users, new models and constant prompt changes are involved.

Spot checks, not evidence.

A handful of prompts tried by hand and a demo that looked good. Nobody can say how often the feature gets it right.

Every change is a gamble.

A new model, a prompt tweak or a retrieval change fixes one case and quietly breaks three others, and nobody finds out until a user does.

Nobody agrees on “good”.

Product, engineering and domain experts each judge outputs differently, so quality arguments go round in circles.

02 · WHAT YOU’LL LEARN

Start with people, then automate.

Automated evals are only as good as the human judgement behind them. We teach both, in that order, the way we run evaluation on our own client work.

PART ONE

Human evaluation.

Agree on what good looks like, then measure it with people who know the domain.

  • Writing a rubric your reviewers can apply consistently
  • Annotation guidelines and worked examples
  • Sampling real outputs so the results mean something
  • Measuring agreement between reviewers, and fixing the rubric when they disagree

PART TWO

Automated evaluation.

Turn those human judgements into evals that run on every change.

  • Building a test set from real usage and your human labels
  • Code-based checks for the things you can assert
  • LLM-as-judge for the things you can’t, calibrated against your human labels
  • Running evals in CI so regressions are caught before release

03 · THE SESSION

Bring a real AI feature. Leave with evals for it.

Bring one AI feature you’re building or running, and some real examples of what it produces. We work on it together.

  • 20 min

    Why evals: how we evaluate the AI features we build for clients

  • 70 min

    Human evaluation, hands on: write a rubric for one of your features and label real outputs together

  • 80 min

    Automated evaluation, hands on: build a test set, write checks and an LLM judge, and compare it with your human labels

  • 30 min

    Putting it in your pipeline: running evals in CI, and next steps for your team

Book your team’s session

We’ll run it at a time and place that suits you. Get in touch with a line on the AI feature you want to evaluate.

Book a session