AN AI WORKSHOP · HALF DAY · FOR TEAMS RUNNING AI IN PRODUCTION

Unbloated

Make the AI features you already run cheaper, faster and more reliable, without the answers getting worse.

Book a session

01 · THE PROBLEM

It worked in the demo. Then it went live.

Getting an AI feature launched is the first half of the job. Real traffic brings costs, wait times and failures the prototype never had to handle.

The bill grew faster than the usage.

Every request goes to the biggest model with the longest prompt, and the monthly invoice is the first anyone hears of it.

Users wait, then leave.

Long chains of model and tool calls add up to answers that take seconds too long for the people using them.

A bad day for your provider is a bad day for you.

Rate limits, timeouts and outages turn straight into errors, because nothing is in place to retry or fall back.

02 · WHAT YOU’LL LEARN

Measure first, then optimise.

Every optimisation is a trade-off with quality. We show you how to find the biggest wins in your traces and check each one against your evals before it ships.

PART ONE

Cost.

Find out where the money goes, then spend less of it without the answers getting worse.

  • Tracking cost per feature and per request from your traces
  • Routing simpler requests to smaller, cheaper models
  • Prompt caching and trimming the context you send
  • Batching work that doesn’t need an instant answer

PART TWO

Speed and reliability.

Get answers to users sooner, and keep the feature working when a provider doesn’t.

  • Streaming and parallel calls so users see results sooner
  • Timeouts, retries and fallback models
  • Handling rate limits and quotas as usage grows
  • Checking every change against your evals so quality doesn’t slip

03 · THE SESSION

Bring a live AI feature. Leave with a cheaper, faster one.

Bring one AI feature that’s running in production, with access to its traces or logs. If you don’t have evals for it yet, start with our No Surprises workshop.

  • 20 min

    Where the time and money go: reading traces from AI features we run for clients

  • 70 min

    Cost, hands on: measure what your feature costs per request, then route, cache and trim to bring it down

  • 60 min

    Speed and reliability, hands on: stream, parallelise and add fallbacks to your feature

  • 50 min

    Proving it: run your changes against evals to check quality held, and plan next steps for your team

Book your team’s session

We’ll run it at a time and place that suits you. Get in touch with a line on the AI feature you want to optimise.

Book a session