AN AI WORKSHOP · HALF DAY · FOR TEAMS RUNNING AI IN PRODUCTION
Unbloated
Make the AI features you already run cheaper, faster and more reliable, without the answers getting worse.
Book a session01 · THE PROBLEM
It worked in the demo. Then it went live.
Getting an AI feature launched is the first half of the job. Real traffic brings costs, wait times and failures the prototype never had to handle.
The bill grew faster than the usage.
Every request goes to the biggest model with the longest prompt, and the monthly invoice is the first anyone hears of it.
Users wait, then leave.
Long chains of model and tool calls add up to answers that take seconds too long for the people using them.
A bad day for your provider is a bad day for you.
Rate limits, timeouts and outages turn straight into errors, because nothing is in place to retry or fall back.
02 · WHAT YOU’LL LEARN
Measure first, then optimise.
Every optimisation is a trade-off with quality. We show you how to find the biggest wins in your traces and check each one against your evals before it ships.
PART ONE
Cost.
Find out where the money goes, then spend less of it without the answers getting worse.
- Tracking cost per feature and per request from your traces
- Routing simpler requests to smaller, cheaper models
- Prompt caching and trimming the context you send
- Batching work that doesn’t need an instant answer
PART TWO
Speed and reliability.
Get answers to users sooner, and keep the feature working when a provider doesn’t.
- Streaming and parallel calls so users see results sooner
- Timeouts, retries and fallback models
- Handling rate limits and quotas as usage grows
- Checking every change against your evals so quality doesn’t slip
03 · THE SESSION
Bring a live AI feature. Leave with a cheaper, faster one.
Bring one AI feature that’s running in production, with access to its traces or logs. If you don’t have evals for it yet, start with our No Surprises workshop.
20 min
Where the time and money go: reading traces from AI features we run for clients
70 min
Cost, hands on: measure what your feature costs per request, then route, cache and trim to bring it down
60 min
Speed and reliability, hands on: stream, parallelise and add fallbacks to your feature
50 min
Proving it: run your changes against evals to check quality held, and plan next steps for your team
Book your team’s session
We’ll run it at a time and place that suits you. Get in touch with a line on the AI feature you want to optimise.