One prompt or model change can quietly break an agent that worked yesterday. Most teams test by trying a few questions by hand, so their users find the regressions first. Live, Param turns real agent runs from Langfuse into a test set, scores them with code checks and an LLM judge, and fails the build when a change makes the agent worse.
Why this matters
A prompt tweak or a model update can quietly break an agent that worked yesterday. Most teams test by trying a few questions by hand, so users find the regressions first. Evals turn the runs you already have into a test set that runs on every change, the same way unit tests protect normal code. Without them, every change to an agent is a guess.
What happens in the hour
The problem4 min
Why testing an agent by hand means your users find the regressions first.
Built live40 min
Real traces from Langfuse turned into labelled cases, code checks for facts and format, a judge for the rest, and the evals run in CI.
Questions16 min
The score that should block a release, and how to start a test set for your own agent.
What you will be able to do
Turn runs into tests
Pick traces from real use and label what a good answer looks like.
Score with checks and a judge
Code checks for facts and format, a judge for the rest, tuned to your labels.
Fail the build
Run the evals in CI on every prompt or model change.
The ideas behind it
A test set from real runs
Good and bad runs from production traces become the cases you test against, so tests match what users actually ask.
Code checks
Fast, exact tests for things that must always hold: valid JSON, the right tool called, no secret in the output.
An LLM judge
A model scores answers against a written rubric for qualities code cannot check, such as whether the answer is grounded.
A gate in CI
The eval runs on every pull request and fails the build when the score drops.
Where it fits
The parts of a production agent system, in the order the buildcamp builds them. The lit tiles are the ones this live lab builds; the Agentic AI Buildcamp for Engineers builds all of them.
Week 1Architect
Roles and modelsOne job per agent, a model chosen for that job, and a token budget.
OrchestrationA queue agents pick work from, with hand-offs a person can follow.
Context and memoryRetrieval with sources, memory across sessions, prompts laid out for the cache.
Tool callingTyped tools that act in GitHub, Linear and Google.
Week 2Build
MCP serversEvery tool behind an MCP gateway, scoped to the role that needs it.
Durable executionRuns that resume after a crash and never repeat a write.
Human in the loopA person approves anything that cannot be undone.
Week 3Secure and deploy
Identity and permissionsEach agent signs in as itself and acts on behalf of a user.
LLM gatewayRouting, caching and a budget on every model call.
Traces, evals and costEvery run traced, scored and charged to the agent that made it.
Multi-tenancyEach team or customer kept apart, in data and in the bill.
Week 4Govern and extend
Policy as codeRules checked on every action, not written in a document.
Signed skillsNew roles built from reviewed, signed skills.
Before you come
For
Engineers who change prompts and models and hope nothing broke, and tech leads who need proof an agent is ready before it ships.
You need
Python helps.
Track
Agentic AI for engineers who build and run it.
Questions
Can an LLM judge be trusted?
Only once you have checked it against human labels on a sample. The lab shows how to do that before relying on it.
How many test cases do we need?
Start with a few dozen real cases that cover the failures you have seen, and grow the set every time a new failure appears.
Does this need Langfuse?
Langfuse is where the runs come from in the lab. Any tracing tool that exports runs works the same way.
Is Python required?
It helps, since the checks are written in Python, but the build is shown and explained step by step.
Is the Live Lab really free?
Yes. Live Labs are free on Maven. You sign up with your email and get the join link and the recording.
What if I cannot make it live?
Sign up anyway. Everyone who signs up gets the recording, so you can watch the build later and reply with questions.
Do I need to code along?
No. Most people watch the build and ask questions. Every step is shown, so you can repeat it on your own afterwards.