Live LabsBuildersEvals3 hours

Add evals and tracing to an agent

You cannot improve what you cannot see. We instrument an agent, turn its failed runs into a dataset, and run an experiment that scores a fix before it ships.

You leave with

Traces, a dataset built from real failures, and an experiment that proves a prompt change helped.

How the time is spent

  1. I do45 minParam builds it live

    I instrument the sample agent and walk through the traces of three failed runs.

  2. We do72 minTogether, step by step

    We build a dataset from those failures and write the scoring together.

  3. You do63 minOn your own code

    You change a prompt, run the experiment and show whether the score moved.

What you will be able to do

Read a trace

Find the step where a run went wrong.

Build a dataset from failures

Turn production mistakes into test cases.

Prove a fix

Run an experiment before a change ships.

Starts from

Before you come

For
Engineers with an agent or LLM feature already running.
You need
TypeScript. A free Langfuse Cloud account (the workshop repo uses Cloud).
Stack
TypeScript, Langfuse
Track
Hands-on agentic AI for engineers who ship.