Live LabsBuildersEvals2 hours

Test an agent like software

Prompts change every week. We put assertions and adversarial tests around an agent so a bad change fails in CI, not in front of a customer.

You leave with

A regression suite and a red-team run for one agent, failing the build when quality drops.

How the time is spent

  1. I do30 minParam builds it live

    I break a working prompt with one innocent edit and show nothing catches it.

  2. We do48 minTogether, step by step

    We write assertions and a red-team config for the same agent.

  3. You do42 minOn your own code

    You add the suite as a CI gate on your own repo.

What you will be able to do

Assert on LLM output

Write checks that hold across model versions.

Red-team before release

Find prompt injection and leaks early.

Gate releases

Fail the build when quality drops.

Starts from

Before you come

For
Engineers responsible for an LLM feature in production.
You need
Node or Python. A GitHub repo you can add a workflow to.
Stack
promptfoo, DeepEval, GitHub Actions
Track
Hands-on agentic AI for engineers who ship.