Top 1% Engineers, issue 1
Every AI system design starts with one question: who decides what runs next?
Workflow, single agent or several agents. How to choose before you write code, and why Season 1 of Live Labs opens with it.

The question under every design review
When a team shows me an AI feature that is slow, expensive or unpredictable, the cause is rarely the model. It is a decision made in the first hour and never revisited: who decides what runs next.
If your code decides, you have a workflow. If the model decides, you have an agent. If one model hands work to other models, you have a multi-agent system. Each answer has a different cost, a different way to fail and a different way to debug.
Start with a workflow, and earn the agent
A workflow is a fixed sequence of steps where one or two of them call a model. The path is known, so you can test it, retry it and explain it to an auditor. Most business automation belongs here: triage an invoice, summarise a ticket, draft a reply for approval.
Move to an agent when the steps cannot be known in advance: the task needs the model to look, decide and look again. Research, debugging and open-ended support are good examples. The price is that every run can take a different path, so you need traces and evals before you need anything else.
A second agent must pay for itself
Splitting work across several agents feels like good engineering. Often it adds handoffs, duplicated context and more places for a run to go wrong. A second agent earns its place when the work splits into parts that need different tools or different permissions, or that can run in parallel.
A useful test: if you cannot write down what each agent owns and what it hands back, you have one agent with extra steps.
A short checklist for your next design review
Before anyone writes a prompt, answer these on one page.
- Who decides the next step. Your code, the model, or a supervisor model. Write it down.
- What happens when a step fails. Retry, ask a person, or stop. Each tool call needs an answer.
- How you will see a bad run. Traces from the first day, and a dataset made from real failures.
- What one request costs. Estimate tokens per run before you build, then measure it after.
- Where a person approves. Anything that sends, pays or deletes waits for a human in the loop.
Why Season 1 opens with this
The first Live Lab, on Tuesday 20 October, builds one support-triage task three ways: as a workflow, as a single agent and as a supervisor with two workers. We trace all three, count the tokens and break each one on purpose.
The labs after it follow the same idea. Design first, then evals and tracing, durable workflows, an MCP gateway, ingestion and vector search at scale, browser agents, AI QA, a company brain and an AI SDR. Each one starts from a public GitHub repo and ends with something running on your own code.
If you take one thing from this issue: before your next AI feature, write down who decides what runs next. Reply and tell me what you wrote.
Param