Cut AI Agent Costs with Prompt Caching
Agents resend the same instructions, tools and history on every step, and you pay for them every time. Most teams notice only when the invoice arrives, with no idea which agent or prompt caused it. Live, Param measures one agent's spend per run, reorders its prompt so the stable parts hit Claude's prompt cache, and moves the simple steps to a cheaper model.
Why this matters
Agents send the same instructions, tool definitions and history on every step of every run, and you pay for those tokens each time. Bills grow faster than usage, and most teams cannot say which agent or prompt is responsible. The fix is rarely a cheaper model for everything. It is measuring each run, laying prompts out so the repeated part is cached, and sending only the hard steps to the expensive model.
What happens in the hour
- The problem4 min
Why the agent bill grows faster than usage, and why nobody can say which agent caused it.
- Built live40 min
Input, cached and output tokens counted per step; the prompt reordered so stable tools and instructions come first; simple steps routed to a smaller model.
- Questions16 min
The before and after cost of the same run, and where to start on your own agents.
What you will be able to do
- Measure each run
- Count input, cached and output tokens per step and per agent.
- Lay out prompts for the cache
- Put stable tools and instructions first so repeat calls bill at the cached rate.
- Pick a model per step
- Send simple steps to a smaller model and keep the large one where quality needs it.
The ideas behind it
Cost per run
Tokens in, tokens out and price for one complete run of one agent. The number every other decision starts from.
Prompt caching
The model provider stores the stable start of a prompt and charges much less to reuse it. It only works if that part comes first and does not change.
Prompt layout
Instructions and tools first, history next, the new question last. The order decides how much of each call can be cached.
A model per step
Simple steps such as routing or extraction go to a small, cheap model; reasoning-heavy steps keep the large one.
Where it fits
The parts of a production agent system, in the order the buildcamp builds them. The lit tiles are the ones this live lab builds; the Agentic AI Buildcamp for Engineers builds all of them.
- Roles and modelsOne job per agent, a model chosen for that job, and a token budget.
- OrchestrationA queue agents pick work from, with hand-offs a person can follow.
- Context and memoryRetrieval with sources, memory across sessions, prompts laid out for the cache.
- Tool callingTyped tools that act in GitHub, Linear and Google.
- MCP serversEvery tool behind an MCP gateway, scoped to the role that needs it.
- Durable executionRuns that resume after a crash and never repeat a write.
- Human in the loopA person approves anything that cannot be undone.
- Identity and permissionsEach agent signs in as itself and acts on behalf of a user.
- LLM gatewayRouting, caching and a budget on every model call.
- Traces, evals and costEvery run traced, scored and charged to the agent that made it.
- Multi-tenancyEach team or customer kept apart, in data and in the bill.
- Policy as codeRules checked on every action, not written in a document.
- Signed skillsNew roles built from reviewed, signed skills.
Before you come
- For
- Engineers running agents whose bill grows faster than their usage, and the tech leads who have to explain and forecast it.
- You need
- None. Code is shown, not required.
- Track
- Agentic AI for engineers who build and run it.
Questions
Does prompt caching change the answers?
No. The model reads the same prompt. Caching changes what you are charged and how fast the first token arrives.
Does this only work with Claude?
The lab uses Claude's prompt cache. OpenAI and Google offer caching too, with different rules, and the layout advice applies to all of them.
How much can we save?
It depends on how much of each call repeats. Agents that resend long instructions and tool lists on every step save the most, and the lab shows how to measure it on your own runs.
Will a cheaper model make the agent worse?
On some steps, yes. That is why the lab picks the model per step and checks the result, rather than switching everything at once.
Is the Live Lab really free?
Yes. Live Labs are free on Maven. You sign up with your email and get the join link and the recording.
What if I cannot make it live?
Sign up anyway. Everyone who signs up gets the recording, so you can watch the build later and reply with questions.
Do I need to code along?
No. Most people watch the build and ask questions. Every step is shown, so you can repeat it on your own afterwards.