AI Engineering in Practice8 min read
Retry patterns for LLM API errors in production
How to build retry logic that handles rate limits, timeouts, and transient failures without burning money. The backoff rules and the 3 errors you must not retry.
Build a Claude Code Verification Harness live Thursday, Oct 8, 12:00 PM ET
Oct 8, Reserve a seatI build complex agentic systems for a living. Co-founder and CTO at Faltara, marketing AI and automation specialist at Wise, and a software engineer at my core. I ship AI that earns its complexity in production, not in demos.
My focus on learnwithparam is the part of agent engineering nobody likes to teach: auth, rate limits, observability, resilience, the dozen little decisions that turn a prototype into a service you can leave running unattended. If a topic makes a staff engineer nod knowingly, it probably belongs in my track.
I cover the patterns I actually use at work: tool design, session and tenant models, circuit breakers and fallbacks, Langfuse and Prometheus for tracing, and the deployment layer that keeps long-running agents alive. The goal is production-grade, not tutorial-grade.
Specialty: Complex agentic systems.
AI Engineering in Practice8 min read
How to build retry logic that handles rate limits, timeouts, and transient failures without burning money. The backoff rules and the 3 errors you must not retry.
AI Engineering in Practice8 min read
How Prometheus metrics surface performance bottlenecks in agentic AI. The 3 queries, the alert rules, and the dashboard that finds hot loops fast.
AI Engineering in Practice8 min read
How to stress test an agentic AI service before it ships. Concurrency, tokens, latency budgets, and the load profile that simulates real traffic.
AI Engineering in Practice8 min read
How to inject API keys and secrets into agent containers without baking them into the image. BuildKit secrets, runtime injection, and the 3 bad patterns.
AI Engineering in Practice7 min read
How to configure CORS for a production agentic API without wildcard origins. The allowlist, the credentials flag, and the preflight that breaks SSE.
AI Engineering in Practice8 min read
Why FastAPI lifespan is the only right place for agent startup code. Per-worker initialization, ordered teardown, and the bugs it kills.
AI Engineering in Practice8 min read
How to wire LangGraph into a FastAPI chatbot API with streaming, persistence, and per-user threads. The production pattern that scales past demos.
AI Engineering in Practice9 min read
How FastAPI Depends makes agent auth testable and composable. The pattern, the chain, and why module-level globals break at scale.
AI Engineering in Practice9 min read
How to persist agent state in Postgres so conversations survive restarts. The schema, the session writer, and the idempotency rule that prevents loss.
AI Engineering in Practice9 min read
How a service layer in an AI agent codebase decouples business logic from HTTP routes. The pattern, the tests, and the refactor from a fat route.
AI Engineering in Practice10 min read
How to structure an agentic AI codebase so it survives growth. The 4-layer pattern, the cut lines, and the refactor that keeps a project shippable.
AI Engineering in Practice10 min read
How to instrument an agentic AI service with Prometheus. The 4 metrics that matter, the histogram trap, and the dashboard that surfaces regressions.
LLM Engineering11 min read
Why linear LangChain chains fall over on real agents and how LangGraph's stateful graphs replace them. The state model, loops, and upgrade path.
AI Engineering11 min read
How circuit breakers prevent LLM outages from cascading through your agent. The 3 states, the failure window, and the 50-line implementation.
AI Engineering in Practice9 min read
How to survive LLM provider outages with Tenacity retries and fallback models. The retry policy, the fallback chain, and the 60-line pattern.
AI Engineering in Practice11 min read
How to sanitize agent API inputs beyond frontend validation. Prompt injection defense, payload limits, and the 4 layers every agent service needs.
AI Engineering in Practice11 min read
How to version an agentic API without breaking clients. The URL prefix pattern, the deprecation playbook, and when to ship v2.
AI Engineering in Practice10 min read
How to wire Langfuse into an agentic AI service for full observability. The trace hierarchy, the decorator pattern, and what to log per span.
AI Engineering11 min read
How coding agents remember context across sessions. The memory store, the recall pattern, and the 3 kinds of memory every agent should keep.
AI Engineering12 min read
How to design custom tools for coding agents that are not read, edit, or bash. The naming rules, the schema patterns, and the 3 custom tools that pay off.