LLM Engineering8 min read
Choosing the LLM judge for evaluation pipelines
How to pick the LLM that grades your LLM. The cost-quality tradeoffs, the calibration check, and why a weaker judge is sometimes the right call.
Build a Claude Code Verification Harness live Thursday, Oct 8, 12:00 PM ET
Oct 8, Reserve a seat
Sr. Data Scientist at ValueMomentum · Databricks certified
I have spent my career in the applied data science trenches, building ML and NLP systems that ship. Senior Data Scientist at ValueMomentum, Databricks certified, with production experience across deep learning, LLMs, NLP, and GenAI on Azure.
I have spent my career in the applied data science trenches, building ML and NLP systems that ship. Senior Data Scientist at ValueMomentum, Databricks certified, with production experience across deep learning, LLMs, NLP, and GenAI on Azure.
My focus on learnwithparam is the model and data layer of AI systems: choosing and evaluating embeddings, designing evaluation frameworks that actually catch regressions, data preprocessing for RAG, and the metrics that matter when you scale beyond a demo.
Everything I teach comes from real projects where the difference between a good model and a bad one was a dollar figure on a dashboard, not a benchmark score. The goal is evaluation-driven, data-first AI engineering, not leaderboard chasing.
Specialty: Data science, ML, NLP, evaluation.
LLM Engineering8 min read
How to pick the LLM that grades your LLM. The cost-quality tradeoffs, the calibration check, and why a weaker judge is sometimes the right call.
LLM Engineering9 min read
Why ground truth and relevancy measure different things in RAG evals. When to use each, how to build both datasets, and the 2 metrics that matter most.
LLM Engineering8 min read
How to use Pydantic models to force your RAG planner LLM to return structured steps. The schema, the retry loop, and why plain JSON prompts break in production.
LLM Engineering8 min read
How to test a RAG pipeline for hallucinations systematically. Adversarial prompts, the out-of-scope set, and the metric that catches confabulation.
LLM Engineering8 min read
How to test a RAG pipeline like real software. Unit, integration, and eval tests that catch regressions before they ship. The 3-layer test strategy.
LLM Engineering8 min read
How to filter irrelevant retrieved chunks with a cheap LLM call before the final answer. The prompt, the batch pattern, and the 40 percent noise reduction.
LLM Engineering8 min read
How to pick the right k value for your RAG retriever. The 3-step tuning process, the failure modes of k=3 and k=20, and the sweet spot in between.
LLM Engineering8 min read
How to combine multiple vector stores in one RAG pipeline. The merge pattern, the deduplication rule, and when multi-source beats a single index.
LLM Engineering8 min read
How to use FAISS for production RAG. Index types, persistence, memory trade-offs, and the 4 settings that decide if FAISS beats a managed vector DB.
AI Engineering in Practice8 min read
How to debug a live agent incident using Langfuse traces. The search patterns, the 5-minute workflow, and the post-mortem that catches the root cause.
AI Engineering in Practice9 min read
How to use Langfuse trace data to find where your agent burns tokens. The 4 queries, the cost-per-user view, and the 50 percent savings patterns.
AI Engineering in Practice8 min read
How to combine Langfuse traces with Grafana dashboards for agent monitoring. The integration, the panels, and the alerting that catches real problems.
LLM Engineering8 min read
How to wire eval pipelines into CI so every agent change is scored automatically. The nightly job, the regression gate, and the dashboard that matters.
LLM Engineering8 min read
How to load evaluation metrics dynamically in a Python eval pipeline. The registry pattern, entry points, and the test override that makes CI fast.
LLM Engineering9 min read
Why LLM judges without explicit reasoning drift, and how chain-of-thought rationales make their scores defensible. The prompt, the parser, the trust.
LLM Engineering9 min read
How to build an LLM-as-a-judge evaluation framework for agentic AI. The prompt, the rubric, the bias controls, and the loop that catches regressions.
LLM Engineering11 min read
How to add chain-of-thought reasoning to a RAG pipeline. The prompt, the parsing, and the cases where CoT beats a straight answer by a wide margin.
LLM Engineering11 min read
How to pick an embedding model for production RAG. The 5 criteria that matter, the benchmarks that lie, and the migration cost nobody warns you about.
LLM Engineering11 min read
How RecursiveCharacterTextSplitter works, why it beats naive chunking, and the separator order that makes or breaks retrieval quality.
LLM Engineering11 min read
How to use RAGAS to evaluate RAG pipelines. The 4 metrics that matter, the eval loop, and the trap that makes most RAG evals dishonest.