Build a Claude Code Verification Harness

Oct 8, Reserve a seat
Sunil Samson Suresh

Sunil Samson Suresh

Author

Sr. Data Scientist at ValueMomentum · Databricks certified

I have spent my career in the applied data science trenches, building ML and NLP systems that ship. Senior Data Scientist at ValueMomentum, Databricks certified, with production experience across deep learning, LLMs, NLP, and GenAI on Azure.

Background

I have spent my career in the applied data science trenches, building ML and NLP systems that ship. Senior Data Scientist at ValueMomentum, Databricks certified, with production experience across deep learning, LLMs, NLP, and GenAI on Azure.

My focus on learnwithparam is the model and data layer of AI systems: choosing and evaluating embeddings, designing evaluation frameworks that actually catch regressions, data preprocessing for RAG, and the metrics that matter when you scale beyond a demo.

Everything I teach comes from real projects where the difference between a good model and a bad one was a dollar figure on a dashboard, not a benchmark score. The goal is evaluation-driven, data-first AI engineering, not leaderboard chasing.

What I cover

Specialty: Data science, ML, NLP, evaluation.

  • Embedding model selection and benchmarking
  • RAG evaluation frameworks (Ragas, custom metrics)
  • Data preprocessing and chunking for retrieval
  • LLM-as-a-judge evaluation pipelines
  • Vector database comparison and tuning
  • NLP and language model fundamentals
  • Feature engineering for GenAI applications
  • ML experiment tracking and reproducibility

Recent articles by Sunil

LLM Engineering9 min read

Ground truth vs relevancy in RAG evaluation

Why ground truth and relevancy measure different things in RAG evals. When to use each, how to build both datasets, and the 2 metrics that matter most.

LLM Engineering8 min read

Hallucination testing for RAG pipelines

How to test a RAG pipeline for hallucinations systematically. Adversarial prompts, the out-of-scope set, and the metric that catches confabulation.

LLM Engineering8 min read

LLM-based content filtering for RAG pipelines

How to filter irrelevant retrieved chunks with a cheap LLM call before the final answer. The prompt, the batch pattern, and the 40 percent noise reduction.

LLM Engineering8 min read

FAISS vector stores in production RAG

How to use FAISS for production RAG. Index types, persistence, memory trade-offs, and the 4 settings that decide if FAISS beats a managed vector DB.

AI Engineering in Practice8 min read

Real-time agent debugging with Langfuse traces

How to debug a live agent incident using Langfuse traces. The search patterns, the 5-minute workflow, and the post-mortem that catches the root cause.

AI Engineering in Practice9 min read

Agent cost optimization from trace data

How to use Langfuse trace data to find where your agent burns tokens. The 4 queries, the cost-per-user view, and the 50 percent savings patterns.

AI Engineering in Practice8 min read

Langfuse + Grafana: agentic AI monitoring

How to combine Langfuse traces with Grafana dashboards for agent monitoring. The integration, the panels, and the alerting that catches real problems.

LLM Engineering8 min read

Dynamic evaluation metric loading in Python

How to load evaluation metrics dynamically in a Python eval pipeline. The registry pattern, entry points, and the test override that makes CI fast.

LLM Engineering11 min read

Choosing an embedding model for RAG

How to pick an embedding model for production RAG. The 5 criteria that matter, the benchmarks that lie, and the migration cost nobody warns you about.