Contributing Writer · Retrieval & Search

Tommy Tao

Contributing Writer, Explore Agentic

About Tommy

Tommy writes the measurement side of retrieval. The RAG evaluation work is the core of it: recall@k and precision against a labeled retrieval set, faithfulness and answer-relevance scoring, LLM-as-a-judge rubrics with the calibration and bias checks that decide whether a judge is worth trusting, and CI gating so a retrieval regression fails a build instead of a customer demo.

Around that sit the applied pieces — intelligent document processing with OCR and multimodal extraction, semantic search over retail inventory where every miss has a dollar cost, and the agentic RAG glossary entry on retrieval that lives inside the planning loop rather than firing once at the start. Tommy also wrote the Claude Haiku versus Sonnet comparison, which routes model choice by task shape on measured cost-per-task rather than rate-card arithmetic, and reviews the prompt caching, agent evaluation, and EDI 270/271 pieces. Expect evaluation harnesses, worked numbers, and a stated limit on what each metric does not cover.

RAG evaluationLLM-as-a-judgeAgentic RAGDocument processingSemantic search