Tommy Tao
Contributing Writer, Explore Agentic
About Tommy
Tommy writes the measurement side of retrieval. The RAG evaluation work is the core of it: recall@k and precision against a labeled retrieval set, faithfulness and answer-relevance scoring, LLM-as-a-judge rubrics with the calibration and bias checks that decide whether a judge is worth trusting, and CI gating so a retrieval regression fails a build instead of a customer demo.
Around that sit the applied pieces — intelligent document processing with OCR and multimodal extraction, semantic search over retail inventory where every miss has a dollar cost, and the agentic RAG glossary entry on retrieval that lives inside the planning loop rather than firing once at the start. Tommy also wrote the Claude Haiku versus Sonnet comparison, which routes model choice by task shape on measured cost-per-task rather than rate-card arithmetic, and reviews the prompt caching, agent evaluation, and EDI 270/271 pieces. Expect evaluation harnesses, worked numbers, and a stated limit on what each metric does not cover.
Pieces written or reviewed by Tommy
12 pieces across 4 formats on this site — 6 written by Tommy and 6 reviewed. A written byline means Tommy researched and drafted the piece; a reviewed byline means Tommy read it against its cited sources and could defend its claims before it published. Every row below is labelled either way.
- Pillar · Reviewed Retrieval is still the hardest part of the stack
- Glossary · Written Agentic RAG
- Comparison · Reviewed Glean Knowledge Graph vs Microsoft Graph: the architectural fork that decides enterprise search at scale
- Insight · Written Claude Haiku vs Sonnet: Which Model for Which Agent Workload (with the Cost-Per-Task Math)
- Insight · Reviewed Prompt Caching on Bedrock and the Anthropic API: The Cost Lever That Actually Moves Agent Bills
- Insight · Written LLM-as-a-Judge: Building Automated Evaluation You Can Actually Trust
- Insight · Reviewed Glean year-zero cost: what a 500-seat deployment actually burns before the first renewal
- Insight · Written Your Documents Are Sitting on a Gold Mine — Here's How AI Unlocks It
- Insight · Written Retail Inventory Search Is Broken — Here's How AI Cuts the Cost of Every Miss
- Insight · Reviewed AI Agent Evaluation: How to Test Agents Before Production
- Insight · Written RAG Evaluation: Measuring Retrieval Quality Before You Ship
- Insight · Reviewed EDI 270/271 Eligibility Checks: Why a 96%-Electronic Standard Still Makes Clinics Call Payers
Other contributors
- Ryo HangEditor-in-Chief
- Alexander GromanContributing Writer
- Kelvin YuContributing Writer
- Soraya ZhengContributing Writer
- Cynthia ZhangContributing Writer
- Chandler BensonContributing Writer
- Elias SaljukiContributing Editor
- Ginny CasuccioContributing Editor
- Gloria Qian ZhangContributing Editor
- Laura Bradley McCoyContributing Writer
- Merve TengizContributing Editor
- Michael CloughEditorial Advisor