Abby Feng
Contributing Writer, Explore Agentic
About Abby
Abby writes the evaluation-program side of shipping LLM systems. The long-form LLM evaluation guide is the anchor: what to measure on a generation task versus a retrieval task versus a full agent trajectory, how to build an LLM-as-a-judge that agrees with human labels often enough to be worth acting on, how RAG evaluation differs from end-to-end agent evaluation, and how to wire the whole thing into CI so a regression fails a pipeline rather than surfacing in a support ticket weeks later.
The recurring argument is that the hard part is rarely the metric — it is the program around it: who owns the eval set, how it grows, what happens when a score moves, and how an engineer defends a quality number in a review meeting. The healthcare review byline extends that to voice agents calling insurance payers, where the escalation threshold is the real quality control. Expect rubrics and operating mechanics, not leaderboards.
Pieces written or reviewed by Abby
2 pieces across insight on this site — 1 written by Abby and 1 reviewed. A written byline means Abby researched and drafted the piece; a reviewed byline means Abby read it against its cited sources and could defend its claims before it published. Every row below is labelled either way.
Other contributors
- Ryo HangEditor-in-Chief
- Alexander GromanContributing Writer
- Kelvin YuContributing Writer
- Soraya ZhengContributing Writer
- Cynthia ZhangContributing Writer
- Tommy TaoContributing Writer
- Chandler BensonContributing Writer
- Elias SaljukiContributing Editor
- Ginny CasuccioContributing Editor
- Gloria Qian ZhangContributing Editor
- Laura Bradley McCoyContributing Writer
- Merve TengizContributing Editor
- Michael CloughEditorial Advisor