BUILDING SYSTEMS THAT TELL YOU WHEN THEY FAIL

PRECISION, IDENTICAL ITEMSRULES0.958B JUDGE0.360.923MICRO F1 · 68 HAND-LABELED TRACESSTALLS9/9 → 0/944 DETECTIONS · NO TRUE-NEGATIVE BUCKET
(About me)(AI / ML Engineer)

I build LLM agent systems and the evaluation infrastructure that keeps them honest. At ClearAgent I built a LangGraph multi-agent service for agent compliance, and the harness that scores whether it actually works. Before that I shipped a production LLM pipeline that cut report delivery from days to hours across 200+ monthly reports. I care most about the part that usually gets skipped: measuring whether the thing you built does what you claimed, and publishing the answer when it doesn't.

  • ClearAgentAI Engineer25'
  • DataCorp Traffic Private LimitedAI/ML Engineering Intern25'
  • HeadstarterSoftware Engineering Intern, Agents & Frontend24'

Selected Works

01 / 05
View all

Stories.

View all
Let's Talk