06
Featured essay
Evaluation
Evals Are Bullsh*t
Most evals measure whether an agent looks good. Good evals measure whether it does the job your business actually needs.
August 1, 2026
6 min read
Technical writer & builder
Field notes / 2026
I write about agents, systems, and how intelligence becomes useful in the real world.
06
Featured essay
Evaluation
Most evals measure whether an agent looks good. Good evals measure whether it does the job your business actually needs.
August 1, 2026
6 min read
Selected writing
About / Now
Building, studying, and writing in public.
Background
I'm a builder and tinkerer. Previously on the Agent team at OpenAI. On leave from CS at the University of Waterloo, KP fellow.
Current questions
Right now I'm exploring multi-agent systems, long-horizon agents, and continual learning.
The common thread is simple: how do capable systems become useful, dependable, and better with experience?
Contact
The fastest way to reach me is a DM on X or a note on LinkedIn. Interesting questions are always welcome.