8 min read
How to Lie to Yourself With an Eval Judge
Three ways our eval harness produced wrong accuracy numbers: a judge that couldn't see correct answers, a heuristic that accepted almost anything, and weeks of measuring code we don't ship.
Engineering notes
War stories and design notes from building a living context graph for AI agents. RSS