What 639,000 Execution Steps Taught Me About How AI Agents Really Fail
I applied the MAST failure taxonomy to 639,000 execution steps from AI agents running in production for five months. My first headline finding turned out to be an infrastructure bug masquerading as agent behavior. This post is about what agents actually fail at in production, and the discipline it takes to not fool yourself with production data.