A 42-of-42 failure rate in our LLM benchmark was a parser bug. How we found it, fixed our scorer without fudging results, and ...