SWE-bench end-to-end testing reveals if an AI agent succeeds at completing tasks across dozens of tool calls, moving beyond ...