Grading an AI agent only on its final response is like grading code solely on whether it compiled. If the agent took five hallucinated steps, queried unauthorized databases, and stumbled onto the right answer, it is a production failure. We examine trajectory auditing.
Step-level validity, state-transition correctness, and tool-call safety auditing.