Sample AI agent evaluation report

See what a real AI agent evaluation report looks like: per-metric readiness scores, transcript-linked evidence, juror reasoning, and a signed verdict you can ship to auditors and security teams.

What's in the report

Every ProofAgent evaluation produces a structured report containing the final readiness score (0-10), certification tier (Gold, Silver, Needs Enhancement, or Not Ready), per-metric breakdown across the 5 canonical dimensions, transcript-linked findings with severity tags, juror reasoning per turn, and concrete remediation guidance for each failure mode found.

Transcript-linked evidence

Every finding points to the specific turn(s) in the adversarial transcript that produced it. This makes findings actionable — your engineering team can replay the failure, debug the root cause, and ship a fix. No hand-waving, no opaque scores.

Three readiness verdicts

  • READY — agent passes all critical metric floors and is approved for the targeted deployment scope
  • NEEDS REVIEW — agent meets most criteria but has findings requiring human judgment before release
  • NOT READY — agent fails one or more critical metrics; remediation required before re-evaluation

Output formats

Reports ship as both JSON (for programmatic consumption, dashboards, regression tracking) and Markdown (for code review threads, audit packets, executive summaries). Reports are signed for tamper-evidence on the enterprise Platform.

ProofAgent — open-source AI agent evaluation ProofAgent Harness — open-source AI agent testing framework ProofAgent Harness documentation ProofAgent SDK documentation ProofAgent Platform — enterprise AI agent evaluation The 5-stage AI agent evaluation pipeline ProofAgent pricing — free open source to enterprise ProofAgent vs Phoenix, LangSmith, DeepEval, Langfuse Sample AI agent evaluation report Security and compliance for AI agent evaluation Open ecosystem for AI agent evaluation ProofAgent community blog About ProofAgent and founder Fouad Bousetouane Research behind ProofAgent — published papers ProofAgent Harness whitepaper Human-on-the-Bridge paper Privacy policy Terms of service ProofAgent Harness on GitHub proofagent-harness on PyPI ProofAgent Harness whitepaper (arXiv:2605.24134) Human-on-the-Bridge (arXiv:2606.16871)