Human-on-the-Bridge: Scalable Evaluation for AI Agents

Human-on-the-Bridge (arXiv:2606.16871) by Fouad Bousetouane is the research paradigm behind the open-source ProofAgent Harness. Instead of putting a human reviewer in the loop on every agent run, it curates human expertise once — upstream — into reusable traps, juror personas, and scoring rubrics, then lets small evaluator models stress-test frontier-grade AI agents at scale.

Not a human in the loop — a human on the bridge

A captain on the bridge sets the course, the rules, and the instruments rather than steering every wave. Human-on-the-Bridge applies the same separation to AI agent evaluation: experts invest judgment once into the evaluation machinery, and that machinery then runs automatically on every agent and every release — rigorous because it carries real expertise, scalable because it runs locally and cheaply without a reviewer babysitting each transcript.

What the paper contributes

  • Human expertise curated upstream into reusable evaluation intelligence: traps, juror personas, and scoring rubrics
  • Small evaluator models reliably assess agents built on frontier LLMs when armed with pre-curated expert judgment
  • Multi-turn adversarial pressure surfaces policy drift, manipulation, and tool-call failures that static benchmarks miss
  • Multi-juror consensus through debate or Delphi reduces single-model judging bias
  • Every mechanism is reproducible and open-source under Apache 2.0 in the ProofAgent Harness

Headline finding

Production-grade agents built on frontier LLMs (GPT 5.5, Claude Opus 4.7) fail under sustained, composite adversarial pressure — and a small, local evaluator model armed with pre-curated expert judgment can reliably catch it. Frontier models alone are not enough; the agent layer needs its own evaluation infrastructure, and that infrastructure does not have to be as large as the agent it tests. The companion core whitepaper (arXiv:2605.24134) documents the full pipeline and experimental results.

Plain language explainer on Medium: Human-on-the-Bridge (HoB): The New Paradigm for Scaling AI Agent Evaluation.

ProofAgent — open-source AI agent evaluation ProofAgent Harness — open-source AI agent testing framework ProofAgent Harness documentation ProofAgent SDK documentation ProofAgent Platform — enterprise AI agent evaluation The 5-stage AI agent evaluation pipeline ProofAgent pricing — free open source to enterprise ProofAgent vs Phoenix, LangSmith, DeepEval, Langfuse Sample AI agent evaluation report Security and compliance for AI agent evaluation Open ecosystem for AI agent evaluation ProofAgent community blog About ProofAgent and founder Fouad Bousetouane Research behind ProofAgent — published papers ProofAgent Harness whitepaper Human-on-the-Bridge paper Privacy policy Terms of service ProofAgent Harness on GitHub proofagent-harness on PyPI ProofAgent Harness whitepaper (arXiv:2605.24134) Human-on-the-Bridge (arXiv:2606.16871)