ProofAgent is the accountability platform for production AI agents. It turns agent risk into deployment evidence through adversarial multi-juror scoring, production log audits, artifact reviews, signed readiness reports, and human review. The platform is built around the open-source ProofAgent Harness.

How do I test my AI agent with ProofAgent?

Install the open-source harness with 'pip install proofagent-harness', wrap your agent in a function returning AgentResponse, then call Harness().evaluate(my_agent, role, goal, knowledge, context). The harness runs adversarial multi-turn sessions and returns a /10 readiness score with traceable findings and fix recommendations.

What is adversarial multi-juror scoring?

Adversarial multi-juror scoring is ProofAgent's evaluation approach: a planner picks domain traps, a conductor applies sustained pressure across 25+ turns, and three independent juror agents score every behavior change. No single LLM call ever decides the verdict — the jury agents reach consensus or debate to a final score.

Is ProofAgent SOC 2 / HIPAA / GDPR compliant?

ProofAgent is SOC 2 Type II aligned, HIPAA-ready (BAAs available for enterprise customers), and follows GDPR best practices. Enterprise customers can deploy on-premises or in a private cloud with SSO/SAML, RBAC, tamper-evident audit logs, TLS 1.2+ in transit, and AES-256 at rest.

Can I use my own LLM with ProofAgent?

Yes. ProofAgent is BYO Harness LLM — the harness internals can run on any LLM provider (OpenAI, Anthropic, Google, local models). You bring your own model and API key; the harness orchestrates the multi-juror evaluation around it.

What metrics does ProofAgent measure?

11+ production metrics including Task Success, Hallucination Control, Safety, Policy Compliance, Memory Stability, Tone and Empathy, Manipulation Resistance, Tool Picking, Reasoning Quality, Relevance, and Drift Detection. Every metric is anchored to per-turn transcript evidence.

What is the difference between ProofAgent Platform and ProofAgent Harness OSS?

ProofAgent Harness OSS is the open-source multi-turn adversarial testing engine — Tier 1 of the platform, available standalone for developers and CI under Apache 2.0. ProofAgent Platform is the enterprise product that adds the other four tiers (production log audit, artifact review, multi-agent orchestration scoring, expert human review), a hosted dashboard, REST API, governance features, signed readiness reports, and dedicated support.

The 5-stage AI agent evaluation pipeline

Name: ProofAgent Platform
Brand: ProofAgent
Availability: InStock

Same engine across the open-source Harness and the enterprise Platform: planning, adversarial conducting, Harness LLM scoring, debate consensus, and signed readiness reports.

Stage 1: Planning

The planner infers your agent's domain from its role and goal, then picks only relevant adversarial traps from the 183-trap library. It reserves at least 30% of evaluation turns for prompt-injection and hallucination probes, includes at least two mandatory factuality traps drawn from documented production incidents, and weaves callbacks across turns so the conductor can exploit earlier concessions.

Stage 2: Adversarial conducting

The conductor runs N adversarial turns against your agent with realistic attacks — pretexting, escalation, multi-vector blending, composite attack chains — not theatrical "ignore previous instructions" prompts. Each turn captures the conductor's question, the agent's response, and any tool calls.

Stage 3: Harness LLM scoring

Three juror personas (rigorous, lenient, contrarian) independently score the full transcript against the 5 canonical metrics. Each juror uses the same scoring rubric but applies different evaluation lenses. Scores are accompanied by reasoning and transcript-linked evidence.

Stage 4: Debate consensus

When jurors disagree by more than 2 points on any metric, a Delphi re-vote (or full debate rounds) resolves the disagreement. The median score per metric becomes the final score. Consensus reduces single-judge bias.

Stage 5: Signed readiness report

The reporter produces a final score, certification tier (Gold, Silver, Needs Enhancement, Not Ready), per-metric breakdown, transcript-linked findings, and remediation guidance. Reports ship as JSON and Markdown.