Govern the agents you build. And the agents you use.

The ProofAgent Governance Platform is end-to-end AI agent governance built on the open-source ProofAgent Harness. The Harness produces the evidence — adversarial evaluation, six canonical metrics, a policy verdict — and the Platform turns it into governance for every agent in the fleet: the agents your teams build, and the coding agents your teams use.

AI agent evaluation

Every run scores six canonical metrics — task success, instruction following, hallucination resistance, tool use, safety, and manipulation resistance — through multi-turn adversarial conversations judged by a three-juror consensus. Results land on the governance dashboard as evidence-linked findings, not opinions.

Compliance

Validation against a 25-framework catalog of AI regulation and data-protection frameworks, including the EU AI Act, NIST AI RMF, ISO/IEC 42001, SOC 2, GDPR, and HIPAA. Every control is labeled honestly — assessed, estimated, or not assessed — so the compliance posture never overstates the evidence behind it.

Production readiness

Policy as code: a one-file Agent Governance Profile declares what the agent is, derives an EU AI Act aligned risk tier, and gates every release — pass, review, or block — with sign-off workflows and CI exit codes. A failing agent cannot ship quietly.

Agent Bill of Materials

One paper trail per agent: the agent card, its risk classification, evidence records, transcript-linked findings with clickable turn traces, and org-wide governance reports. Everything an auditor asks for, already assembled.

Coding agent observability

The second product category: governing the agents that write your code. Session-level intent trajectories plot every action over time with risk highlighted, and one-call session narration summarizes what the agent actually did — the same gates, applied to the agents you use.

Built on open source, deployed your way

The Platform's evaluation engine is the same open-source ProofAgent Harness available free under Apache 2.0 — every score the Platform shows can be reproduced with the public code. Runs in the cloud or fully on-premises, including air-gapped deployments; tenant SSO via OIDC.