EIO-Agents: the semantic layer for agent evaluation

EIO-Agents is a framework-agnostic semantic layer for AI-agent evaluation. It uses the Evaluation Intelligence Ontology (EIO) to normalize an evaluation framework's outputs into standardized, evidence-linked claims and a Portable Evaluation Record (PER). The package also validates, verifies and explains the record.

Where EIO-Agents fits

  1. Agent infrastructure: agent, tools, context and runtime
  2. Evaluation framework: ProofAgent Harness is one example, not a requirement
  3. EIO-Agents semantic layer: Evaluation Intelligence Ontology (EIO) and standardized evidence-linked claims
  4. Portable Evaluation Record (PER): versioned output

How the workflow works

A producer supplies observed turns, tool results, source context and evaluator signals. EIO-Agents projects claims and cited evidence into a versioned, content-addressed PER; it includes scores only where the source inputs support them. A consumer validates the PER structure, independently verifies its proof and digest against the bundle, and explains the release recommendation or a finding. Missing evidence can withhold a score or readiness.

Try the included synthetic evaluation

From the EIO-Agents repository, use Python 3.10 or newer. The source bundle and generated PER remain local; the test score is not a safety certification.

git clone https://github.com/ProofAgent-ai/eio-agents.git
cd eio-agents
python3 -m venv .venv
.venv/bin/pip install -e .
.venv/bin/eio-agents project tests/data/native/v0_6/source-complete.bundle.json -o native.per.json --jcs native.per.jcs
.venv/bin/eio-agents validate native.per.json
.venv/bin/eio-agents verify native.per.json --bundle tests/data/native/v0_6/source-complete.bundle.json
.venv/bin/eio-agents explain native.per.json --list

PyPI installation will be linked after package publication. See the source repository for the full quick start.

EIO-Agents structures and checks an evaluation; it does not run the evaluation or certify the agent.

EIO 0.6.0 ontology: claims, evidence and decisions

The Evaluation Intelligence Ontology gives evaluators shared names and rules for claims, evidence, risk, decisions and release views. Its public manifest pins 39 YAML files covering core concepts, risk and domain modules, mappings, governance, scoring, reliability and explanation templates. Exact module hashes and an ontology digest identify the bytes a PER used.

Claims and predicates

A claim applies a versioned predicate to specific turns, records a decision state and cites evidence. Findings, scores, control views and the release recommendation derive from claims rather than adding new facts about agent behaviour.

Evidence and proof

Evidence refs point into the evaluation archive. Their type, source, compatible pairing and anchor determine whether they witness agent behaviour; a user message or retrieved text alone cannot. Deterministic or identified human decisions also need a witnessing citation and exact fidelity or confirmed recurrence to be PROVEN under EIO rules. Semantic jury findings support review but are not proof by themselves.

Framework mappings describe evidence relevance from an evaluation, not legal conformity or certification.

Framework and regulatory mappings

The ontology bundles 30 EIO-Agents mapping targets across regulations, standards, frameworks and security taxonomies. These titles come from the versioned EIO framework registry; they are not the ProofAgent Harness framework catalogue.

Scope: All 30 mappings are provisional evidence-relevance links only. Listing a target does not assert legal applicability, compliance, certification, audit attestation or endorsement by the framework owner. Canada AIDA is classified as a proposal in this EIO release catalogue; this page does not determine its current legal status.

AI laws, regulations and proposals (5)

Privacy and data protection (11)

Sector rules and guidance (3)

AI risk frameworks (1)

AI management standards (1)

Security standards and frameworks (4)

Assurance criteria and industry standards (2)

Security risk taxonomies (3)

PER 2.0.0: what the portable record contains

The schema requires 14 top-level blocks: header, provenance, subject, scope, evidence, claims, coverage, findings, controls, scores, reliability, release_recommendation, limitations and telemetry. It carries cited claims and derived views, not the entire raw bundle.

For an example dashboard-style view, interpretation of readiness versus the release recommendation, and what is or is not an industry standard, read the Portable Evaluation Record guide.

header and provenance identify versions and sources; evidence and claims locate and decide observations; findings, controls and reliability are claim-derived views; limitations and telemetry bound interpretation. Canonical PER bytes have a content digest that can be checked when the source bundle is supplied.

Q, E, C, G and readiness

The four score axes are Q (context), E (behaviour), C (compliance) and G (governance). Metric, axis and readiness values may be withheld when required source inputs are missing. The separate claim-derived release recommendation is PASS, REVIEW or BLOCK, not a guessed grade.

Verification and privacy boundaries

Validation checks PER structure. Verification additionally checks cited evidence and scores against the local bundle, re-derives the record and compares digests. This checks consistency, not whether the original run or evaluator judgment was correct.

The raw source bundle stays local. PER refs can use hashes and fingerprints instead of copying protected text, but review a PER before sharing it. Local explanation or resolution with the bundle may reveal withheld text; keep that output under the bundle's access controls.

Versioned EIO 0.6.0 and PER 2.0.0 source files

Producers and validators use these JSON Schemas, JSON-LD context and ontology modules to interpret the same concepts and PER fields. Select the version named by your package or record. Each released URL identifies immutable bytes; the Python package embeds offline copies. Structural schema validation alone does not prove that evidence is genuine or that an agent is production-ready.

JSON Schema files declare $schema: https://json-schema.org/draft/2020-12/schema; “draft” names the JSON Schema standard version, not the EIO 0.6.0 release status.

EIO-Agents, including its ontology and schemas, is maintained by ProofAgent / ProofAI LLC under Apache-2.0. Read the LICENSE and NOTICE, browse the source repository, or contact support@proofagent.ai.