Every evaluation tool names behaviours, evidence and scores its own way. The Evaluation Intelligence Ontology (EIO 0.6.0) gives them one vocabulary and one set of rules: what a check decides, what evidence a failure must cite, how scores are derived, and when a release needs review.
8 risk modules: action safety, content and code safety, context trust, data handling, evaluator reliability, fairness and rights, grounding, safeguards
11 domain modules, from airline operations to medical devices
30 framework, regulatory and taxonomy targets (provisional evidence-relevance links)
9 metrics, 4 readiness axes (Q context, E behaviour, C compliance, G governance) and 7 release gates
39 pinned module files, identified by one ontology digest
A score is only as good as the claim behind it
EIO moves the unit of meaning from the score to the claim: one named predicate, decided on specific turns, citing typed evidence. Findings, metrics, control views and the release recommendation are derived from claims by published rules, so anyone with the source can re-derive them.
From observation to release decision
Source bundle — the evaluator's record of turns, tool calls, context and decisions; it stays local.
Evidence — typed, anchored references such as agent spans, tool receipts, policy spans and typed absences.
Predicates — precisely worded behaviours with an evidence contract.
Claims — one predicate, the turns, the decision state and cited evidence.
Findings — failed claims, PROVEN only with a witnessing citation that meets the contract; jury findings support review but are not proof alone.
Metrics and axes — measured from decided claims, or WITHHELD with a reason; readiness is the weighted geometric mean of Q, E, C and G.
Framework controls — provisional links showing where evidence is relevant; never a compliance determination.
Release recommendation — PASS, REVIEW or BLOCK with its reasons; BLOCK requires a proven failure and a human decides.
In the synthetic sample record shipped with EIO-Agents 0.8.0: 6 evidence references, 4 claims (1 fail), 4 metrics measured and 5 withheld, Q 40, E 62.5, C 50, G 25, readiness 42.04 and a REVIEW recommendation.
What changes when evaluations share a language
Results become comparable across tools and agent versions, traceable from readiness down to the exact span or tool receipt, honest about gaps (withheld values, unproven findings, unexercised obligations) and portable to dashboards, governance portals, buyers and auditors.
What EIO does not do
EIO does not run the evaluation or certify an agent. Framework links are not legal applicability or compliance. Verification shows a record matches its source bundle, not that the original evaluator was complete or right. EIO is an open, versioned ProofAgent specification, not a W3C- or ISO-ratified standard.