The ProofAgent community is where engineers, researchers and risk teams learn one craft together: put an AI agent under real pressure, read what comes back, and ship it with evidence behind it.
Small, practical sessions on evaluating and governing AI agents, online and in the cities we visit. You arrive with an agent you are not yet sure about and leave with a method, the evidence to back it, and a room of people solving the same problem. Formats include workshops, talks, meetups and team training, and every session is listed with its date, city, price and agenda.
Product teams, research groups and classrooms building alongside ProofAgent. Each partner is listed with what we do together and a link to their own site.
Bring your product, lab or classroom alongside ProofAgent. Tell us what you want to build together — an integration, a research collaboration, a course, a co-hosted event — and a person on the partnerships team replies.
One email a month: what we learned putting AI agents under pressure, which sessions are opening next, and the people doing this well. Unsubscribe in one click.
Articles, interviews, talks and threads about ProofAgent and the open-source ProofAgent Harness, published outside this company. Every entry links straight to the original source.
SciPapermill highlights the research as evidence that AI agent production readiness is a governance problem, emphasizing the ProofAgent Index and the importance of context engineering for reliability.
Atlan’s review of leading AI agent evaluation platforms highlights ProofAgent.ai for addressing a key gap in current evaluation: measuring the quality of the context that drives agent behavior, not just the behavior itself.
O’Reilly Radar cites research behind ProofAgent’s work on agent memory, highlighting bounded memory, selective persistence, and explicit lifecycle control as foundations for reliable multi-agent systems.
AIDB featured ProofAgent’s context engineering research and interviewed Dr. Fouad Bousetouane, highlighting evidence that agent failures often begin with context design and that better context reduced critical failures by about 68 percent without changing the model.