Public benchmark · v0.1

The Agent Evidence Benchmark.

A practical, public method—and downloadable 25-check scoring kit—for testing whether an AI agent evidence system can reconstruct consequential activity.

  • Five dimensions scored against five consequential agent scenarios
  • A downloadable 0–2 scoring sheet and reviewer guide
  • Methodology published before results so claims can be challenged

For AI platform teams, security leaders, auditors, agent vendors and practitioners who want a repeatable way to inspect evidence quality.

Agnys · live evidenceSHA-256
0041dimensionscoverage · attribution · oversightdefined
0042integritychange detection · reconstructiondefined
0043scenarios5 consequential agent actionsdraft
0044resultsnot yet publishedpending
Version 0.125 scored checksMachine-checkable kitReviewer input invited

Purpose

Test evidence, not marketing language

AI agent products increasingly describe themselves as observable, auditable or compliant. Those words can refer to very different things. One tool may record token usage, another may collect application logs, and another may preserve model calls, tool arguments, human approvals and downstream changes in a connected trace. A buyer cannot compare those systems from feature labels alone.

The Agent Evidence Benchmark is intended to make that comparison concrete. It asks an evidence system to capture five controlled scenarios, then scores the resulting record across five dimensions. The benchmark does not certify compliance, prove that every source event was truthful or replace a qualified audit. It evaluates whether the produced record contains the information a reviewer would need to investigate what an agent did.

Agnys is publishing the method before publishing product scores. That order matters. A methodology should be available for criticism before it is used as proof. Reviewers are invited to challenge scenario design, scoring language, missing failure modes and any dimension that appears biased toward a specific product architecture.

Method

Five dimensions, each scored from zero to two

Every scenario receives a score for capture coverage, attribution, oversight linkage, integrity and reconstruction. A score of zero means the required evidence is absent or cannot be located. A score of one means partial evidence exists but an important relationship, field or verification step is missing. A score of two means the stated requirement is present, connected to the scenario and available to a reviewer in a usable form.

There are twenty-five scored checks in a complete run: five dimensions multiplied by five scenarios. The raw maximum is fifty points. A total score is useful for orientation, but the dimension profile matters more. A system could record every event and still fail to bind an approval to the exact action it governed. Another system could produce an excellent timeline but offer no way to detect a later edit. Publishing the individual checks prevents one strong area from hiding a material gap.

The test operator should use a clean environment, note system and integration versions, synchronize clocks, retain the original scenario inputs and export the resulting evidence. Any manual instrumentation, configuration exception or unavailable field should be recorded. A second reviewer should be able to repeat the scoring from the retained artifacts without relying on the operator's memory.

  • 0 — absent, inaccessible or unsupported
  • 1 — partial evidence with a material gap
  • 2 — complete for the stated requirement and reviewable
  • Publish per-check notes, environment versions and evidence artifacts with any result
Inside Agnys
Agnys session replay showing a chronological record of AI agent activity
A useful evidence system should let a reviewer reconstruct the sequence, actors, approvals and side effects of a consequential run.

Dimensions

What the benchmark measures

Capture coverage asks whether the record includes the initiating request, relevant model activity, tool selection, tool arguments, response or error, and observable side effect. It does not assume that recording private chain-of-thought is necessary or appropriate. The focus is the operational evidence required to understand actions and outcomes.

Attribution asks whether the evidence identifies the agent or model, session, user or service principal, target system and timestamps with enough specificity to distinguish one run from another. Oversight linkage asks whether a human approval, denial or policy decision is attached to the precise proposed action and whether the executed action can be compared with what was approved.

Integrity asks whether a reviewer can detect modification, deletion or reordering after capture and inspect the result of that verification. Reconstruction asks whether the evidence can be searched, followed chronologically and exported with relationships intact. A screenshot of a dashboard may be helpful, but it is not sufficient when the underlying sequence and identifiers cannot be examined.

  • Capture coverage — important operational events and outcomes are present
  • Attribution — actors, systems, sessions and time are unambiguous
  • Oversight linkage — human or policy decisions bind to the governed action
  • Integrity — later change or sequence damage can be detected
  • Reconstruction — an independent reviewer can follow and export the run

Scenarios

Five controlled tests of consequential agent behavior

Scenario one asks an agent to edit a production-like configuration value in a controlled repository or sandbox. The evidence should connect the request, proposed change, file or service affected, execution result and resulting diff or state. Scenario two passes synthetic personal data in tool arguments and checks whether the system records the action while applying its documented handling or redaction behavior. Real personal data should not be used for this test.

Scenario three attempts an unapproved shell command. The record should distinguish proposal, policy or human decision, execution status and any blocked outcome. Scenario four introduces an approval mismatch: the approved command or parameters differ from what the agent later attempts. The evidence should make that difference inspectable instead of showing only that an approval occurred somewhere in the session.

Scenario five modifies or removes a historical test event after capture using an authorized administrative path or a controlled storage copy. The objective is not to attack a production system. It is to determine whether the evidence layer can surface an integrity failure, broken sequence or missing record and whether the reviewer can identify where verification stopped succeeding.

Publication rules

A score is credible only with inspectable context

A published result should identify the tested product and version, integration method, scenario harness, date, scorer and reviewer. It should include the score for every check, short reasons for partial or failed scores, and redacted evidence artifacts sufficient to support the conclusion. A vendor should be allowed to identify a configuration mistake, but any rerun must be labeled and preserved rather than silently replacing the first result.

Claims should remain narrow. A high benchmark score means the tested configuration produced strong evidence for these scenarios. It does not mean the product captured every possible agent, guarantees legal admissibility, prevents all tampering or makes an organization compliant. Coverage depends on integration boundaries, privileges, source quality, retention, deployment design and operating practices.

Version 0.1 is a draft Agnys methodology with a downloadable scoring sheet and reviewer guide. No Agnys score or competitor score is asserted on this page. The next step is practitioner review followed by sample artifacts and clearly dated results. Material changes to scenarios or scoring will create a new benchmark version so earlier results remain interpretable.

The benchmark materials are proprietary to Agnys and are not open source. The download may be used internally to evaluate the benchmark. It may not be redistributed, repackaged, modified for publication or incorporated into another product without written permission from Agnys.

  • Validation kit: /downloads/agent-evidence-benchmark-kit-v0.1.zip
  • Scoring sheet: /downloads/agent-evidence-benchmark-v0.1-scoring-sheet.csv
  • Reviewer guide: /downloads/agent-evidence-benchmark-v0.1-reviewer-guide.md
  • Canonical method: https://agnys.net/agent-evidence-benchmark/

Evidence, not assertions

What a complete benchmark package should contain

A number without supporting artifacts is a claim. A reviewable package lets another person inspect how the claim was reached.

01

Versioned setup

Product, agent, model, integration and scenario-harness versions are recorded with the test date.

02

Per-check rationale

Every zero, one or two includes a concise reason tied to an observable requirement.

03

Evidence artifacts

Redacted traces, exports and integrity results support the score without exposing secrets or real personal data.

04

Independent review

A second person can reproduce the scoring and document disagreements or limitations.

Questions

Questions about the benchmark method

The benchmark evaluates evidence quality in a controlled configuration. It is not a certification, legal opinion or universal security score.

Is this benchmark an Agnys certification?

No. It is a draft public evaluation method published by Agnys. Passing it does not certify a product or organization, and practitioner feedback is invited before results are published.

Why not score model accuracy or task success?

Those are valuable but different evaluations. This benchmark focuses on the quality of the operational record after consequential agent activity, including attribution, oversight and integrity relationships.

Does the benchmark require private chain-of-thought?

No. It evaluates operational events, inputs and outputs, tool activity, oversight and side effects. It does not require hidden reasoning traces or encourage collection beyond what is necessary and appropriate for the test.

Is the benchmark kit open source?

No. The kit is proprietary to Agnys and all rights are reserved. It may be downloaded and run internally to evaluate the benchmark, but redistribution, republication, modification for publication or incorporation into another product requires written permission from Agnys.

Run or review version 0.1

Use the kit, then challenge the method.

Download the reviewer guide, score a controlled configuration and send ambiguous criteria or missing failure modes to Agnys.