FREE: Talk to our live AI audit agent.
SYSTEMS

Benchmark methodology · v1

Value, with evidence.

The Agentic 100 is a curated benchmark, not a popularity chart. A workflow must clear the operating gate first. It is then rewarded for expected business value, its likelihood of working in practice, and the strength of the evidence behind it.

Reward = weighted operating value × evidence confidence − unmitigated risk.

Stage I · Eligibility

The operating gate.

A high projected return cannot compensate for a workflow that is vague, unsafe, or not meaningfully agentic.

  1. 01A named trigger, accountable owner, and observable success metric
  2. 02Enough specificity for a competent operator to reproduce it
  3. 03A genuine observe → decide → act loop, not a prompt dressed up as a system
  4. 04Human approval for irreversible, sensitive, financial, legal, or customer-facing actions
  5. 05Source evidence and explicit failure modes

Stage II · Reward

What earns a place.

  • Business value

    30%

    Does it materially move revenue, margin, risk, customer outcomes, or execution capacity?

  • Frequency

    15%

    Does the underlying work recur often enough for the system to compound?

  • Agency fit

    15%

    Can an agent observe, decide, act, and return an exception or result inside explicit boundaries?

  • Reliability

    15%

    Are the inputs, rules, checkpoints, and failure paths stable enough to operate?

  • Time to value

    10%

    Can a team reach a measurable first outcome without a transformation programme?

  • Adoption likelihood

    10%

    Does it fit the tools, incentives, ownership, and review habits people already have?

  • Portfolio novelty

    5%

    Does it add a useful operating pattern rather than duplicate a stronger entry?

Confidence & risk

Claims get discounted.

An elegant workflow with no operating evidence cannot outrank one that repeatedly returns value. Unmitigated privacy, security, legal, financial, or irreversible-action risk removes up to twenty points; a hard safety failure removes the workflow entirely.

×1.00

Reproduced with measured outcomes

×0.85

Running internally with directional evidence

×0.70

Corroborated by credible implementations

×0.50

Editorial estimate awaiting reproduction

Stable core, flexible lenses.

The global benchmark keeps one safety and evidence standard. Industry, company-stage, and function lenses may adjust value, frequency, and adoption by at most five percentage points each. They do not weaken safety, reliability, or evidence requirements. That keeps the benchmark comparable without pretending the same workflow has equal value everywhere.

This first edition is curated against the published rubric. Individual scores will become public only after the evidence records are backfilled and reviewed; we will not manufacture precision from editorial estimates.

Explore The Agentic 100 →