Evidence-linked benchmark report
Automated and human results remain separate. Missing evidence is shown rather than estimated.
Agent Readiness Benchmark
Important: These scores come from deterministic mock runs, not live OpenAI, Claude or Gemini browser agents. They demonstrate the measurement and reporting workflow only.
| Product | Agent Readiness | Success | Median steps | Recovery | Errors/action |
|---|
The product remains the unit of analysis. Provider views expose consistency; they are not an LLM leaderboard.
Failure labels come from structured run records. Live implementations should attach screenshots, replay references and observable page state.
Public customer evidence is secondary and remains separate from agent behavioral results. Quotes are never fabricated.
Confidence labels prevent demo-generated hypotheses from being presented as established facts.
Automated and human results remain separate. Missing evidence is shown rather than estimated.