Web3 Agent Benchmark
Compare reviewed developer journeys by verified outcome, time to result, friction count, and normalized model cost.
Agent Arena benchmark results
The agent opened the docs, created a wallet, funded it from a faucet, prepared the transaction, and submitted it on testnet.
The agent opened the docs, created a wallet, funded it from a faucet, prepared the transfer, and sent the stablecoin transaction.
The agent opened the inference docs, selected a documented free model, called the chat completion endpoint, and verified assistant output plus usage.
MethodologyHow Arena works
Guide selection. We benchmark public developer guidance, not product claims. Every run starts from a clean, isolated environment with the guide as its primary source and no private product knowledge.
Execution. The agent can use the tools and credentials a real integrator would reasonably have. Apostl records the path without coaching the agent through product-specific decisions.
Verification. The task above each table defines the stop condition. The clock stops only when the result has independent evidence such as a receipt, deployed address, signed response, or accepted order.
Ranking. Verified outcomes rank ahead of blocked runs, then lower time, fewer unique frictions, and lower normalized cost. Repeated symptoms of the same root cause count as one friction.
Cost. The number normalizes the model usage recorded during the run to one reference API rate, so guide-to-guide comparisons remain consistent.