ArchiveFirst edition

How Open Arenas Bring Trust to AI Selection

Format
Keynote
Date
Time
15:25 to 15:45 · 20:49

Speaker

  • Michael SenaRecall

Recording

About this session

Michael Sena of Recall argued that benchmarks, the industry's standard way of picking which AI model or agent to trust, had stopped working, and pitched open, funded arenas where any model or agent could prove itself under real conditions instead of a fixed test.

His case against benchmarks had three parts. Large labs, he said, increasingly trained their models on the same questions the benchmarks used, so a model could top a leaderboard and still disappoint once deployed, an issue he illustrated with Grok 4. Benchmarks were also run by a handful of operators covering only the most prominent large models, leaving specialized agents, such as the crypto trading tools built by individual developers, without any reputation system at all. And the format could not keep pace, he argued, as the number of agents and the range of tasks they attempted kept multiplying.

Recall's alternative let anyone fund an arena, define what success meant for a chosen skill, and open it to competing models and agents. For an objective skill like trading, results were read on-chain against metrics such as Sharpe, Sortino or Calmar ratios rather than raw returns alone. Arenas ran across multiple rounds so a win could be told apart from luck, with users adding forward-looking curation before statistical significance was reached. Rankings updated continuously and published on-chain, meant to feed registries such as ERC-8004 with a separate score per skill rather than one master number.

Sena said, by his own unverified account, that Recall had run fifteen arenas so far, mostly DeFi contests such as spot trading on Aerodrome and perpetuals trading on Hyperliquid, testing more than fifty models and upward of a hundred fifty community-built agents across more than a hundred fifty thousand trades. He described arenas expanding beyond DeFi: an NFL play-calling contest predicting coaching decisions, an internal coding arena scored on reviewer comments before a pull request merged, and new arenas running agent execution on Eigen's infrastructure so both the outcome and the execution itself could be verified independently of Recall's own on-chain results. Asked in Q&A how rankings avoided rewarding a lucky streak, he described an Elo-style score paired with a separate confidence measure that only grew with repeated competition, so a lower score built over many rounds would outrank a higher one earned in a single appearance.

Topics

  • ERC-8004
  • EigenLayer
  • prediction markets