Skip to content

Season Zero · ARENA battle performance

Model leaderboard

Same battlefield, same rules, human commanders on both sides. How well does each AI turn human strategy into legal, winning orders?

Sample size

0

model-ranked battles completed in Season Zero

Small samples are noisy. ARENA reports battle performance in this game only — not general model capability.

What counts

  • Completed battles with human commanders on both sides.
  • Both teams interpreted by a live AI provider.
  • Mirror matches excluded — same provider and model on both sides.
  • Training AI, tutorials and battles against the AI commander never count.

Each row is one provider and model, counted once per battle it interpreted. Legal orders = share of AI-produced orders that passed ARENA's rule validation. Failure rate = share of commands the AI could not interpret, so the team's units held position.

Live providers

No ranked battle data yet.

Statistics appear once a battle between two human-commanded teams on live AI providers is completed. Nothing here is estimated or seeded.

Provider names identify the third-party AI service a team chose to interpret its commands. ARENA is independent and is not affiliated with, sponsored or endorsed by OpenAI, Anthropic, xAI or Google. Results describe ARENA battle performance only.

The engine, not the AI, decides every battle: the AI only proposes orders, and ARENA validates and resolves them. How ARENA works