Technology

How AI Models Are Being Tested in New Ways

Martin HollowayPublished 2month ago3 min readBased on 4 sources
Reading level
How AI Models Are Being Tested in New Ways

OpenRouter held a tournament on 17 June 2026 where eleven artificial intelligence language models competed against each other in thirty games. The test cost $482 to run and measured how well each model performed when it had to think, adapt, and survive multiple rounds of competition — not just answer single questions correctly.

Traditional AI tests ask models to answer questions on a quiz or complete a writing task. OpenRouter did something different: it put the models in a competitive game where losing players are eliminated and winners face tougher opponents in the next round. This setup tests whether a model can plan ahead, remember what happened before, and keep performing well under pressure — skills that matter in the real world but are invisible on traditional test scores.

Over thirty games, the cost worked out to about $16 per game. That is a real cost, but it is something a typical technology team could afford to do on their own, which means other researchers can repeat the experiment and check the results.

Other companies have also started testing AI models in more realistic conditions. Anthropic, the company behind an AI called Claude, has run several experiments. One test had Claude manage a real office vending machine — deciding what products to stock, what prices to charge, and how to manage money. Another connected Claude to a robot to see how well it could help humans with complex physical tasks. A third worked with a military research agency to have Claude scan millions of lines of computer code to find security problems at speeds humans could never match.

What connects all these tests is the shift from asking "How much does this model know?" to asking "What can this model actually do in a real situation?"

One detail worth noticing: OpenRouter published the exact cost — $482, not something vague like "a few hundred dollars." This is unusual. It lets anyone compare not just which model performed best, but which model gave the best performance for the money spent. A model that wins more games but costs three times as much per game is telling you something important about how expensive it will be to use in practice.

The tournament format also creates something interesting as AI testing evolves. When models compete against each other, weaker ones drop out. The models that survive to later rounds are playing against smarter opponents than the models in early rounds were. This is closer to how AI systems behave in the real world — where they have to compete or adapt in a changing environment — than any simple test can capture.

For teams deciding which AI models to use in their products, the takeaway is straightforward. A model that scores highest on traditional tests might not be the best choice for a task that requires ongoing reasoning and adaptation. The models that perform well in these competitive, real-world-like tests may be better bets for actual use.