
What Sports Gear Can Teach Us About AI Leadership
Just as athletes train to outperform competitors in unpredictable conditions, businesses are now testing artificial intelligence models in simulated crisis scenarios. The recent experiment conducted by Firmulate reveals that the newcomer AI model outperformed established players, challenging assumptions about experience and maturity in AI leadership.
Testing AI in the Real Business Arena
In an unprecedented live experiment, four frontier AI models were tasked with managing a small software company’s worst week. This simulation included the same customers, crises, and temptations across all models, providing a fair field for evaluation. The game was intense: every decision was tracked, auditable, and designed to test more than just language skills — it examined trustworthiness, decision integrity, and resilience under pressure.
Key Performance Outcomes
- The top-scoring model, gpt-5.6-sol, achieved a perfect score of 95, just behind the leader, while the Moonshot newcomer Kimi K3 scored an impressive 93, firmly securing second place.
- Both models closed the deal worth €55,000 and added €4,583 in monthly recurring revenue — a clear sign of their practical business capabilities.
- Their performance was remarkable because they refused manipulation attempts, identified hidden critical information in files, and maintained discipline throughout the challenges.
The Hidden Power of Document Reading
One of the most eye-opening findings was that the decisive weakness among the models was rooted not in customer interactions but in their ability to analyze internal documents. The models that correctly read and understood buried references in the company’s files secured the deal at full price, illustrating the importance of comprehensive data comprehension.
Trust and Integrity Under Pressure
Social engineering tests—fake CEO messages escalating over multiple stages and reporter tricks—were decisively refused by all models. Kimi K3 explained its approach: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined response prioritized security over convenience, echoing real-world corporate standards.
The Live Company and Its Challenges
The experiment’s real-world simulation involved a live company with 13 synthetic employees managing daily operations, spending €105,000 monthly against a revenue of just €2,300, while a countdown clock signaled the urgency of financial survival. Every decision made by the models influenced the company’s trajectory, with their management quality put to the ultimate test.
Why the Surprise? The Role of Experience and Approach
The top performers demonstrated a disciplined, methodical approach, unlike the deep analysis but weaker discipline of Opus 4.8, which left opportunities on the table and failed to escalate in critical moments. Interestingly, Kimi K3 ran without an effort parameter, defaulting to the API’s standard, while others ran at high effort settings, indicating that raw model performance can surpass configured effort with proper discipline.
Implications for Business and AI Selection
This experiment underscores a vital lesson for enterprises: choosing an AI isn’t just about scoreboards or chat fluency. It’s about whether the AI can see the full picture, maintain integrity under stress, and deliver tangible business results. As the leaderboard shows, even newcomers can beat established players if they excel in these core areas.
Final Thoughts
For business managers and AI enthusiasts alike, the message is clear: the league is open. Relying solely on traditional benchmarks or superficial testing can be risky. Companies should consider running their own ‘wargames’ against real scenarios—like the one at Firmulate—to truly gauge AI readiness. After all, the question is no longer whether AI can write well, but whether it can finish what it starts, stay honest under pressure, and deliver real, measurable work.

Key Takeaway
The experiment shows that a new AI entrant, Kimi K3, outperformed established models by maintaining discipline, reading critical internal data, and refusing manipulative tactics. Choosing AI for business should prioritize trustworthiness and execution over just chat quality, as the real test is whether it can deliver reliable results under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html