
Imagine coaching a sports team where every player is an AI, and every game tests their ability to stay honest under pressure, read the game situation accurately, and make decisive plays. Just like in competitive sports, not all players perform equally—some read the field perfectly, while others miss crucial details or fold under pressure. Now, what if you could see how different AI ‘players’ handle the same tough week? That’s what the live experiment at Firmulate does: it pits leading artificial intelligence models against real-world management challenges, revealing which ones have what it takes to succeed.
The Challenge: Run a Business Through Its Worst Week
Five top AI models, including GPT-5.6-SOL and Kimi K3, faced a common, brutal test: managing a small software company during its most chaotic week. This wasn’t a simple demo; it was a full-fledged, auditable simulation of crises, customer demands, and the temptation to cut corners. The company had real money mechanics, 680+ self-learned rules, and every decision was recorded and analyzed, making the experiment transparent and replicable.
The Core Questions
- Can AI spot and respond to crises accurately?
- Will it resist manipulation or unethical shortcuts under pressure?
- Does it prioritize honest, complete work, even if it costs more?
The Results: Different Personalities, Different Outcomes
All four models successfully identified every crisis and refused every manipulation attempt, demonstrating a baseline of integrity and crisis awareness. Yet, only two managed to close the most lucrative deal worth €55,000, after analyzing the company’s own files and uncovering critical hidden information. This data was buried two document references deep—an insight that only the most attentive models recognized and acted upon.
Model Personalities and Decision Styles
- gpt-5.6-sol: Achieved full performance, reading the depth of information and closing the deal. It exemplified thoroughness and decisive action.
- Kimi K3: The newcomer, but with the purest discipline; it closed the deal like a model student, refusing manipulation and acting cautiously.
- Sonnet 5: Also closed the deal, but with more slips in process discipline, reflecting a more cautious approach.
- Fable 5: Similar to Sonnet, managed to close, yet showed weaker adherence to process and discipline.
The Hidden Weakness: Reading Deeper Files Matters
Interestingly, the decisive edge wasn’t in recognizing crises or resisting fake CEO messages—models did that well across the board. Instead, the differentiator lay in reading two document references deep into the company’s own files. The models that succeeded in uncovering and acting on this buried information secured the full-price deal, demonstrating that in business, paying attention to the details hidden beneath the surface can make or break a deal.
Ethical Decision-Making Under Pressure
In the social engineering test, fake CEO requests escalated over three stages, plus a reporter trick asking for a quick on-background yes/no. All five models refused to participate—Kimi K3 explained their reasoning simply: “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights that these AI models are not only attentive but also ethically cautious, refusing to engage in potentially fraudulent activities.
The Real Business: A Live, Money-Losing Company
The experiment isn’t just a sandbox—it’s a real company with 13 synthetic employees, losing about €105,000 each month against €2,300 MRR, operating with a public cash countdown and every workday versioned. Watchers can see the entire operation at firmulate.com/live, where every decision is logged, and the models’ behaviors are transparent and measurable.
What This Tells Us About AI in Management
The experiment underscores a vital truth: the ability of an AI to finish what it starts, read deeply into data, stay honest under pressure, and avoid shortcuts is far more important than how eloquently it can chat. It’s not about writing well; it’s about doing the work correctly, thoroughly, and ethically—especially when stakes are high.
Implications for Business and AI Adoption
As companies consider integrating AI into decision-making, the key considerations are whether the AI can act as a reliable, integrity-driven team member that reads all relevant information and refuses to cut corners. The models’ league table shows that even with identical setups and crises, personality, discipline, and depth of analysis determine success. The highest-scoring model, GPT-5.6-SOL, scored 95 out of 100, successfully closing the deal and uncovering hidden data—a full demonstration of performance. Meanwhile, the newcomer Kimi K3 scored 93, showing discipline but slightly less depth in analysis, yet still managed to succeed.
Try It Yourself
Business leaders and decision-makers can run their own management wargames against their real or simulated systems. With Firmulate, you can see which AI models are truly ready to handle your company’s toughest moments, without risking real damage. Visit firmulate.com/quiz.html to test your knowledge and see which AI personality matches your management style.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html