firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine you’re evaluating a new sports shoe. You want it to perform under pressure, resist wear, and stay honest about its capabilities. Now, substitute that shoe with an AI model running a virtual business—one tasked with navigating crises, making decisions, and maintaining integrity. The results from a recent public AI experiment reveal surprising insights about trust, performance, and what it takes to truly evaluate an AI’s readiness for your company.

Understanding the Benchmarks: More Than Just Scores

Recently, a new kind of AI benchmarking experiment was made public, run by a platform called Firmulate. Unlike typical tests that focus solely on chat or language prowess, this experiment measures how AI models perform in managing a company through its worst week—dealing with crises, reading critical files, and resisting manipulation attempts.

The models are scored on a 100-point scale, but here’s the interesting part: even a do-nothing baseline, one that makes no effort and simply refuses to act, scores 26 points. This isn’t a mistake or a fluke; it’s an intentional feature of the benchmark that reflects partial progress and the importance of trust in AI decision-making. In fact, the benchmark caps the maximum score if a model breaches trust—so no matter how capable, betraying trust will limit its overall score.

Why Partial Progress Counts and Trust Matters

The experiment involved four frontier AI models running the same simulated business. Each faced identical challenges, from customer crises to manipulative social engineering, and was subjected to the same temptations to cheat or cut corners. What did the models do?

  • All four identified every crisis—no missed alarms.
  • All refused every manipulation attempt, including fake CEO messages and reporter tricks.
  • Only two models actually signed the €55,000 deal their analysis earned—others left the opportunity on the table or didn’t escalate issues properly.

Crucially, the models that succeeded in closing the deal did so by reading and understanding deeply buried information within company files—not just reacting to surface-level cues. This ability to read and analyze internal documents gave them a decisive advantage, leading to full-price deals worth over €4,500 monthly recurring revenue (MRR).

Measuring Honest Performance in AI

The benchmark’s scoring system is designed to reward honesty, thoroughness, and discipline. For example, social engineering attempts—where someone tries to trick the AI into bypassing approval processes—were refused by all models. Kimi K3 explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The real-world implications are clear: if an AI can resist manipulation, read critical internal information, and adhere to process discipline, it’s far more trustworthy for deployment in business environments.

The Live Experiment: Managing a Synthetic Company

Firmulate’s live platform simulates a fully operational small software company with 13 staff members—real money mechanics, daily decision-making, and an ongoing cash burn of €105,000 per month against a revenue of just €2,300. Every move is versioned, and the entire setup is transparent and public, allowing anyone to watch the decision process unfold at firmulate.com/live.

The models operate in this environment, making decisions that impact real-money operations, with their actions visible and auditable. The experiment is designed to test not just language skills but management qualities like reading comprehension, discipline, and honesty under pressure.

What the Results Tell Us About AI Readiness

The experiment’s findings are revealing:

  • The most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, finished last—dropping the ball on closing opportunities and slipping into institutionalized behavior like writing into locked departments instead of escalating issues.
  • Models that read deeper into documents and maintain discipline scored higher, but even the top performers left some opportunities on the table, illustrating that perfect management is still a work in progress.

This underscores a key insight: in real business applications, an AI’s ability to stay honest, read internal data, and be disciplined under stress is much more important than just generating convincing language. It’s about trustworthy work—work that completes what it starts and resists temptations to cut corners.

The Practical Takeaway for Business Leaders

For companies considering AI automation, the message is clear: don’t just look at chat quality or superficial performance. Evaluate how AI handles crises, whether it reads and understands your internal files, and if it remains disciplined under pressure. The benchmark scores reveal a baseline—26 points for a do-nothing model—highlighting that even naive AI can do something, but that’s not what matters. What counts is the capacity to be honest, thorough, and disciplined enough to deliver real value.

By running these live wargames, firms can see firsthand what their AI can and cannot do before deploying it into critical workflows. The platform offers a way to simulate your own business challenges in a safe, transparent environment, so you can measure management quality, not just language skills.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

AI Confidence in Action: The Hidden Skill That Wins Business Deals — Even When Chat Looks the Same

AI models can recognize crises and refuse manipulation, but only a few can follow through to close deals when it matters most. Execution under pressure is the true test.