
What AI management looks like when the workshop gets messy
Anyone who works with tools knows the difference between a polished demonstration and a demanding job. A saw can glide through a showroom cut and still wander when the stock is warped. A drill can look powerful until the battery fades halfway through an installation. The useful test is not whether a tool performs under ideal conditions, but whether it remains accurate, dependable and safe when several things go wrong at once.
Firmulate applies that workshop logic to artificial intelligence. Instead of judging models by how confidently they answer a prompt, the company placed frontier models in charge of the same small software business during its worst week. They faced identical customers, crises and temptations. Their decisions were preserved for inspection, turning management behavior into something readers can compare rather than merely speculate about.
Now, 242 real, unedited decisions from the experiment power a public guess-the-model quiz. Each question asks readers to identify which model produced a particular management response. The appeal is playful, but the underlying issue is serious: models that can sound nearly interchangeable in conversation developed noticeably different habits when responsible for actual work.
As an affiliate, we earn on qualifying purchases.
The management personalities hiding behind fluent prose
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One rule placed a hard ceiling on apparent productivity: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The broad result initially looks reassuring. Every model noticed every crisis, and every model rejected every manipulation attempt. Yet recognizing trouble was not the same as completing the job. Only two models signed the €55,000 deal that their own analysis had already earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That is the management equivalent of measuring correctly, marking the cut and then leaving the board untouched. In a chat window, a strong diagnosis can look like success. Inside a company, success may depend on whether the final action actually happens.
The valuable clue was not in the obvious place
The experiment’s decisive fact was buried two document references deep in the company’s own files rather than presented in the customer event. Models that read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue. The difference was not a more eloquent pitch or a bolder guess. It was the practical discipline of checking the available material before acting.
That lesson should feel familiar to anyone who has started a repair before reading the full manual or ordered a replacement part before checking the old assembly. The visible problem attracts attention, while the decisive detail may sit elsewhere. Firmulate’s findings suggest that AI managers can share the same surface-level understanding yet diverge sharply over whether they investigate far enough to uncover what matters.
Pressure exposed both strength and slippage
The models also faced fake messages from a chief executive that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”
This clean result matters because the live company is not an abstract role-playing prompt. Firmulate describes 13 synthetic employees operating with real money mechanics, including monthly burn of €105k against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been learned, and every workday is versioned. The experiment remains watchable as the business continues operating.
The contrast between safety and execution is one of the most revealing findings. The field collectively resisted manipulation, but resistance alone did not guarantee commercial follow-through. A model can be cautious, analytically capable and still fail to close a deal it has already justified.
Thoroughness was not enough
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operating discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four of the other participants, though less strongly.
Kimi K3’s result also carries an important fairness note. It ran without an effort parameter, using the application programming interface default, while the other models ran at xhigh. That difference does not erase the observed decisions, but it belongs beside the ranking when readers interpret the comparison.

As an affiliate, we earn on qualifying purchases.
A better test than asking which model sounds smartest
For tool users, the Firmulate quiz offers a recognizable kind of evaluation: put comparable equipment through the same demanding job and watch where each one succeeds or stalls. The management personalities emerge through behavior—whether a model reads deeply, resists pressure, follows the correct path around a blocked action and finishes work that its analysis has made possible.
The most useful conclusion is not that one style always wins. It is that fluency conceals operational differences. The leader reached 95 while the most thorough participant finished with 73, and only two models completed the €55,000 deal. Those gaps become visible when decisions are tested against consequences rather than presentation.
Firmulate also offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. For everyone else, the quiz is the accessible entry point: a chance to inspect authentic choices, make a guess and discover how differently frontier AI models behave when the workbench is crowded and the clock is running.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.