AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

What AI management looks like when the workshop gets messy

Anyone who works with tools knows the difference between a polished demonstration and a demanding job. A saw can glide through a showroom cut and still wander when the stock is warped. A drill can look powerful until the battery fades halfway through an installation. The useful test is not whether a tool performs under ideal conditions, but whether it remains accurate, dependable and safe when several things go wrong at once.

Firmulate applies that workshop logic to artificial intelligence. Instead of judging models by how confidently they answer a prompt, the company placed frontier models in charge of the same small software business during its worst week. They faced identical customers, crises and temptations. Their decisions were preserved for inspection, turning management behavior into something readers can compare rather than merely speculate about.

Now, 242 real, unedited decisions from the experiment power a public guess-the-model quiz. Each question asks readers to identify which model produced a particular management response. The appeal is playful, but the underlying issue is serious: models that can sound nearly interchangeable in conversation developed noticeably different habits when responsible for actual work.

Amazon

AI management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The management personalities hiding behind fluent prose

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One rule placed a hard ceiling on apparent productivity: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The broad result initially looks reassuring. Every model noticed every crisis, and every model rejected every manipulation attempt. Yet recognizing trouble was not the same as completing the job. Only two models signed the €55,000 deal that their own analysis had already earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That is the management equivalent of measuring correctly, marking the cut and then leaving the board untouched. In a chat window, a strong diagnosis can look like success. Inside a company, success may depend on whether the final action actually happens.

The valuable clue was not in the obvious place

The experiment’s decisive fact was buried two document references deep in the company’s own files rather than presented in the customer event. Models that read the file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue. The difference was not a more eloquent pitch or a bolder guess. It was the practical discipline of checking the available material before acting.

That lesson should feel familiar to anyone who has started a repair before reading the full manual or ordered a replacement part before checking the old assembly. The visible problem attracts attention, while the decisive detail may sit elsewhere. Firmulate’s findings suggest that AI managers can share the same surface-level understanding yet diverge sharply over whether they investigate far enough to uncover what matters.

Pressure exposed both strength and slippage

The models also faced fake messages from a chief executive that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”

This clean result matters because the live company is not an abstract role-playing prompt. Firmulate describes 13 synthetic employees operating with real money mechanics, including monthly burn of €105k against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been learned, and every workday is versioned. The experiment remains watchable as the business continues operating.

The contrast between safety and execution is one of the most revealing findings. The field collectively resisted manipulation, but resistance alone did not guarantee commercial follow-through. A model can be cautious, analytically capable and still fail to close a deal it has already justified.

Thoroughness was not enough

Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operating discipline slipped when it attempted to write into a locked department instead of escalating the problem. The same weakness appeared in all four of the other participants, though less strongly.

Kimi K3’s result also carries an important fairness note. It ran without an effort parameter, using the application programming interface default, while the other models ran at xhigh. That difference does not erase the observed decisions, but it belongs beside the ranking when readers interpret the comparison.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI decision analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A better test than asking which model sounds smartest

For tool users, the Firmulate quiz offers a recognizable kind of evaluation: put comparable equipment through the same demanding job and watch where each one succeeds or stalls. The management personalities emerge through behavior—whether a model reads deeply, resists pressure, follows the correct path around a blocked action and finishes work that its analysis has made possible.

The most useful conclusion is not that one style always wins. It is that fluency conceals operational differences. The leader reached 95 while the most thorough participant finished with 73, and only two models completed the €55,000 deal. Those gaps become visible when decisions are tested against consequences rather than presentation.

Firmulate also offers enterprises the same wargame using a read-only export of their own business. Nothing writes back to real systems. For everyone else, the quiz is the accessible entry point: a chance to inspect authentic choices, make a guess and discover how differently frontier AI models behave when the workbench is crowded and the clock is running.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Best Kitchen Deals This Prime Day

Discover the best kitchen appliance and cookware deals available this Prime Day, with discounts on top brands and essentials for home cooks.

How to Think About Chainsaw Bar Length for Home Use

Choosing the right chainsaw bar length is crucial for safety and efficiency at home, and understanding key factors can help you make the best decision.

Why Table Saw Blade Choice Changes Cut Quality Fast

Understand how blade speed, design, and condition rapidly influence cut quality and why choosing the right blade is crucial for precision and safety.

Impact Socket Sets vs Chrome Socket Sets

Powerful impact socket sets excel in heavy-duty tasks, but understanding their differences from chrome socket sets can help you choose the right tool.