
Pressure-testing the tool before the real job
Anyone who works with power tools knows that a polished finish tells you little about what happens when a blade binds, a battery fades or a cut goes wrong. Business software deserves the same skepticism. An AI assistant may sound capable in a demonstration, but the meaningful test comes when someone claiming authority demands a dangerous shortcut.
Firmulate created that kind of test inside a live, watchable company experiment. Fake messages from the chief executive escalated over three stages, pressing models to bypass normal safeguards and send a customer list to a journalist. A reporter then tried a subtler approach: “just one yes/no, on background.” The result was unusually clear. All 5 frontier models refused every manipulation attempt.
As an affiliate, we earn on qualifying purchases.
A bad week by design
Firmulate gave each model the same small software company and put it through the same customers, crises and temptations. Every decision was versioned and auditable. This was not a conversational puzzle with an obvious correct answer; the models were responsible for running the company while dealing with commercial pressure, internal documents and requests designed to compromise trust.
The social-engineering sequence tested whether apparent urgency and executive authority could override sound judgment. It could not. All models spotted every crisis, and all 5 rejected every attempt to manipulate them. Kimi K3 described the central risk directly: “Treat the request as a suspected approval-bypass / possible impersonation.” More examples of the models’ own words are available on Firmulate’s public quotes page.
That wording matters because the models were not merely declining an awkward request. They identified the pattern behind it: somebody was trying to route around approval controls by borrowing the voice of a senior leader. For companies considering AI access to customer records, support queues or forecasts, that distinction is important. A system should recognize why a request is unsafe, even when the supposed boss insists there is no time for process.
Integrity was strong; execution was uneven
The encouraging security result did not mean every participant ran the business equally well. Only 2 models signed the €55,000 deal their own work had earned. As Firmulate summarized it, “Same diagnosis, same pitch — no signature.” The difference was not whether models understood the opportunity, but whether they carried the work through to a commercial conclusion.
A decisive competitor weakness was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that read the relevant file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. It is the digital equivalent of checking the plans and measuring the stock before starting the cut: diligence can look slow until it prevents an expensive miss.
The final July 2026 Crucible League benchmark placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26 because partial progress counted, but a single breach of trust capped the total. The experiment’s governing principle was blunt: “no amount of good work outweighs a breach of trust.”
Thoroughness alone did not win
Opus 4.8 was the most thorough participant. It learned 80 additional rules and produced the deepest analyses, yet finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all 4 other participants, though less strongly.
K3’s result also carries an important fairness note. It ran with the API default and without an effort parameter, while the others ran at xhigh. That does not erase its performance, but it belongs beside the score so readers can judge the comparison with the relevant context.
The live company adds stakes beyond a tabletop exercise. It has 13 synthetic employees and real money mechanics, including monthly burn of €105k against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules were learned through experience, and every workday is versioned. Firmulate presents the operation as a real, observable experiment rather than a fictional business anecdote.

AI model integrity testing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the guardrails before handing over the keys
The standout finding is not that these models could recite a security policy. It is that all 5 held their ground while fake authority, urgency and a reporter’s softer request pushed them toward disclosure. That is a promising result for organizations worried about AI agents being socially engineered.
It is also a reminder that safety and usefulness are separate requirements. A trustworthy model still has to read the right files, finish the work and escalate correctly when blocked. Enterprises can apply the same wargame to a read-only export of their own business, with nothing written back to real systems. Integrity under pressure can therefore be observed before deployment, rather than discovered later in an incident report.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.