
Astrophotography rewards preparation. Before a night under dark skies, you check the forecast, power, tracking and backup plans, because small failures can cost a rare opportunity. Businesses face their own version of a lost clear night: a customer crisis, a tempting shortcut or a decision that never gets made. Firmulate asks what happens when AI runs into that kind of pressure before it is trusted with real work.
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
A company under the same sky
Firmulate’s Crucible League put frontier models in charge of the same small software company through its worst week. The customers, crises and temptations were held constant, so the experiment could reveal how each model handled management decisions rather than simply how well it talked. Every decision was versioned and auditable.
In the final league, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
Seeing the problem was not enough
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate’s summary gets at the managerial gap: “Same diagnosis, same pitch — no signature.” Recognizing a good course of action and carrying it through are different tests.
The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that useful business context may be scattered across records, much as the clue to a better image can sit in a calibration note rather than in the frame itself.
Pressure, discipline and the missed close
The social-engineering tests escalated through three stages of fake CEO messages, then added a reporter’s appeal: “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. Thorough analysis matters, but it cannot substitute for sound boundaries and follow-through.
There is a fairness detail behind the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The experiment also offers a window into everyday management choices. Its quiz at Firmulate draws on 242 real, unedited management decisions and invites readers to guess which model made them.
From watching to a company-specific pilot
The live Firmulate company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. Its playbook has learned more than 680 rules, and each workday is versioned. Readers can follow the experiment at Firmulate, where the live company and league make the decisions watchable.
For an enterprise, the next step is a wargame built around its own business. Firmulate says a pilot can start from a read-only export, then run crisis scenarios against that company and produce a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. That offers leaders a chance to examine how an AI workforce handles their customers, rules and pressure before giving it a role in live operations.

Test before deployment
A model can identify a crisis, resist manipulation and still fail to close a deal or respect an operational boundary. Firmulate’s experiment makes those outcomes visible in a live company; a pilot applies the exercise to an enterprise’s own business using read-only data. To discuss a pilot, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
