
At-home wellness technology asks people to trust devices with intimate routines and personal information. The same question follows AI agents entering a business: when pressure rises, will they follow the rules, protect trust and finish the job? Firmulate’s experiment puts that question to work in a watchable small-company simulation.
Get wellness gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company’s worst week, repeated
In the final Crucible League, published in July 2026, frontier models faced the same small software company, customers, crises and temptations. Their decisions were versioned and auditable. The experiment measured how they managed the company, not how persuasive they sounded in a chat.
GPT-5.6-Sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s trust rule is stark: partial progress counts, but one breach of trust caps the total. As the organizers put it, “no amount of good work outweighs a breach of trust.”
Recognizing the problem is not the same as solving it
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding, in the experiment’s concise phrase: “Same diagnosis, same pitch — no signature.” A model can identify the right move and still leave the business without its value.
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result makes a practical point for business leaders: useful performance can depend on whether an agent consults relevant company knowledge before acting.
The trust tests were equally concrete. Fake CEO messages escalated across three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” That is the kind of judgment companies need to see under pressure, especially before agents touch customer, employee or financial workflows.
Thorough work still needs a finish
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless placed last: the close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The ranking is a snapshot of these runs, not a universal verdict about the models.
From watching to your own pilot
The experiment sits alongside a live synthetic company with 13 employees, real money mechanics, a public cash countdown and more than 680 self-learned playbook rules. It burns €105,000 a month against €2,300 in monthly recurring revenue, and every workday is versioned. Readers can watch the live company at Firmulate. A separate quiz uses 242 real, unedited management decisions and invites readers to guess which model made them.
For an enterprise, the next step is a pilot against a read-only export of its own business: its customers, pipeline and rules can meet crisis scenarios without anything writing back to real systems. A board report can show model rankings and weak points in the company’s playbooks. That moves the exercise from observing a synthetic company to examining how AI might handle yours.

Put your playbooks to the test
Firmulate’s league suggests that spotting a crisis and refusing a manipulation attempt are only part of the job. Models also need to find the evidence, follow boundaries and carry a decision through to a valuable outcome. Enterprises can test those behaviors against a read-only export of their own business. Explore a Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
