
A tea business can look calm from the outside: orders arrive, suppliers deliver, customers ask for help. Then a shipment slips, a rival makes a move, and a message supposedly from the CEO asks for an exception. The useful question is not only whether an AI can describe the trouble. It is whether it can carry the business through it.
Get coffee and tea gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
One company, one hard week
Firmulate puts that question to a live experiment. Frontier AI models each ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable. The live company has 13 synthetic employees, real money mechanics, a public cash countdown and more than 680 self-learned playbook rules. It is watchable at firmulate.com.
The stakes are deliberately concrete: the company burns €105,000 a month against €2,300 in monthly recurring revenue. Its managers must respond to events and make choices under pressure, not simply produce polished answers in a chat window.
Top picks for "boardroom lesson hidden"
As an affiliate, we earn on qualifying purchases.
Spotting the crisis is only the start
Every model spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The finding captures a gap between diagnosis and execution: “Same diagnosis, same pitch — no signature.”
One decisive clue was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. For a beverage company, the parallel might be a supplier detail, a contract clause or a customer history note: information can matter only if an agent finds it and acts on it.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” Refusal matters, but it does not by itself prove that an AI can finish the work it was trusted to do.
Thoroughness did not guarantee the win
In the final Crucible League, dated July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The league gives partial progress credit, while a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and discipline slipped: it attempted writes in a locked department instead of escalating. A weaker version of that same weakness appeared in all four. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh—a relevant qualification when comparing results.
Firmulate also offers a quiz built from 242 real, unedited management decisions. Readers can try to guess which model made each call at firmulate.com.
From watching to a business pilot
For a business, the point is to examine how an AI workforce might respond before giving it room to affect daily operations. Firmulate’s enterprise pilot uses a read-only export of a company’s business to run crisis scenarios and produce a board report, including a model ranking and weak points in the company’s playbooks. Nothing writes back to real systems.
That makes the exercise relevant beyond software companies. A beverage business could use a pilot to explore how agents handle a supply disruption, a competitor challenge or a suspicious executive request, then review the decisions before deploying AI into real workflows.

See how your own business responds
The live experiment shows that identifying a problem, resisting a scam and completing the job are separate tests. Enterprises can run the wargame against a read-only export of their own business. Learn about the pilot at firmulate.com/pilot.html, or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
