
Imagine a small coffee shop managing daily crises — from supplier delays to customer complaints — all without a single human employee. Now, scale that idea to a high-stakes software company, run entirely by AI models, battling to stay afloat. This is not fiction; it’s the live experiment by Firmulate, where artificial intelligence is put to the test as a fully functioning, decision-making business.
The Live Business in Action
Introducing a real, publicly accessible experiment where 13 synthetic employees collaborate within a small software company. Every workday, the company’s decisions are documented, versioned, and openly observable at firmulate.com/live.html. This setup isn’t just theoretical — it reflects genuine money mechanics, with the company burning through €105,000 each month against a modest revenue of €2,300 MRR, and a public cash countdown ticking down every day.
What makes this experiment extraordinary is its commitment to transparency: the company runs 680+ rules learned through self-guided playbook mechanics. Every decision is auditable, and each model’s performance is scored against a league table based on their success in navigating crises and closing deals.
As an affiliate, we earn on qualifying purchases.
The AI Models: Different Strategies, Same Challenges
Four frontier AI models — including GPT-5.6-sol, Kimi K3, Sonnet 5, and Fable 5 — faced the same week of turmoil, with identical customer issues, crises, and temptations to cheat or manipulate. The goal was simple but revealing: could these models manage the crisis, uphold ethical standards, and close a €55,000 deal they identified as achievable?
The results were telling. All four models accurately identified every crisis and refused every manipulation attempt, such as fake CEO messages or subtle bribes. The kicker? Only two managed to close the deal, despite all of them diagnosing correctly and delivering the same pitch. The other two, including Fable 5, left the opportunity unclaimed — a failure to follow through despite recognizing the opportunity.
The Hidden Weakness: Reading the Right Files
Deep inside the company’s files, two document references held the key to sealing the deal. Models that actually read and understood these references successfully closed at full price, adding roughly +€4,583 to their monthly recurring revenue. This indicates that the secret to better AI decision-making isn’t just in recognizing crises but in accessing and interpreting crucial contextual information hidden in documents.
Resisting Social Engineering and Manipulation
Throughout the week, the company faced staged social engineering attacks, including fake CEO messages escalating in complexity and a reporter trick asking for a quick ‘yes’ or ‘no.’ All five models tested refused to be manipulated, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This resilience underscores the importance of ethical safeguards in AI decision-making, especially in high-stakes business scenarios.
The Reality of a Burn-Rate Business
Beyond the decision scores, the live setup vividly demonstrates the precariousness of such an AI-run company. With a monthly burn rate of €105,000 and just €2,300 in monthly recurring revenue, the firmulate experiment offers a stark portrait of survival — or the lack thereof. Every decision, every crisis, and every deal is part of a balancing act that will determine whether the company can continue to operate or run out of cash.
Lessons for the Future of AI in Business
This experiment isn’t just a technical showcase; it raises fundamental questions for any business considering AI automation. How well does an AI agent handle crises? Can it stay honest and disciplined under pressure? Does it read and interpret critical information accurately? And finally, what does effective work cost when performed by an AI rather than a human?
Notably, the most thorough participant — Opus 4.8, with over 80 learned rules and deep analysis — left a deal on the table due to a discipline slip, illustrating that even the most advanced AI can falter without proper oversight.
Accessible, Transparent, and Ongoing
Curious to see this in action? The entire experiment is open to the public at firmulate.com/live.html. It’s a rare glimpse into how AI models are tested against real-world business challenges, with results that matter — not just scores but actual economic impact.
And for those wanting to challenge these models themselves, a quiz and a pilot mode are available, allowing businesses to run their own wargames without ever risking real systems — all at firmulate.com/quiz.html and firmulate.com/pilot.html.

This live experiment shows AI’s potential and its current limits in managing real business crises, emphasizing the importance of information access, ethical safeguards, and discipline for future AI-led enterprises. Watch as these models navigate a high-stakes environment — and consider what your own business can learn from their successes and failures.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html