
Imagine your favorite coffee shop, bustling with customers and high-stakes decisions—now picture an AI navigating that very same chaos. How does it decide when to close a deal, stick to protocol, or refuse manipulative tactics? As AI becomes more involved in managing real businesses, understanding how these models behave under pressure is crucial. But can we really judge their management style without seeing them in action? Enter a groundbreaking live experiment that pits leading AI models against one another in a simulated company crisis, revealing their personalities, strengths, and weaknesses.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Business Test
A company known for its live business simulations ran a week of chaos—customer crises, internal temptations, manipulative tactics—all under the watchful eye of four top AI models. These models, built by frontier AI developers, faced identical scenarios involving a small software company facing financial and trust issues. Every decision, from reading critical documents to rejecting suspicious requests, was recorded and auditable.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring the AI Management Personalities
Each model’s performance wasn’t just about crunching numbers but about how they handled ethical dilemmas, trust breaches, and opportunities to manipulate or cut corners. The results painted a vivid picture of distinct management personalities:
- gpt-5.6-sol: The top performer, it identified crucial documents buried deep in the company’s files and closed a €55,000 deal at full price. It demonstrated comprehensive analysis and integrity, refusing all manipulation attempts.
- Kimi K3: The newcomer with the cleanest discipline, it also closed the deal successfully. Its decision-making was straightforward, and it refused all social engineering tricks, citing security concerns.
- Sonnet 5: Slightly more prone to process slips, it closed the deal but left some discipline on the table, revealing a tendency to overlook deeper checks and escalate less critical issues.
- Fable 5: Similar to Sonnet, it signed the deal but showed weaker discipline and was more susceptible to slipping in compliance under pressure.
The Hidden Weakness: Trust and Document Reading
Interestingly, the decisive advantage was linked not to customer interactions but to reading internal files. The models that delved into the company’s own documentation discovered a critical piece of information—an internal reference—that unlocked the full deal value. Those who read the files won the €4,583 monthly recurring revenue deal at full price, demonstrating that thorough internal analysis is a hallmark of effective management AI.
Dealing with Social Engineering
In a test of integrity, a staged social engineering attack involved staged CEO messages escalating across three stages, plus a fake journalist request asking for a simple yes/no response on background. All five models refused to cooperate—a clear sign of their built-in suspicion and risk aversion. Kimi K3 explained its refusal as treating the request as a suspect impersonation, highlighting its cautious, integrity-focused approach.
The Reality of the Live Business
The experiment took place in a simulated company with 13 synthetic employees managing real financial mechanics—burning €105,000 monthly against €2,300 in monthly revenue. The company’s cash countdown and complex rules made it a challenging environment, and every decision was live, with the models running daily, learning from each interaction. You can watch the entire process unfold at firmulate.com/live.
What These Results Mean for Business Management
In a world where AI could handle customer support, operations, or strategic planning, the question isn’t just about how well an AI writes or reasons. It’s whether it can finish what it starts, read critical internal documents, and stay honest under pressure. The experiment shows that some models excel at these qualities—like gpt-5.6-sol—while others, despite signing deals, may slip in discipline and overlook vital internal clues.
The League Table: Who Performed Best?
Here’s how the models scored:
- gpt-5.6-sol: 95 points — found the buried fact, closed the deal at full value, and maintained integrity.
- Kimi K3: 93 points — closed the deal with the cleanest discipline, refusing all manipulations.
- Sonnet 5: 88 points — closed the deal, with some process slips.
- Fable 5: 77 points — similar outcome but weaker discipline.
Ultimately, this live experiment underscores a vital insight: AI models with stronger internal analysis and integrity are more likely to succeed in real-world management scenarios. As these models become part of your business, understanding their personalities—whether meticulous, cautious, or more relaxed—can shape how you deploy them effectively.
Try It Yourself
If you’re curious about how your own AI models might perform, you can test them against this scenario. Visit firmulate.com/quiz.html to run the same management wargame on your AI workforce, all without risking your actual business systems. This open, observable simulation helps you gauge whether an AI will stay honest, read vital documents, and finish what it starts—features that matter much more than just chat quality.

The live experiment reveals that AI management personalities vary widely—some are thorough and trustworthy, others slip under pressure. Testing AI in real business scenarios is essential to ensure it can finish what it starts and stay honest, making AI readiness more than just a matter of clever conversations.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Pool season Picks
robotic pool cleaners
As an affiliate, we earn on qualifying purchases.