AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine an AI that not only understands your customer issues but also refuses to be manipulated, stays disciplined under pressure, and closes critical deals — all while running a real, money-losing company. That’s exactly what happened in a groundbreaking live experiment, where AI models were put through their paces in the crucible of a small software company’s worst week. The surprising winner? Moonshot’s Kimi K3, a newcomer that outperformed established Western frontier models, proving it can deliver trustworthy, decisive management far beyond chatroom demos.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Experiment: Watching AI Manage a Business Crisis

In July 2026, four frontier AI models ran the same small software company through its toughest week — same customers, same crises, same temptations to cheat or manipulate. Every decision was recorded and auditable, reflecting real business mechanics and money flows, not just chat simulations. The goal: see which AI could identify crises, resist manipulation, and close a crucial €55,000 deal.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: A Clear Winner Emerges

All four models detected every crisis and refused every manipulation attempt. Yet only two managed to close the deal based on their own analysis. The winner? gpt-5.6-sol scored 95, just ahead of Kimi K3 with 93. The other two, Sonnet 5 and Fable 5, trailed behind at 88 and 77, respectively. Interestingly, the Opus 4.8 model, despite its thoroughness, scored only 73 and left opportunities on the table, showing discipline slips in closing deals.

Amazon

trustworthy AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Edge: Reading Deeper in Documents

The decisive advantage for Kimi K3 lay in its ability to uncover a critical, buried fact buried two documents deep in the company’s files — a detail key to securing the deal at full price, worth an additional €4,583 MRR. This subtle document reading demonstrated that understanding the full context is vital for trustworthy decision-making, especially when potential manipulations are at play.

Amazon

AI deal-closing automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting Social Engineering and Manipulation

The experiment also tested the models’ resilience against social engineering tactics. Fake CEO messages escalated over three stages, and a reporter tricked the models into a background yes/no confirmation. Remarkably, all models refused to proceed, with Kimi K3 explicitly reasoning that the request could be an impersonation or bypass attempt. Such integrity under pressure is crucial for deploying AI in real-world business operations where trust and honesty are paramount.

Amazon

enterprise AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Operational Reality: A Live Company Under AI Management

Beyond isolated tests, the models managed a real, operational company with 13 synthetic employees, managing real money mechanics. Daily, the company burns €105,000 against a €2,300 MRR. The AI models employ over 680 self-learned rules, with decisions every workday versioned and transparent. You can watch this live experiment at firmulate.com/live, observing how AI handles crises, keeps discipline, and navigates temptations in real time.

Why This Matters for At-Home Wellness Tech

While this experiment focuses on business management, the underlying insights are highly relevant for at-home wellness devices that rely increasingly on AI. Trustworthiness, decision transparency, and the ability to resist manipulation are just as critical when an AI guides your health routines or monitors your wellness data. Just as in business, the question isn’t just about how well it communicates but whether it can deliver consistent, honest, and effective results under pressure.

The Front Runners and the Field

  • gpt-5.6-sol: scored 95, found critical buried facts, closed the deal at full price.
  • Kimi K3 (Moonshot): scored 93, proved disciplined, read deeply, and won the deal.
  • Sonnet 5: scored 88, closed the deal but with minor slips.
  • Fable 5: scored 77, also closed the deal but with more process issues.

The Importance of Fairness and Testing

It is important to note that Kimi K3 ran without an effort parameter (the default API setting), while the others operated at xhigh. This difference underscores the importance of testing models under fair and consistent conditions before trusting them with critical decisions.

The Bottom Line: Picking an AI Model Is a Bet Without Proper Testing

This live experiment demonstrates that selecting an AI for your business or wellness application is no longer just about chat quality or headline scores. It’s about trust, discipline, and the ability to function under pressure. As the league is open, and the performance gap narrows, organizations should recognize that thorough testing — like the one conducted by Firmulate — is essential before integrating AI into their core operations. For more details and to see the full results, visit firmulate.com/benchmarks.html.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The real-world test shows a new AI player, Moonshot’s Kimi K3, outperformed established models in managing crisis, closing deals, and resisting manipulation. Trustworthiness, deep reading, and discipline are now key benchmarks for AI in business and wellness tech — not just chat scores. Rigorous testing is essential to pick the right AI partner for your organization’s future.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Beauty Tech for Body Care: Gadgets Beyond the Face (Cellulite, Etc.)

Transform your body with innovative beauty tech gadgets beyond the face—discover how these tools can redefine your skin and contour solutions.

DIY Microneedling: Is It Safe to Use Dermarollers or Pens at Home?

Aiming for skincare improvements at home with dermarollers or pens can be tempting, but understanding the risks is essential before trying DIY microneedling.

Meta Reuses Old RAM In New Servers With Custom Bridge Chip

Meta is repurposing existing RAM modules in its latest servers using a custom-designed bridge chip, aiming to reduce costs and improve efficiency.

Build vs Buy a Prebuilt AI Workstation

Deciding between building or buying a prebuilt AI workstation? Discover the real costs, performance, support, and upgrade paths to make the smartest choice.