
Imagine a company run entirely by artificial intelligence, making real business decisions, facing genuine crises — yet losing €105,000 every month. This is not science fiction but a real, live experiment where AI models operate as a full-scale business with real money mechanics, daily challenges, and public performance metrics. For consumers of at-home wellness tech, this story reveals how AI’s capabilities and limits unfold in the wild — and why trust in AI-managed systems matters more than ever.
The Live Experiment: A Company Without Employees and a Cash Countdown
At the heart of this experiment is a small, digital company managed entirely by AI models, dubbed Firmulate. It features 13 synthetic employees, each guided by thousands of self-learned rules, working daily to run a business that’s actively losing money — burning through €105,000 each month against a monthly recurring revenue of just €2,300. This setup is a build-in-public showcase: every workday, the decisions are versioned, auditable, and publicly observable at firmulate.com/live.html.
The goal isn’t to demonstrate AI chat prowess but to evaluate management quality, crisis handling, honesty, and discipline in real-time decision-making. This is where AI models are tested not in isolated chat demos but as functioning components of a business, with real financial stakes and operational pressures.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Do AI Models Fare in a Business Crisis?
Four frontier AI models — including GPT-5.6-sol, Kimi K3, Sonnet 5, and Fable 5 — were subjected to identical scenarios, simulating the company’s worst week. They faced the same customers, crises, and temptations to cut corners or manipulate. Every decision was documented and possible to audit later, ensuring transparency.
Remarkably, all four models identified every crisis and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. For instance, when fake approval requests were made through staged messages, all models correctly refused, with Kimi K3 explicitly reasoning that the request could be an impersonation or bypass of approval protocols.
However, only two of these models managed to close a deal worth €55,000 based on their analysis, despite diagnosing the same issues. The other two, including the most thorough participant, Opus 4.8, detected the opportunity but failed in execution — leaving the deal on the table. This critical gap highlights that even when AI understands the problem, it doesn’t always follow through to completion.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Files
Digging into the company’s own documentation revealed a decisive weakness. The models that closed the deal had read and understood a specific buried fact in the company’s files — a detail not apparent in customer interactions but essential for the sale. This underscores a vital point: AI’s effectiveness depends heavily on how well it can access and interpret internal knowledge, not just respond to surface-level cues.

HEALTHCARE A System in CRISIS!: AI is reshaping healthcare leadership in the U.S. and in Canada
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Discipline and Decision-Making in Practice
Opus 4.8, the most disciplined model with over 80 learned rules, showed a tendency to avoid risky decisions but slipped in discipline near the close, writing attempts into a locked department instead of escalating them as required. This slip cost the opportunity. Notably, this pattern persisted across the other models, emphasizing that even the most rule-informed AI isn’t immune to lapses under pressure.

Exploring Internal Communication
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and Wellness Tech
For consumers and companies invested in at-home wellness devices and digital health solutions, the story behind Firmulate offers a cautionary tale. If AI models are entrusted with managing customer data, support, or decision-making, their ability to stay honest, thorough, and disciplined in real-world crises is crucial. It’s not enough for AI to generate convincing chat messages — they must reliably finish what they start, access vital internal information, and resist manipulative tactics.
Why Trust and Transparency Matter
The live experiment’s transparency, with every decision versioned and auditable, provides a blueprint for how AI-managed systems should operate in sensitive contexts. It also highlights that current AI models excel at recognizing problems and resisting fraud but can falter in execution, especially when discipline slips or when hidden, critical data is overlooked.
Takeaway: Building Trust with AI in Business
As the AI ecosystem evolves, understanding its real-world performance — not just demo chats — is vital. The Firmulate experiment demonstrates that AI can identify crises and refuse manipulation, but execution and internal knowledge access determine success. For at-home wellness tech and beyond, the lesson is clear: trust in AI depends on transparency, disciplined decision-making, and thorough access to internal knowledge — not just surface-level responses.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html