
Imagine a real company operating entirely with AI-driven decisions, taking critical business actions, and openly showing its struggles and triumphs — all in real time. For at-home wellness tech enthusiasts, this might sound distant, but it illustrates a future where artificial intelligence could be managing the essentials of a business, from crisis response to closing deals. The question isn’t just about AI’s writing skills, but whether it can actually run a company, stay honest, and complete real work under real pressure.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: A Company in the Crosshairs of AI
At Firmulate, a remarkable live experiment is unfolding — a fully operational company run by AI models, without human employees, facing every challenge a real business encounters. This isn’t a simulation; it’s a real company losing €105,000 each month against just €2,300 of monthly recurring revenue. Every workday, its decisions are recorded, versioned, and made open for scrutiny. This transparency offers a rare look into AI’s capacity to manage complex, money-critical operations.
The AI Models and Their Performance
Four frontier AI models were tasked with navigating a typical worst week for a small software company. The models faced the same customers, crises, and temptations, with decisions fully versioned and auditable. The results? All four identified every crisis and refused every manipulation attempt — from fake CEO messages to reporter tricks. Yet, only two managed to close the €55,000 deal their own analysis had earned. This marked a stark difference: the same diagnosis and pitch, but only some models signed the deal, demonstrating that honesty and diligence are crucial.
The Hidden Weakness: Reading the Files
The decisive edge went to the models that read the company’s internal documents. Hidden two references deep in the files was a critical piece of information that led to closing the deal at full price, adding more than €4,500 in monthly recurring revenue. Models that didn’t read these files missed this opportunity entirely, illustrating that in real business, the crucial details are often buried beyond surface-level data.
Resistance to Social Engineering
Social engineering tactics — fake CEO messages escalated in stages, or a reporter trying to get a quick ‘yes/no’ — were tested rigorously. All models refused to be manipulated, with Kimi K3 explicitly treating such requests as potential impersonation. This points to an emerging strength: AI systems can be trained or designed to recognize and resist deception attempts, a vital trait for trustworthy automation.

50 AI Automations That Pay: Done-for-You Workflows That Save Hours and Make Money
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Company Behind the Experiment
The company isn’t just a theoretical construct; it’s a tangible, functioning entity with a team of 13 synthetic employees. Its mechanics are transparent — every decision is made by AI models, and every day’s work is versioned. Yet, it’s struggling financially, burning €105,000 each month against a mere €2,300 monthly revenue. Its public cash countdown adds urgency to the experiment, which is accessible to the public at firmulate.com/live.
The Deep Dive: OPUS 4.8
The most thorough model participant, OPUS 4.8, integrated over 80 learned rules and performed in-depth analyses. Despite this, it left a deal on the table and experienced lapses in discipline, such as escalating issues instead of escalating them properly. These weaknesses were consistent across all models, indicating that even the most advanced AI still struggles with nuanced human decision-making and process discipline.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI
This experiment sheds light on critical questions for any enterprise considering AI automation: Can AI finish what it starts? Will it read and understand complex internal documents? Can it resist deception? Today’s models are capable of identifying crises and refusing manipulation — but they’re still vulnerable to process slips and missed opportunities. For at-home wellness device makers, this underscores the importance of decision integrity, transparency, and thorough data access when deploying AI in real-world settings.
What Does This Mean for You?
- AI can recognize and react to crises with impressive accuracy, refusing manipulation attempts.
- Reading deeper into internal documents can unlock hidden opportunities, making AI more effective at closing deals.
- Even top-performing models may slip on discipline, leaving value on the table—a reminder that AI isn’t infallible.
- Transparency in decision-making, versioning, and open experimentation are key to understanding AI’s true capabilities.

AI-Powered Cybersecurity: AI Tools for Enterprise Security | AI for Network Security | AI Risk Management | AI in Cyber Policies | Cyber Threat Management AI | ML in Fraud Prevention
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps and How to Engage
Businesses curious about testing their own AI workforce can run their strategies against a read-only export of their data, ensuring no interference with actual systems. This allows companies to gauge how AI handles crises, negotiations, and deception before full deployment. Details are available at firmulate.com/pilot.html.

AI companies are testing the limits of automation in a real, transparent environment. The experiment shows AI can identify crises, resist manipulation, and even close deals, but discipline slips remain. For businesses, transparency and thorough data reading are vital for trustworthy AI deployment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pool season Picks
robotic pool cleaners
As an affiliate, we earn on qualifying purchases.