AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine an AI system so unreliable that even doing nothing at all earns it 26 points out of 100. For businesses considering AI for critical tasks, understanding what this score means is essential. Welcome to a new era of AI transparency, where honest benchmarks reveal not just what AI can do, but what it won’t do — especially under pressure.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get wellness gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Purpose of a Benchmark: More Than Just a Score

In evaluating AI models, it’s tempting to focus solely on high scores and shiny demos. But the core purpose of a benchmark is to measure real-world performance, including honesty, discipline, and commitment. The recent Firmulate experiment sets a new standard by testing frontier models in a simulated yet realistic business environment — a small software company facing its worst week.

Why Does the Do-Nothing Baseline Score 26?

The experiment revealed that even a completely inactive AI, one that refuses to make decisions or take actions, scores 26 out of 100. This might seem odd, but it reflects that some minimal engagement — like reading documents or acknowledging crises — is automatic and not necessarily indicative of competence. Partial progress counts towards the final score, emphasizing that even small steps matter in measuring AI reliability.

The Cap on Trust: One Breach Limits the Score

A key finding is that a single breach of trust — such as succumbing to a manipulative social engineering attempt — caps the maximum score an AI can achieve, regardless of other strengths. This principle underscores that honesty under pressure is non-negotiable. No matter how many crises are handled perfectly, a breach erodes the entire evaluation.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment in Action: Testing AI Under Pressure

The firmament of models tested included GPT-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8. All faced the same simulated week: identical customers, crises, and temptations. Every decision was explicitly versioned and auditable, mirroring real business decision-making.

Key Findings: Honesty and Diagnosis

  • All models detected every crisis and refused manipulation attempts.
  • Only two models signed the €55,000 deal based on their own analysis — the same diagnosis and pitch, but only the trusted models sealed the deal.
  • Deeper insights emerged from internal documents: models that read and understood these files secured the full deal, worth more than €4,583 monthly recurring revenue.

Social Engineering and Trust

Models faced staged social engineering, including fake CEO messages escalating through multiple stages and a reporter trick. All models refused to be manipulated, demonstrating robust security. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI decision-making audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Wargame: Real Money, Real Decisions

The experiment took place in a live, watchable environment called the Firmulate company emulator, where 13 synthetic employees managed real-money mechanics — burning €105k monthly against €2.3k MRR, with continuous updates and version control. The platform allows enterprises to run the same scenario against their own data without affecting actual systems.

Model Profiles and Performance

  • Opus 4.8, the most thorough participant with over 80 learned rules, failed to close the deal — demonstrating how discipline slipped when the focus was on analysis rather than execution.
  • K3 ran without effort parameters, achieving the highest score of 93, and closed the deal with integrity.
  • Other models scored between 77 and 88, with some slipping in process discipline but still closing deals.
Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business and AI Adoption

For companies considering AI for decision-making, support, or automation, the benchmark underscores critical questions: Will the AI finish what it starts? Will it read and understand your data before acting? Will it maintain honesty under pressure? These are the metrics that matter, far more than chat quality or superficial scores.

Why Trust Matters

A single breach of trust caps the maximum performance of an AI, regardless of other strengths. This emphasizes the importance of selecting models that demonstrate consistent honesty, especially when dealing with sensitive information or high-stakes decisions.

Amazon

business AI transparency platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Learn More and Test Your AI Readiness

Businesses can run their own wargames using the same framework, testing how their AI models perform under simulated crises. The process is transparent, versioned, and designed to reveal weaknesses before deploying AI into real-world operations. Details and demonstrations are available at firmulate.com/benchmarks.html.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The latest AI benchmarks highlight that honesty, discipline, and thorough understanding are vital for trustworthy AI. A zero-score do-nothing model still earns 26 points, illustrating that minimal engagement is automatic — but trust under pressure is what truly counts. Businesses should evaluate AI not just on what it writes, but on what it *does* under stress.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can a Red Light Face Mask Fit Into a Realistic Night Routine?

With just 10 to 20 minutes, a red light face mask can transform your night routine—discover how to enhance your relaxation and skin benefits.

Can AI Models Make Better Business Decisions Than Humans? A Live Experiment Shows the Answer

Live AI management wargames reveal which models truly excel under pressure—reading critical internal info, staying disciplined, and closing deals—beyond chat quality.

Himax Technologies Surges In Global Coverage

Himax Technologies experiences a sharp increase in international media mentions, with 20 mentions in a recent window, indicating rising global interest.

Google’s Hyper-personalized ‘Dreambeans’ Feed Is Now Free To Test

Google’s personalized ‘Dreambeans’ feed is now available for free testing, allowing users to explore highly tailored content experiences. Details remain limited.