AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Your next AI upgrade needs a tougher test than a clever demo

A new phone can make an assistant feel remarkably capable. But the harder question for businesses is what happens when AI moves beyond the screen: faced with pressure, confusing instructions and a decision that could win a customer, can it follow through? Firmulate is putting that question to a live experiment. Its small synthetic software company has employees, business decisions and a public cash countdown. The aim is to see how AI behaves as a workforce, not just how polished its answers sound.

One company, the same worst week

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations. Their decisions were versioned and auditable, making the results a record of what each model did, rather than a collection of impressive-sounding responses.

The league placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment also applies a blunt rule to trust: one breach caps the total. As its principle puts it, “no amount of good work outweighs a breach of trust.”

The gap between knowing and doing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.” Recognizing a good move and actually making it turned out to be different tests.

The sale hinged on a detail tucked two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. For a business considering AI agents, that is a practical distinction: an agent may understand the situation in front of it and still miss context that changes the outcome.

Pressure tests beyond the customer call

The experiment also tried to lure models into bypassing approval. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a more complicated result. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. Diligent analysis, it seems, does not automatically make for disciplined execution.

There is a caveat in comparing the league: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The ranking is a record of this experiment, with that difference part of the conditions.

A live company, then a company like yours

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown. Its playbooks contain 680+ self-learned rules, and every workday is versioned. The company is synthetic; the financial pressure and decision trail make the experiment watchable at firmulate.com.

There is also a quiz built from 242 real, unedited management decisions. Readers can try to guess which model made each call at firmulate.com. It offers a more revealing challenge than judging AI by its best one-line answer: can you recognize the difference between a model that diagnoses, one that refuses a risky shortcut, and one that closes the deal?

For companies, the next step is to run the wargame against a read-only export of their own business. They can put crisis scenarios against their own playbooks and receive a board report with model rankings and weak points. Nothing writes back to real systems. That takes the experiment from watching a synthetic company to examining how AI might respond to the pressures and procedures of an actual one.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks under pressure

A model that spots every crisis can still miss the sale, and a thorough analysis can still end with poor discipline. Firmulate’s live experiment makes those gaps visible; a pilot lets enterprises examine them against their own business using a read-only export. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Smartest AI in the Test Still Came Last — Here’s Why That Should Worry You

Opus 4.8 had the deepest analysis and 80 learned rules — and still finished last. Diligence, it turns out, doesn’t equal impact.

Half-Life 2 Running Natively On HaikuOS

A developer has successfully ported Half-Life 2 to run natively on HaikuOS, marking a significant milestone for the operating system’s gaming capabilities.

Python 3.15’S Ultra-Low Overhead Interpreter Profiling Mode

Python 3.15 releases a new profiling mode with minimal performance impact, enhancing developers’ ability to optimize code.

The Safari MCP Server For Web Developers

Apple introduces the Safari MCP server, offering new tools for web developers to improve testing and deployment workflows.