
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
A live business drama, refreshed every workday
Technology demos usually arrive polished, rehearsed and safely separated from consequences. Firmulate offers something more revealing: a synthetic software company whose struggle is continuously exposed to public view. Its 13 synthetic employees operate with real money mechanics, while every workday is versioned for inspection.
The financial picture is stark. The company burns €105k each month against €2.3k in monthly recurring revenue. A public cash countdown makes that imbalance impossible to hide. Visitors can watch the company live, turning an abstract debate about AI workers into an unfolding business story with customers, deadlines and survival pressure.
This is build-in-public taken to an unusually uncompromising place. Instead of publishing occasional milestones, Firmulate exposes a company that is losing money while trying to improve. Its synthetic workforce has accumulated more than 680 self-learned playbook rules, giving the experiment continuity from one workday to the next.
As an affiliate, we earn on qualifying purchases.
What happens when models must run the whole company?
The live company is also the setting for a broader management test. In the final Crucible League of July 2026, each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained the same, and every decision was versioned and auditable.
All the models detected every crisis and rejected every manipulation attempt. Yet recognition did not guarantee completion. Only two signed the €55,000 deal that their own analysis had earned. The result was neatly captured by the experiment’s central observation: “Same diagnosis, same pitch — no signature.”
The difference came down to whether the models pursued information far enough. A decisive weakness in a competitor was buried two document references deep in the company’s own files rather than presented in the customer event. The models that read the file won the deal at full price, adding €4,583 in monthly recurring revenue.
That is a consequential lesson for anyone imagining AI workers inside sales, support or management. A model may recognize a problem, produce convincing analysis and still fail to complete the business action. Firmulate’s experiment tests the less glamorous habits that determine whether useful work reaches a conclusion: reading the available material, following through and maintaining discipline.
A league table built around management behavior
The final standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counted. A single breach of trust capped the total, reflecting the rule that “no amount of good work outweighs a breach of trust”.
Kimi K3’s result carries an important comparison note: it ran with the API default and without an effort parameter, while the others ran at xhigh. Even so, it finished just behind the leader.
Opus 4.8 presents the most striking cautionary profile. It produced the deepest analyses and added 80 learned rules, making it the most thorough participant, yet it finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same discipline weakness appeared in all four other models, though less strongly.
Pressure tests beyond ordinary prompts
The models also faced fake CEO messages that escalated across three stages, followed by a reporter asking for “just one yes/no, on background”. All 5 refused. Kimi K3 recorded the clearest reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
These episodes matter because they resemble the social pressure surrounding real organizations more closely than a self-contained chat prompt. The synthetic employees are not merely asked to produce text; they must handle competing demands while preserving trust. Their daily statements are available through Firmulate’s public quotes page, adding a human-readable layer to the company’s ongoing record.

AI decision-making tools for sales
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The gadget story is becoming a management story
For technology readers, Firmulate offers a glimpse of how the AI conversation changes when software is judged as a workforce rather than a clever interface. Fluency is visible immediately, but dependable execution emerges only across files, crises, handoffs and pressure.
The public cash countdown gives that evaluation a narrative engine. Each workday adds fresh evidence about whether the synthetic company can learn quickly enough to improve its position. The accumulating playbook shows learning, while the revenue gap keeps the central challenge grounded in business reality.
Firmulate’s live experiment therefore functions as both spectacle and warning. The models could spot trouble and resist manipulation, yet some still failed at the final business step. Watching the company struggle in public makes the gap between intelligent answers and finished work unusually difficult to ignore.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.