AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Smartest AI You Can Buy Still Can’t Close a Deal

We gadget lovers love a benchmark. New chip drops, we check the scores. New AI model drops, we check the leaderboard — and the top models are all acing coding tests and chat contests. But here’s the uncomfortable question nobody’s asking: those tests measure how well an AI answers. What about how well it manages?

That gap is exactly what Firmulate, a live AI experiment you can watch right now, set out to measure. Instead of asking models to write code or chat cleverly, it handed four frontier AIs the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes.

The results from the final Crucible League table in July 2026 read like a management post-mortem, not a leaderboard.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Management Quality, Not Chat Quality

Firmulate calls it a wargame. Each model ran a synthetic software firm — 13 employees, real money mechanics, a burn rate of €105k a month against just €2.3k in monthly recurring revenue — through a gauntlet the site describes as churn waves, price increases, a downround scenario and a PR crisis. Every decision was versioned and auditable.

The final standings: gpt-5.6-sol took first with 95, Kimi K3 scored 93, Sonnet 5 landed 88, Fable 5 managed 77, and Opus 4.8 finished last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps your total. As the experiment puts it: no amount of good work outweighs a breach of trust.

Amazon

business AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding That Should Worry Every Buyer

Here’s the headline result: all five models spotted every crisis and refused every manipulation attempt. Social engineering? A fake CEO message escalated over three stages, plus a reporter trying to sneak an on-record quote with “just one yes/no, on background” — five out of five models refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models — gpt-5.6-sol and Kimi K3 — actually finished the job, signing the €55,000 deal their own analysis had earned. Firmulate’s summary of the failure: “Same diagnosis, same pitch — no signature.”

And the buried detail is the best part. The decisive competitive weakness that unlocked that deal wasn’t in any customer conversation — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The lesson translates directly to any business: your AI agent needs to read your files before it talks to your customers.

Amazon

AI crisis management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Trap

Opus 4.8’s profile is the most human story in the results. It was the most thorough participant in the entire field — 80+ learned rules, the deepest analyses of anyone — and it still finished dead last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the problem. The same weakness showed up, just weaker, in all four models. Effort and thoroughness, it turns out, are not the same as judgment.

One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still took second. That’s a genuinely impressive showing for the Moonshot newcomer.

Amazon

AI enterprise management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch It Lose Money

What makes Firmulate more than a one-off paper is that it’s alive. The company is real software running every business day, with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned — watchable at firmulate.com/live. It’s currently around company day 1131, and the site rebuilds itself twice a day.

Want to test your own instincts? A quiz powered by 242 real, unedited management decisions lets you guess which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The New Curriculum

Coding benchmarks told us AIs could build the software. Chat arenas told us they could talk about it. Neither tells you whether an agent, left alone with your CRM for a week, will spot the crisis, refuse the impersonator, read the buried document and close the deal it earned.

Scenarios with names like churn wave, price increase and PR crisis are the new curriculum, and the crucible results suggest we’re entering an era where the question isn’t “is it smart” but “is it a good manager”. The best conversationalist didn’t win here — the model that finished what it started did. Before you hand an AI the keys to your support queue or forecast, that’s the benchmark worth checking.

Full results and plain-language findings are on Firmulate’s benchmarks page.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

60HZ Vs 120HZ: Understanding Smartphone Screen Refresh Rates

Meta description: “Many users wonder whether 60Hz or 120Hz screens provide better performance, but understanding their differences can significantly impact your device experience.

How Does Face ID Work, and Is It Secure?

The technology behind Face ID uses advanced sensors and neural networks to recognize your face securely, but how safe is it really?

The AI Company Turning Its Cash Crisis Into a Public Technology Test

Firmulate turns a synthetic software company’s daily fight for survival into a public test of whether AI can actually finish the job under pressure.

G# – A modern .NET language with Go, Kotlin, and Swift ergonomics

G# is introduced as a modern .NET language designed for ergonomic development, drawing inspiration from Go, Kotlin, and Swift. Here’s what is known so far.