AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Spec Sheets Are Easy. Running a Business Is Hard.

Gadget fans know the ritual: a new flagship lands, the benchmarks drop, and within a day we all know whose silicon wins. But the AI models increasingly woven into our phones, laptops and apps are about to get a much harder job than generating a clever reply. They’re being lined up to answer support tickets, update CRMs and chase invoices — real work, with real money attached. So how do you benchmark that?

One live experiment, run by Firmulate, has an answer: hand five frontier AI models the same small software company, wreck its week, and watch who keeps the business alive — and honest. The final league table from the July 2026 Crucible has a surprise near the top: Moonshot’s Kimi K3, the newcomer, scored 93, beating three of four Western frontier models and landing just two points behind the leader.

Amazon

business AI assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Standings

The final Crucible league: gpt-5.6-sol in first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”

The test itself is refreshingly concrete. Each model ran the same small software company — 13 synthetic employees, real money mechanics, burn of €105k a month against just €2.3k in MRR — through its worst week. Same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable, and the whole thing is watchable, with a public cash countdown and 680+ self-learned playbook rules accumulated along the way.

Amazon

AI customer support automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Models

Here’s the part no chat demo would ever reveal: all five models spotted every crisis and refused every manipulation attempt. On a spec sheet, they’d all look equal. The difference showed up in execution.

Only two models — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. The others delivered what the experiment calls “same diagnosis, same pitch — no signature.” The decisive competitive weakness was buried two document references deep in the company’s own files, not in the customer conversation. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

Then came the social engineering: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s Clean Sheet

The newcomer’s week: it found the buried security needle, won the €55k deal, saved the churning customer, and resisted all three baits — with just one deviation, the cleanest discipline in the field.

The Cautionary Tale

Opus 4.8 is the profile every buyer should study. It was the most thorough participant — it learned over 80 new rules and produced the deepest analyses — and still finished last. It left the close on the table, and its discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort and eloquence, it turns out, don’t guarantee follow-through.

A Note on Fairness

One caveat deserves the same prominence as the scoreboard: Kimi K3 ran without an effort parameter, using the API default, while the other four models ran at their maximum “xhigh” effort setting. Even so, the newcomer nearly topped the table — which makes the result more striking, not less.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Try It Yourself

The experiment is refreshingly hands-on. A public quiz at Firmulate’s benchmarks page draws on 242 real, unedited management decisions and challenges you to guess which model made which call. Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI cybersecurity and fraud detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League Is Open

The gadget world got comfortable with a simple hierarchy of AI brands. This experiment suggests that hierarchy doesn’t hold up under real workloads. A newcomer from Moonshot outperformed three of four Western frontier models at the messy, unglamorous job of actually running a company — closing deals, reading files, and staying honest under pressure.

The practical lesson isn’t “switch to K3.” It’s that model rankings are now workload-specific, and picking a model for business-critical work without testing it on your own scenarios is a bet, not a decision. Chat quality was never the whole story. As AI moves from our phones into our operations, the benchmark that matters is the one Firmulate is running live: does it finish what it starts, does it read your files first, and does it stay honest when nobody’s watching?

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Safari MCP Server For Web Developers

Apple introduces the Safari MCP server, offering new tools for web developers to improve testing and deployment workflows.

The GNU Emacs Architecture: Unlocking The Core [Pdf]

A new detailed PDF explores the core architecture of GNU Emacs, revealing insights into its design and future development directions.

Is Grok 4.6 The Next Big Thing In AI? Find Out Here

xAI has announced Grok 4.6, the latest in its AI series, but details on capabilities, availability, and performance remain undisclosed as of now.

Solid Queue 1.6.0 Now Supports Fiber Workers

Solid Queue 1.6.0 now supports fiber workers, enhancing concurrency and performance for JavaScript applications.