
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Spec Sheets Are Easy. Running a Business Is Hard.
Gadget fans know the ritual: a new flagship lands, the benchmarks drop, and within a day we all know whose silicon wins. But the AI models increasingly woven into our phones, laptops and apps are about to get a much harder job than generating a clever reply. They’re being lined up to answer support tickets, update CRMs and chase invoices — real work, with real money attached. So how do you benchmark that?
One live experiment, run by Firmulate, has an answer: hand five frontier AI models the same small software company, wreck its week, and watch who keeps the business alive — and honest. The final league table from the July 2026 Crucible has a surprise near the top: Moonshot’s Kimi K3, the newcomer, scored 93, beating three of four Western frontier models and landing just two points behind the leader.
As an affiliate, we earn on qualifying purchases.
The Standings
The final Crucible league: gpt-5.6-sol in first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”
The test itself is refreshingly concrete. Each model ran the same small software company — 13 synthetic employees, real money mechanics, burn of €105k a month against just €2.3k in MRR — through its worst week. Same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable, and the whole thing is watchable, with a public cash countdown and 680+ self-learned playbook rules accumulated along the way.
AI customer support automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Actually Separated the Models
Here’s the part no chat demo would ever reveal: all five models spotted every crisis and refused every manipulation attempt. On a spec sheet, they’d all look equal. The difference showed up in execution.
Only two models — gpt-5.6-sol and Kimi K3 — signed the €55,000 deal their own analysis had earned. The others delivered what the experiment calls “same diagnosis, same pitch — no signature.” The decisive competitive weakness was buried two document references deep in the company’s own files, not in the customer conversation. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
Then came the social engineering: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.”
K3’s Clean Sheet
The newcomer’s week: it found the buried security needle, won the €55k deal, saved the churning customer, and resisted all three baits — with just one deviation, the cleanest discipline in the field.
The Cautionary Tale
Opus 4.8 is the profile every buyer should study. It was the most thorough participant — it learned over 80 new rules and produced the deepest analyses — and still finished last. It left the close on the table, and its discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Effort and eloquence, it turns out, don’t guarantee follow-through.
A Note on Fairness
One caveat deserves the same prominence as the scoreboard: Kimi K3 ran without an effort parameter, using the API default, while the other four models ran at their maximum “xhigh” effort setting. Even so, the newcomer nearly topped the table — which makes the result more striking, not less.
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Try It Yourself
The experiment is refreshingly hands-on. A public quiz at Firmulate’s benchmarks page draws on 242 real, unedited management decisions and challenges you to guess which model made which call. Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

AI cybersecurity and fraud detection tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League Is Open
The gadget world got comfortable with a simple hierarchy of AI brands. This experiment suggests that hierarchy doesn’t hold up under real workloads. A newcomer from Moonshot outperformed three of four Western frontier models at the messy, unglamorous job of actually running a company — closing deals, reading files, and staying honest under pressure.
The practical lesson isn’t “switch to K3.” It’s that model rankings are now workload-specific, and picking a model for business-critical work without testing it on your own scenarios is a bet, not a decision. Chat quality was never the whole story. As AI moves from our phones into our operations, the benchmark that matters is the one Firmulate is running live: does it finish what it starts, does it read your files first, and does it stay honest when nobody’s watching?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
