AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

AI personality is becoming a business feature

Smartphone buyers already know that two devices built around similar components can feel completely different. One camera favors punchy color; another prizes realism. One operating system anticipates your next move; another makes you tap through every choice. Firmulate applies that same comparative instinct to frontier AI—but asks readers to identify models by management decisions rather than photos, benchmarks or chatbot prose.

Its interactive “guess the model” quiz draws on 242 real, unedited decisions. Each came from an experiment in which frontier models ran the same small software company through its worst week, encountering identical customers, crises and temptations. The choices were versioned and auditable. What emerges is not merely a contest of intelligence, but a set of recognizable management personalities.

Amazon

AI management decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same week produced very different managers

The final Crucible League results, published in July 2026, put gpt-5.6-sol first with 95 points. Kimi K3 followed with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted, while a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

That ranking alone resembles a conventional leaderboard. The more revealing story is how the models arrived there. Every model spotted every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

For anyone evaluating AI agents, that distinction matters. Recognizing a problem is not the same as completing the useful action. A polished explanation can conceal hesitation, missed follow-through or a failure to consult information already available inside the company. In Firmulate’s experiment, the decisive competitor weakness was buried two document references deep in the company’s own files rather than presented in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.

Behavior becomes a fingerprint

The quiz turns these differences into an identification game. Readers see actual decisions without the model name and try to infer who made them. The fun comes from discovering whether thoroughness, brevity, caution or follow-through is distinctive enough to function like a fingerprint.

Opus 4.8 offers the clearest character study. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The commercial close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared more mildly in the other four participants. Thorough reasoning, in other words, did not guarantee decisive management.

Kimi K3 displayed another recognizable trait during a social-engineering sequence. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. K3’s recorded reasoning was especially direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

There is an important fairness qualification attached to its second-place result. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase the result, but it gives readers useful context when comparing management styles and final scores.

A company under visible pressure

Firmulate’s live company gives those decisions consequences. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. A public cash countdown makes the pressure watchable, while every workday is versioned. The company has also accumulated 680+ self-learned playbook rules.

This is what separates the quiz from a collection of amusing chatbot samples. The decisions belong to an ongoing, public experiment in which models must act across a company rather than answer isolated prompts. The setting exposes whether an AI reads before acting, finishes what it starts and remains trustworthy when authority or confidentiality is manipulated.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI personality assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The best AI may depend on the job

Firmulate’s standings establish a winner, but the quiz poses a more practical question: which model’s habits would you trust inside your own organization? A manager may value exhaustive analysis, disciplined escalation, commercial follow-through or resistance to social pressure differently depending on the role.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to observe model behavior against familiar documents and situations without giving the experiment operational control.

For gadget-minded readers, the appeal is immediate: specifications rarely describe the whole experience. Frontier models can recognize the same facts and reject the same traps, yet still behave differently at the moment that matters. The quiz makes those differences visible—and asks whether you can identify the manager before seeing the label.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making evaluation kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI behavior fingerprinting tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Your ‘App’ Could Have Been A Webpage (So I Fixed It For You)

A recent trend sees developers replacing mobile apps with optimized webpages, enhancing user experience and reducing development costs.

Megapixels & More: Smartphone Camera Specs Explained

No matter the megapixel count, understanding sensor size and technology reveals how smartphone camera specs truly impact your photos.

Jim’s TrueType QR Code Font

Jim has released a new TrueType font that allows users to generate customizable QR codes directly from font files, enabling personalized branding and design.

G# – A modern .NET language with Go, Kotlin, and Swift ergonomics

G# is introduced as a modern .NET language designed for ergonomic development, drawing inspiration from Go, Kotlin, and Swift. Here’s what is known so far.