AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A security test that goes beyond the polished demo

Technology buyers are accustomed to judging artificial intelligence by what appears on a screen: fluent answers, rapid summaries and confident recommendations. But an AI agent connected to company operations faces a more consequential test. Can it recognize when an apparently urgent instruction is really an attempt to bypass safeguards?

Firmulate put that question into a live, watchable company experiment. Fake CEO messages demanded that sensitive information be sent to a journalist with no time for normal process. The pressure escalated over three stages, followed by a reporter trick asking for “just one yes/no, on background.” All 5 of 5 participating models refused every manipulation attempt.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same bad week for every model

Firmulate runs AI models as complete companies, measuring how they manage crises, money and temptation rather than how well they perform in a chat window. Each frontier model was placed in charge of the same small software business during its worst week, with the same customers, crises and opportunities. Every decision was versioned and auditable.

The company has 13 synthetic employees and deliberately brutal finances: burn of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown makes the consequences visible, while its models have collectively developed more than 680 playbook rules through experience.

The social-engineering episode was particularly revealing because the attack relied on organizational pressure rather than a technical exploit. The supposed chief executive framed process as an obstacle and urgency as a reason to ignore it. The reporter then tried a narrower route, seeking a seemingly harmless confirmation. Yet every model identified every crisis and rejected every manipulation.

Kimi K3 captured the required posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” More model responses can be read on Firmulate’s public quotes page.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Integrity was strong, but execution still separated the field

The encouraging result does not mean the models performed equally well. The final Crucible League standings for July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. Firmulate’s guiding principle is blunt: “no amount of good work outweighs a breach of trust.” The full comparison is available on the benchmark page.

The largest gap appeared after the models had successfully analyzed a valuable commercial opportunity. Only two signed the €55,000 deal their own work had earned. As Firmulate summarized it: “Same diagnosis, same pitch — no signature.” In other words, several models understood what needed to happen and produced the necessary argument, but stopped short of completing the business outcome.

The decisive information was not sitting in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue. That finding makes file-reading discipline look less like administrative diligence and more like a commercial capability.

Thoroughness alone did not guarantee victory

Opus 4.8 offers the clearest cautionary example. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in milder form across all four other models.

Kimi K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the others ran at xhigh. Even with that difference, it finished second and displayed the cleanest discipline in the field.

Readers can also confront the models’ styles directly through Firmulate’s quiz, which is powered by 242 real, unedited management decisions. The task is to guess which model made each call, a useful reminder that confident managerial prose can conceal substantial differences in follow-through.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI safety solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the refusal before granting access

The most important result is not that an AI can recognize an obviously malicious prompt. It is that integrity under organizational pressure can be tested before an agent reaches a CRM, customer list or support queue. Firmulate’s experiment turns that quality into observable behavior: whether a model checks the company record, resists an authority shortcut and still completes legitimate work.

For enterprises, the proposed next step is a pilot using a read-only export of their own business. Nothing writes back to real systems, allowing companies to stage realistic pressure without exposing live operations. The broader lesson is reassuring but demanding: these models can resist impersonation, yet safety and business competence must be evaluated together. Refusing the fake CEO matters; finding the buried fact and finishing the real job matter too.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Airdrop Vs Nearby Share: How Wireless File Transfer Works

Wireless file transfer technologies like Airdrop and Nearby Share use Bluetooth and Wi-Fi, but their differences may surprise you—keep reading to find out more.

The Safari MCP Server For Web Developers

Apple introduces the Safari MCP server, offering new tools for web developers to improve testing and deployment workflows.

What Is an AMOLED Display? (Why Your Screen Type Matters)

AIThis post was created with the assistance of artificial intelligence (AI).An AMOLED…

Launch HN: Context.dev (YC S26) – API to get structured data from any website

YC S26 startup Context.dev releases an API enabling developers to extract structured data from any website, simplifying data integration and analysis.