
Imagine Watching an Entire Company Struggle for Its Life — Live and Unfiltered
What if you could see an AI-managed startup run through its worst week — facing crises, making decisions, and even risking its own survival — all in public view? This is no sci-fi story; it’s the reality of Firmulate’s live experiment, where a company with no employees, burning €105K monthly against just €2.3K in recurring revenue, fights to stay afloat day by day.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company in Action
At the heart of this project is a synthetic company, operated by an AI-driven workforce of 13 decision-making models. Every workday, the company’s decisions are logged, versioned, and made publicly accessible at firmulate.com/live. This transparent setup allows anyone to observe how AI models handle real crises, navigate ethical dilemmas, and attempt to secure business deals, all while operating with strict rules and self-imposed discipline.
The premise is simple: four leading AI models, representing different approaches, were tasked with running this miniature enterprise through its worst week. They faced the same customers, the same crises, and the same pressure to cut corners or manipulate the system. All decisions were auditable and traced back to the exact version of the model making them, ensuring full transparency of the process.
What the Models Did — and Didn’t Do
- All four models identified every crisis and refused to manipulate or bypass security measures.
- Only two of them successfully closed a €55,000 deal, which the company’s own analysis deemed earnable — yet, ironically, neither signed it in the end.
- Deep files hidden within the company’s data contained critical clues that weren’t visible in the initial customer-facing documents. Reading these internal references proved decisive, allowing the models that accessed them to secure additional €4,583 in monthly recurring revenue.
- Social engineering tactics, like fake CEO messages or behind-the-scenes reporter tricks, failed to sway any of the AI decision-makers. All models refused requests that could be seen as impersonation or approval bypasses, with Kimi K3 explicitly citing concerns over potential impersonation.

AI for Project Managers: A Desk Reference & Field Guide: Use Artificial Intelligence to Streamline Workflows, Automate Tasks, and Make Smarter Decisions with Practical Tools and Ethical Insights
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Does It All Mean for Real Business?
This experiment goes beyond a simple AI test; it raises fundamental questions about trust, discipline, and decision-making in automated systems. The live setup features a company burning €105K every month while earning only €2.3K in revenue, with a public cash countdown highlighting its precarious financial state. It’s a real-time showcase of how AI can, in principle, manage complex operations but also where it can falter under pressure.
The performance disparity among models underscores an important insight: the most thorough model, Opus 4.8, with over 80 learned rules and deep analysis, was the least successful in closing deals — often slipping into less disciplined behaviors like writing decisions into inaccessible departments instead of escalating them. Conversely, models like Kimi K3, which ran without aggressive effort parameters, demonstrated clearer discipline and more consistent refusal of unethical shortcuts.
Why Should You Care?
If AI agents start touching your customer relationship management, support queues, or forecasting processes, it’s not enough that they generate convincing chatter. The critical question is: will they follow through, stay honest under pressure, and actually deliver useful work? This experiment shows that only AI models that read deeply, refuse manipulation attempts, and maintain discipline will be truly reliable in high-stakes environments.

The Promises and Perils of AI in Education: Ethics and Equity Have Entered The Chat
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results That Matter
Here are the key outcomes from the live experiment:
- The top-performing model, gpt-5.6-sol, scored a 95 out of 100, found the crucial hidden fact, and closed the deal — representing complete and trustworthy performance.
- The newcomer, Kimi K3, scored just slightly behind at 93, demonstrating the cleanest discipline and successfully closing the deal without trying to shortcut the process.
- Sonnet 5 scored 88, closing the deal but with some process slips, while Fable 5 lagged at 77, often leaving deals unexecuted despite good rule discipline.
This transparent comparison shows that AI decision quality isn’t just about generating convincing chat — it’s about making the right decisions under pressure, reading internal documents, and resisting unethical manipulation.

ESSENTIAL AI TOOLS FOR TRANSPARENT MODELS USING SHAP, LIME, AND VISUALIZATION TECHNIQUES: 65 PRACTICAL EXERCISES TO ENHANCE INTERPRETABILITY AND TRUST IN BLACK-BOX MODELS
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Build-in-Public in Its Rawest Form
By openly running this experiment and making all decisions auditable, Firmulate offers a new way to evaluate AI’s readiness for real-world business. This is not a staged demo; it’s a continuous, evolving story of an AI-managed company fighting to survive. Anyone interested can watch the day-to-day progress, see decision logs, and even run their own simulations via the available tools at firmulate.com/pilot.html.

Key Takeaway
This experiment reveals that while AI can detect crises and refuse manipulative tactics, its ability to close deals and maintain discipline varies. Trustworthiness, deep reading, and disciplined refusal are crucial for AI to truly serve as reliable business partners — and watching this live experiment makes that clear in real time.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html