AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Imagine an AI managing warranty claims during a product recall, handling a supplier dispute and answering a journalist—all in the same week. For a home appliance or smart home company, a polished answer in a demo is not enough. You need to know how the system behaves when money, customers and trust are on the line.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That is the question behind Firmulate, a live experiment that puts AI models in charge of the same small software company and lets people watch their decisions unfold. The experiment points toward a practical next step for businesses: test AI against your own company before giving it real responsibilities.

A shared crisis, different outcomes

In the final Crucible League, published in July 2026, five models faced the same customers, crises and temptations. The experiment ran each frontier model through the company’s worst week, with every decision versioned and auditable. The top results were gpt-5.6-sol at 95, Kimi K3 at 93 and Sonnet 5 at 88. Fable 5 scored 77, while Opus 4.8 scored 73.

The striking finding was not that models failed to notice trouble. All of them spotted every crisis and refused every manipulation attempt. The gap came when it was time to finish the work: only two signed the €55,000 deal their own analysis had earned. The project captured that disconnect in a line: “Same diagnosis, same pitch — no signature.”

For a company selling connected appliances, that distinction matters. An AI might identify a churn risk or recommend a response to a service issue and still fail to carry the decision through. A convincing conversation alone cannot show whether an agent will follow a company’s playbook when the pressure is real.

The detail hidden in the company’s own files

The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a useful reminder that business judgment can depend on context tucked away in existing documents, not just the latest message or customer interaction.

The test also included a staged social-engineering attempt: fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a more complicated result. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the close on the table and discipline slipped: it attempted to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

A live company to watch—and a pilot to run

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays make the experiment watchable at firmulate.com. The site also offers a quiz built from 242 real, unedited management decisions, where readers can guess which model made each choice.

Watching the public experiment can show how the models handle one company’s dilemmas. A pilot asks a more direct question: how would they handle yours? Firmulate says enterprises can run the same kind of wargame against a read-only export of their business. That can put company-specific crises and playbooks into the test while keeping the exercise from writing back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The experiment shows why AI readiness is about follow-through as much as crisis recognition. Models can spot risks and resist manipulation, yet still miss a deal or mishandle a boundary. For a business weighing AI agents, a pilot against its own data offers a way to examine those decisions before deployment. Explore a Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Roborock S8 MaxV vs Roborock Saros 10: Which Robot Is Better?

Compare the Roborock S8 MaxV and Saros 10 to find out which robot vacuum suits your needs best, focusing on suction power, navigation, and features.

Best Roborock Dust Bags & Accessories in 2026

Discover the top Roborock dust bags accessories of 2026. Find the best options for durability, compatibility, and value to keep your robot vacuum running smoothly.

Roborock S8 Pro Ultra Review: Top Features, Pros & Who It’s For

A detailed review of the Roborock S8 Pro Ultra, highlighting its strengths, weaknesses, and ideal users. Find out if it’s the right vacuum for you.

Countertop Dishwasher vs. Hand-Washing: Small Kitchen Cleanup Showdown

Join us as we compare countertop dishwashers and hand-washing to reveal which method truly saves water and energy in your small kitchen.