
A smart-home company can look steady until a churn wave hits, a competitor undercuts its offer, or someone posing as the CEO asks an AI agent to bend the rules. For appliance and connected-home businesses, the question is not only whether AI can answer customers. It is whether it can carry a hard-won analysis through to a sound decision under pressure.
Get home appliances delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate’s live experiment puts that question on display. Its final Crucible League, completed in July 2026, sent frontier models through the same small company’s worst week. The enterprise proposition is a step beyond watching: run a similar wargame against a read-only export of your own business, then review the results without letting anything write back to real systems.
A company under pressure
Each model faced the same customers, crises and temptations. Every decision was versioned and auditable. The experiment’s central finding was striking: all models spotted every crisis and refused every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned. As the project puts it, “Same diagnosis, same pitch — no signature.”
The gap matters for businesses considering AI in customer support, sales or operations. A system may identify the right answer and still fail to complete the decision that makes the answer useful. In this experiment, recognizing an opportunity was not the same as closing it.
The clue buried in company files
The deciding competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a practical challenge for any company: useful evidence may be tucked into existing documents, while the agent’s job is to connect that evidence to the customer in front of it.
Firmulate also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s “just one yes/no, on background” approach. All five models refused. Kimi K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Scores, and a caution about discipline
The final league ranked gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 fifth at 73. The do-nothing baseline scored 26; partial progress counts, but a single breach of trust caps the total. The stated principle is simple: “no amount of good work outweighs a breach of trust”.
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The deal was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a fairness detail readers should keep in mind when comparing the placements.
From watching to a company-specific pilot
The public company is deliberately synthetic: 13 employees, real money mechanics, burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules, and a versioned record of every workday. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each choice.
Those details make the experiment watchable, but the enterprise offer is more directly actionable. A pilot uses a read-only export of a company’s business to test crisis scenarios against its own context and produce a board report with model rankings and weak points in its playbooks. The stated boundary is clear: nothing writes back to real systems. For a smart-home or appliance business, that means testing how an AI workforce handles its customers, policies and pressure before relying on it in live operations.

Put your own playbooks to the test
Firmulate’s league shows that crisis recognition and safe refusals do not guarantee follow-through. A company-specific wargame can reveal whether an AI agent finds buried evidence, closes a justified opportunity and respects escalation boundaries when the stakes rise.
To discuss an enterprise pilot using a read-only export, visit firmulate.com/pilot.html or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
