
Imagine a smart home device that not only manages your lights and climate but also makes every decision about your business operations—without human intervention. It sounds futuristic, but one live experiment is testing just that: an AI-powered company with no employees, losing money every day, yet striving to stay honest and effective. This radical setup offers insights into how artificial intelligence could shape the future of automation, management, and trust in both homes and industries.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Live Company: A Digital Business Under Pressure
At the heart of this experiment is a real, functioning software company with an unusual twist: it is entirely run by 13 synthetic (AI-driven) employees. Every decision, from crisis management to strategic deals, is made by cutting-edge AI models tested against real-world scenarios. The company faces the harsh reality of burning €105,000 each month against a modest €2,300 monthly recurring revenue, with a public cash countdown reminding everyone of its fragile existence. Every workday, its actions are versioned, tracked, and displayed for the world to see at firmulate.com/live.html.
This setup aims to answer a fundamental question: can artificial intelligence serve as a reliable, honest management team? To probe this, four frontier AI models—each with different strengths—were tasked with navigating the company’s worst week, facing the same crises, customer temptations, and internal challenges. Their decisions are all transparent, with every choice recorded and auditable, creating a rare window into AI decision-making in a business context.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Experiment Revealed
The results are illuminating. All four models successfully identified every crisis, from customer complaints to internal risks, and refused every attempt at manipulation—such as social engineering tricks designed to bypass approval processes. For instance, when fake CEO messages escalated over three stages, all models refused to act on these requests, with Kimi K3 explaining, “Treat the request as a suspected approval-bypass / possible impersonation.”
However, a critical difference emerged when it was time to close a deal. Only two of the models signed a €55,000 contract that their own analysis had earned, despite all making the same diagnosis and pitch. The missing piece was in the company’s internal files—hidden references that, if read carefully, would have revealed a buried fact leading to the deal at full price (+€4,583 MRR). The models that read this buried information successfully closed the deal, illustrating that reading comprehension and digging into internal documents can be decisive in real business outcomes.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Deep Dive into Decision-Making and Flaws
Among the models tested, Opus 4.8 stood out for its thoroughness, analyzing over 80 learned rules and conducting the deepest analyses. Yet, it finished last—leaving the close on the table and slipping into discipline lapses, such as writing attempts into a locked department instead of escalating them. This underscores an important insight: even the most detailed AI models can falter if they do not prioritize or follow disciplined decision pathways.
Interestingly, the experiment also tested fairness by running some models without an effort parameter (a default setting that influences the model’s diligence), while others ran at very high effort. The results showed consistent performance differences, emphasizing that configuration choices matter when deploying AI in real-time management scenarios.
AI cybersecurity and fraud detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for AI in Everyday Life
While this experiment showcases a peculiar and extreme case—the entire company is a digital construct—its lessons extend far beyond. For consumers managing smart home devices, the key takeaway is that AI’s true value lies not just in generating convincing chat or automation, but in reliably completing tasks, reading relevant internal data, and resisting manipulation under pressure. Will your home AI, or a future AI assistant, be capable of making trustworthy decisions in complex scenarios?
Moreover, the experiment demonstrates that AI models can be tested in simulated yet realistic business environments before deployment. Through a dedicated platform, enterprises can run their own ‘wargames’ against copies of their operations, ensuring AI systems will perform ethically and effectively before they go live, with no impact on actual business data (more at firmulate.com/pilot.html).
As an affiliate, we earn on qualifying purchases.
The Big Takeaway
This radical build-in-public experiment offers a stark reminder: AI’s potential in business hinges on more than just language skills or superficial performance. It must demonstrate honesty, thoroughness, and resilience under pressure. For smart home enthusiasts, it underscores the importance of trustworthy automation—where decisions are based on comprehensive understanding rather than surface-level commands. As AI continues to evolve, real-world tests like this make clear that trust and reliability are the true currencies of a digital future.

Watching a live AI-driven company struggle for survival reveals much about trust, decision-making, and the true potential of automation—lessons that resonate in homes and industries alike.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.