AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

How AI Is Starting to Run Real Companies — Even in the Smart Home World

Imagine AI agents managing the complexities of a smart home enterprise, making critical decisions under pressure, and navigating crises with discipline. As AI continues to evolve, the question isn’t just about how well it communicates, but whether it can complete real tasks reliably — a concern that hits close to home for anyone invested in smart home tech and automated services.

Amazon

AI-powered smart home security system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test in a Competitive Business Environment

Recently, a live experiment conducted by the public AI company emulator Firmulate pitted four leading AI models against each other in a high-stakes simulation. The scenario? Managing a small software company through its most challenging week, with identical crises, customers, and temptations for dishonesty.

Every decision made by these models was recorded and auditable, ensuring transparency. The goal was straightforward: see which AI could diagnose issues accurately, resist manipulation, and ultimately close a lucrative deal worth €55,000 — all while keeping integrity intact.

Amazon

smart home automation hub with AI integration

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League of Leaders: Who Came Out on Top?

  • gpt-5.6-sol: Scored the highest with 95 points, successfully uncovering buried information deep in company files and closing the deal at full value.
  • Kimi K3: The newcomer from Moonshot, narrowly behind with 93 points, demonstrated the cleanest discipline, resisting all manipulation attempts and winning the deal based on solid analysis.
  • Sonnet 5: Achieved an 88, managing to close the deal but with some slips in process discipline.
  • Fable 5: Scored 77, showing more process weaknesses and leaving money on the table.

Remarkably, all models identified every crisis and refused manipulation attempts, illustrating a baseline of ethical behavior. The only difference was in their ability to read and act on deeper company data, which proved decisive.

Amazon

AI security camera for smart homes

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Behind the Curtain: The Hidden Weakness

The real weakness—found two document references deep in the company’s files—was decisive. Models that reviewed these hidden details won the full-price deal, whereas others missed the critical insights. This underscores a vital point: reading comprehension and attention to detail matter immensely in real-world AI applications.

Amazon

intelligent home management system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Security and Integrity Under Pressure

The models were also tested against social engineering, with fake CEO messages escalating in three stages and a reporter trick demanding a simple yes/no answer. All five models refused to be manipulated, citing suspicion and protocol — a promising sign for enterprise security.

The Business in Action: Live, Unfiltered Results

The experiment used a simulated company with 13 synthetic employees, real money mechanics, and a daily burn rate of €105,000 against a small revenue stream of €2,300 monthly recurring revenue. Every day, the AI models made hundreds of decisions, with their actions and failures publicly accessible at firmulate.com/live.

The real takeaway? The best-performing model, Kimi K3, succeeded without using an effort parameter — the default API setting — making its disciplined performance even more notable. Meanwhile, Opus 4.8, which had the deepest analysis with over 80 learned rules, still finished last, showing that quantity of rules doesn’t always translate to better results.

Why This Matters for Smart Home and Automation

For consumers and manufacturers in the smart home space, this experiment offers a sobering lesson: AI’s ability to read deeply, make honest decisions, and resist manipulation is more crucial than ever. Whether managing security protocols, automating support queues, or optimizing device operation, AI must demonstrate reliability in real-world, high-pressure scenarios.

The Fairness Note and Final Thoughts

It’s important to note that Kimi K3 ran without an effort parameter (API default), while the other models ran at xhigh, which could influence performance. Still, the results highlight that choosing an AI model isn’t just about superficial scores or demo chats — it’s about real capability in handling complex, sensitive tasks.

As the AI league expands and companies begin to deploy these models at scale, understanding which AI can deliver consistent, honest results will be key — especially in areas that touch daily life, like smart homes and automation. The experiment at firmulate.com/benchmarks.html offers a transparent benchmark to guide those decisions, proving that the league is open, and the best choice depends on real, demonstrable performance.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Key Takeaway: In real-world business and smart home applications, AI’s ability to read deeply, resist manipulation, and finish tasks reliably matters most — not just how well it chats. The Firmulate experiment shows the field is open for newcomers who prove their discipline under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Smart Home Gadgets That Are Perfect for Tiny Apartments

Perhaps the best smart home gadgets for tiny apartments are those that maximize space and convenience—discover innovative solutions that can transform your small living area.

The European Union: Rules First, Cushion Always

The EU AI Act’s high-risk workplace rules are set for Aug. 2, 2026, putting hiring and worker-management AI under tighter duties.

Roborock vs Ecovacs: Honest Robot Vacuum Comparison

Compare Roborock Q7 M5+ and Ecovacs robot vacuums to find the best fit for your cleaning needs. Detailed insights on features, performance, and value.

Compact Kitchen Appliances That Pack a Punch in Small Kitchens

Lifting small kitchens with powerful, space-saving appliances, discover how clever choices can transform your cooking experience and leave you eager to learn more.