firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

In the quest for AI that can manage real-world businesses, recent testing reveals some unexpected leaders.

Just as astrophotographers seek clarity amid cosmic chaos, business leaders increasingly turn to AI to navigate the turbulence of real-company decisions. The latest experiment from Firmulate offers a rare glimpse into how advanced AI models perform under severe pressure, with surprising results that challenge assumptions about what makes an effective AI manager.

Amazon

AI business decision simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing AI in the Real-World Business Arena

In July 2026, four leading frontier AI models were pitted against each other in a rigorous, publicly observable experiment. Each model was tasked with running a small software company through its most challenging week—facing the same customers, crises, and temptations. This setup offers a unique, real-time look at how these models handle decision-making, trust, and integrity under pressure.

The League Table: Who Came Out on Top?

  • gpt-5.6-sol scored highest with a 95, catching the buried security fact and closing a lucrative deal.
  • Kimi K3, the newcomer from Moonshot, scored just slightly below at 93, demonstrating exceptional discipline and honesty.
  • Sonnet 5 followed with an 88, while Fable 5 finished at 77, and Opus 4.8 at 73.

This ranking emphasizes the importance of integrity and thorough analysis over just surface-level performance. The baseline, a do-nothing approach, scored a mere 26, underscoring how much these models have advanced.

Amazon

AI model testing tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Honesty and Deep Reading Win

While all models successfully identified crises and refused manipulative tactics, only two models—gpt-5.6-sol and Kimi K3—actually signed the deal, which was validated by their own analysis. Interestingly, the decisive advantage for K3 was reading two documents deep into the company’s files, uncovering a buried security needle that others missed. This ability to dig deeper into company files made the difference between a failed pitch and a full-price deal worth over €4,500 in monthly recurring revenue (MRR).

The Power of Trust and Integrity

During the simulation, all models demonstrated an ability to refuse social engineering attempts, including staged messages from a fake CEO and a reporter trick. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is critical in real-world applications where manipulation can cost millions.

Amazon

AI decision-making analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real Company, Real Money, Real Risks

The live experiment isn’t just a theoretical test. It involves a functioning company with 13 synthetic employees, managing real money mechanics—burning €105,000 monthly against just €2,300 MRR. The company operates daily with over 680 self-learned rules, all versioned and observable at firmulate.com/live. This setup provides an unfiltered view of how AI manages complex, high-stakes business decisions.

Lessons from the Field

The most thorough participant, Opus 4.8, with over 80 learned rules, showed strong analytical depth but slipped in discipline during the final moments, leaving the close on the table. This highlights a key insight: even the deepest analyses can’t compensate for lapses in discipline or oversight.

Amazon

AI security and trust assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business Leaders

Choosing an AI model for management tasks isn’t just about raw performance scores. The ability to read relevant documents deeply, resist manipulation, and finish what it starts is crucial. And, as the results show, the best model is not always the one with the highest score—it’s the one with integrity and thoroughness.

Fairness and Testing Conditions

It’s important to note that Kimi K3 was tested without an effort parameter (the API’s default setting), whereas the other models operated at xhigh. This fairness consideration underscores that performance isn’t solely based on configuration but also on inherent capabilities.

The Open League: The Future of AI in Management

The current league table—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—illustrates how the field is evolving. The gap between the top models and baseline is stark, and the open competition signals that any new entrant can challenge established leaders with the right approach.

For organizations contemplating AI-driven management tools, these findings demonstrate that success hinges on honesty, diligence, and the ability to uncover hidden risks in company data. It’s no longer enough for AI to generate good-looking reports; it must finish the job with integrity and depth.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Key Takeaway

In the real-world management simulation, the top AI models demonstrated honesty, thoroughness, and discipline—traits that matter more than just scores. The winner, Kimi K3, uncovered hidden risks and secured a lucrative deal while maintaining integrity, highlighting the importance of deep analysis and trustworthiness in AI management systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Advice Made People 3X Less Accurate But 2X Confident, Researchers Found

Research shows people become less accurate but more confident when following AI advice, raising questions about reliance on AI guidance.

Crustc: Entirety Of `Rustc`, Translated To C

A project called ‘crustc’ has translated the entire rustc compiler into C, sparking discussions on compiler development and language interoperability.

Why Do We Need Human Mathematicians Anymore?

Exploring the ongoing relevance of human mathematicians amid advances in AI and automation, amid rising interest and ongoing debates.

When AI Benchmarks Plateau: A Systematic Study Of Benchmark Saturation

Research reveals AI benchmarks are reaching saturation, raising concerns about the future of AI progress and evaluation methods.