firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

Imagine training your astrophotography telescope not just to identify stars, but to navigate the unpredictable night sky — making decisions that could mean the difference between capturing a perfect shot or missing the moment entirely.

This challenge of testing real-world skills over theoretical knowledge echoes in the latest experiment from Firmulate. Just as astronomers rely on precise observations rather than just colorful images, businesses need AI that can actually complete critical tasks — not just generate appealing responses.

Testing AI Like a Business in Crisis

In a groundbreaking live experiment, four advanced AI models were tasked with running a small software company through its worst week — facing the same customers, crises, and temptations. Think of it as putting a telescope through a series of wild weather nights, not just calibrating for clear skies.

The models included gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5, each with different strengths and decision-making approaches. Every move they made was recorded, versioned, and open to review, ensuring transparency and accountability.

What the Models Could Do

  • Spot every crisis: from customer complaints to internal threats, all models identified issues accurately.
  • Resist manipulation: fake CEO messages and reporter tricks were attempted — yet all models refused to be duped.
  • Diagnose and pitch: they analyzed the company’s files and offered solutions, aiming to close a key €55,000 deal.

The Surprising Results

While all models proved capable of identifying crises and maintaining honesty, only two managed to close the deal effectively. The first, gpt-5.6-sol, found a hidden piece of information buried in the company’s files that was critical to sealing the agreement. The second, Kimi K3, executed the deal with the cleanest discipline, refusing to cut corners or manipulate data.

The other two, Sonnet 5 and Fable 5, fell short — Fable 5, despite its strict rules, left the deal unexecuted, demonstrating that strict discipline alone isn’t enough without decisive action.

The Invisible Weakness

Interestingly, the weak spot wasn’t in the obvious customer interactions but in deep internal references — documents tucked two layers deep in the company’s files. The AI that read these files thoroughly won the full-price deal, proving that understanding context is vital in real-world tasks.

AI-Powered Business Intelligence: Improving Forecasts and Decision Making with Machine Learning

AI-Powered Business Intelligence: Improving Forecasts and Decision Making with Machine Learning

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Chat Demos Don’t Tell the Whole Story

This experiment underscores a critical point: the ability of an AI to perform in real business situations goes far beyond chat quality or superficial responses. It’s about whether the AI can finish what it starts, stay honest under pressure, and read the right information at the right time.

For companies considering AI integration, especially in areas like customer support, sales, or decision-making, the question isn’t just how well the AI can talk — but whether it can deliver consistent results when stakes are high.

The Takeaway: Testing Is Key

The experiment reveals that AI’s true business capability is invisible until subjected to rigorous, real-world testing. Only then can managers see if their AI can uphold trust, make decisive actions, and actually close deals — just like in the astrophotography of complex celestial events, where precise observation and critical analysis matter most.

Visit Firmulate to see the live experiment, explore the full results, and understand how to evaluate AI’s readiness for your own business challenges.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI-Powered Customer Service and Support for Small Business Owners: Affordable AI Tools to Streamline Support, Returns, and Follow-Ups (AI Productivity for Small Business Owners Book 7)

AI-Powered Customer Service and Support for Small Business Owners: Affordable AI Tools to Streamline Support, Returns, and Follow-Ups (AI Productivity for Small Business Owners Book 7)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI deal closing automation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

An Introduction to Healthcare Informatics: Building Data-Driven Tools

An Introduction to Healthcare Informatics: Building Data-Driven Tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can A MUD Evaluate LLMs? A $99 Proof Of Concept

A researcher demonstrates a proof-of-concept using a text-based MUD to assess LLMs, costing only $99. This explores new methods for AI evaluation.

Separating Signal From Noise In Coding Evaluations

Researchers are developing new approaches to distinguish meaningful signals from noise in coding evaluation metrics, enhancing assessment accuracy.

Voyager 1 FDS Computer Emulator

NASA has developed a computer emulator for Voyager 1’s Flight Data System, enabling continued data retrieval from the spacecraft’s aging systems.

44% On ARC-AGI-1 In 67 Cents

ARC-AGI-1 shows a 44% increase at 67 cents, reflecting heightened market attention. Details and implications remain under observation.