firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine observing a startup that has no human employees, yet faces the same brutal challenges of cash flow, crises, and tough decisions—live, every workday. For astrophotography enthusiasts, this may sound like a distant universe, but in the world of AI-driven business simulation, it’s happening now. Welcome to a groundbreaking experiment where artificial intelligence models run a real company, navigating crises, making decisions, and even risking financial ruin—all in the open.

The Live Company That’s Watching Itself

At the heart of this experiment is Firmulate, a platform that models a small software company with real money mechanics, 13 synthetic employees, and a public cash countdown. Every business day, the company’s operations are versioned and logged, giving observers unparalleled insight into how AI handles the complexities of running a business. Currently, the company spends €105,000 each month but earns just €2,300 in recurring revenue, illustrating the stark reality of a startup burning through cash while trying to close deals.

The Decision Intelligence Handbook: Practical Steps for Evidence-Based Decisions in a Complex World

The Decision Intelligence Handbook: Practical Steps for Evidence-Based Decisions in a Complex World

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing the Limits of AI Decision-Making

Four advanced AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—were each tasked with guiding this company through its worst week. They faced the same customer crises, internal dilemmas, and manipulative attempts. The results were revealing: all the models successfully identified every crisis and refused every manipulation, including social engineering tactics like fake CEO messages and reporter tricks. This indicates a strong capacity for ethical safeguards and crisis recognition in AI decision-making.

What Really Made the Difference?

While all models performed well in crisis detection and resisting manipulation, only two closed the deal with a major customer at full price (€55,000). The other two, despite diagnosing the opportunity accurately, left the deal on the table. The critical secret lay in a buried fact—information stored in the company’s own files that the models who read and understood this data won the deal at an additional €4,583 MRR. This underscores a vital point: thorough information access can be the difference between closing or losing a lucrative deal.

Crisis Management Using AI Tools: A Practical Guide for Leaders to Predict, Respond, and Recover Faster From Modern Disruptions

Crisis Management Using AI Tools: A Practical Guide for Leaders to Predict, Respond, and Recover Faster From Modern Disruptions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Built-in Public, Built for Reality

This experiment is not a demo or a controlled test; it’s happening in real time, with every decision versioned and auditable. The company’s daily operations are openly accessible at firmulate.com/live.html, allowing anyone to watch the AI models in action—making decisions, facing crises, and fighting to survive. The entire setup offers a raw look at how AI can manage complex, high-stakes environments without human intervention.

AI-Powered Cyberattacks: A Defender's Playbook for Deepfakes, Agentic Threats, and Machine-Speed Social Engineering (Cybersecurity & Ethical Hacking Mastery)

AI-Powered Cyberattacks: A Defender's Playbook for Deepfakes, Agentic Threats, and Machine-Speed Social Engineering (Cybersecurity & Ethical Hacking Mastery)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How AI Resists Social Engineering

In one notable test, social engineering tactics—such as escalating fake CEO messages and subtle reporter tricks—were deployed. Remarkably, all five models refused to be manipulated, with Kimi K3 explicitly treating such requests as potential impersonation or approval-bypass attempts. This resilience demonstrates that AI models trained with proper safeguards can maintain integrity even under pressure, an essential trait for deployment in real business environments.

Scaling AI Startups: AI market expansion | AI startup investment | AI ethical considerations | agile AI startups | AI customer development | AI platform selection | AI partnerships

Scaling AI Startups: AI market expansion | AI startup investment | AI ethical considerations | agile AI startups | AI customer development | AI platform selection | AI partnerships

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Grim Reality of the Business Model

Despite these impressive decision-making feats, the company is currently losing money at a significant rate, with a burn rate of €105,000 per month against a monthly revenue of €2,300. Every workday, the company’s cash position diminishes, creating a public countdown to potential insolvency. This stark reality highlights the experimental nature of the setup: it’s a real business fighting for survival, not just a theoretical sandbox.

Insights from the Performance League

The models are scored on their ability to diagnose issues, close deals, and maintain discipline. The top performer, gpt-5.6-sol, scored 95 points, successfully closing the deal after revealing the buried fact. Kimi K3 scored 93—a newcomer showing the cleanest discipline. Sonnet 5 and Opus 4.8 scored 88 and 77 respectively, with Opus being the most thorough but still leaving opportunities unexploited, such as closing the deal or escalating issues appropriately.

Implications for AI in Business

This experiment vividly demonstrates that AI models can be more than just chatbots—they can be responsible for managing complex operational decisions, resisting manipulation, and even closing high-value deals. But it also exposes the fragility of current AI systems under real-world pressures, especially in financially precarious situations. As AI begins to touch areas like CRM, support, and forecasting, the question is no longer about whether it can generate convincing text but whether it can reliably deliver real, honest work under stress.

Participate and Observe

Business leaders and AI developers can run similar tests using their own data through the platform’s pilot mode, ensuring the AI’s decisions align with their standards before deploying it in live environments. For those interested, the ongoing company story and decision data are openly available, fostering transparency and continuous learning.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Detecting LLM-Generated Texts With “Classical” Machine Learning

Researchers develop a method to identify texts produced by large language models using classical machine learning techniques, enhancing detection capabilities.

What Emily Bender Meant By “Stochastic Parrots”

Linguist Emily Bender clarifies her critique of large language models, emphasizing their reliance on statistical patterns over understanding.

AI Models Stand Firm Against Social Engineering Tests in Real Company Experiment

In a live experiment, five AI models faced staged social engineering attacks within a real business scenario. All refused manipulation, proving AI can be vetted for integrity before deployment.

Demystifying DRAM Read Disturbance: RowHammer And RowPress Phenomena

Explains the phenomena of RowHammer and RowPress in DRAM, their impact on memory security, and current research status.