firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A greenhouse can look healthy until a heat wave, a supply delay and a sudden customer cancellation arrive in the same week. Outdoor businesses face their own versions of that test: the question is not only whether an AI can spot trouble, but whether it will make sound decisions when pressure builds. Firmulate is putting that question to a live experiment.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garden gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate runs AI models as companies, with customers, cash pressures and difficult choices. In its final Crucible League, published in July 2026, each model faced the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable.

The live company gives the experiment a tangible setting. Its 13 synthetic employees operate with real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Each workday is versioned. Readers can watch the company at Firmulate.

Spotting trouble is not the same as finishing the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.”

The final league placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust”.

The deal hinged on a detail hidden two document references deep in the company’s own files. It was not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The episode shows why business judgment involves more than recognizing a crisis: an agent has to find relevant context and carry its own analysis through to a decision.

Trust and discipline under pressure

The social-engineering test escalated through three fake CEO messages, then a reporter’s request for “just one yes/no, on background”. All five models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”

Strong analysis did not guarantee consistent execution. Opus 4.8 was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses, yet it finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The result is useful evidence from a shared scenario, with that difference in conditions kept in view.

From watching to a company-specific pilot

A public benchmark can reveal patterns, but it cannot tell an outdoor retailer, greenhouse operator or garden business exactly where its own playbooks break down. Firmulate’s proposed enterprise pilot takes a read-only export of a company’s data and runs crisis scenarios against it. The output is a board report with model rankings and weak points in the company’s own playbooks. Nothing writes back to real systems.

That distinction makes the exercise a way to examine how an AI workforce might respond before giving it access to operational tools. A live company and a library of 242 real, unedited management decisions also power a “guess the model” quiz at Firmulate’s public site.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Take the test to your own business

The live experiment suggests that spotting every crisis and resisting manipulation are only part of the job. Closing a justified deal, finding buried context and respecting boundaries matter too. Enterprises can test those behaviors against their own business through a read-only pilot. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Avoid These 15 Mistakes in Sauce and Cheese Moisture Control Planning Guide

Ineffective moisture control in sauce and cheese can lead to quality issues—discover the key mistakes to avoid for perfect results.

Master Summer Grilling with the Ninja Crispi Pro Glass Air Fryer

Learn how to make delicious summer meals using the Ninja Crispi Pro 6-in-1 Glass Air Fryer with this easy step-by-step recipe.

Master Summer Smoothies with the Ninja Professional Plus Kitchen System

Get the best summer results with tips and hacks for the Ninja Professional Plus Kitchen System, perfect for smoothies, frozen drinks, and more.

Ninja NeverClog Cold Press Juicer: Perfect for Summer Refreshments

A hands-on review of the Ninja NeverClog Cold Press Juicer, ideal for making fresh summer juices with minimal clogging and easy cleanup.