Accuracy

What it can and cannot tell you

This page is written before you are asked for money, and it is linked from the footer of every other page. A model that looks certain is the cheapest way to make an expensive mistake, so here is where this one is strong, where it is weak, and the five ways it can mislead you while appearing to work.

How the ledger worksWhat is built and what is not

Nothing here is softened. If a claim on another page contradicts this one, this one is right.

640 runs of one scenario

The accuracy of a model is the accuracy of its assumptions. This one shows you every assumption it has.

That is not a disclaimer, it is the whole design. The engine is arithmetic, so it is exactly as good as the numbers it is fed and the equations it was given. Every number appears in one ledger with its origin, and the origins are not blended: read from a document with the line quoted, worked out from something that was, estimated by a language model, set by you, or an industry default. When an answer is standing on defaults, the ledger says so at the top.

The strong side

What a simulation of this kind is genuinely good at

All five of these survive fairly large errors in the inputs, because they come from structure rather than from a forecast.

Shape

Whether the line goes up then down, dips before it recovers, or never recovers. A price rise that helps for four months and then stops helping has a shape, and the shape is visible long before the exact numbers are trustworthy.

Order of magnitude

Whether the decision is worth tens of thousands or millions. That answer decides whether you spend another week on it, and it is usually right when revenue and margin came out of a document rather than a default.

Timing

That margin arrives in month one, the customer answer arrives over a year, and the competitor lands between the two. Sequence is the part spreadsheets get most wrong and the part the agents get most right.

Which actor causes the damage

The events list names the agent, the month and how often it happened across the runs. Knowing that two thirds of the loss is one anchored account leaving, rather than broad churn, changes what you do next. This is structural, not predictive.

The width of the uncertainty

A wide band is information. It tells you the decision is not knowable from what you have, and the sensitivity view then tells you which single assumption would narrow it most.

Comparison between options

Six scenarios against one baseline on one set of seeds. Shared errors cancel in the difference, so a ranking is more reliable than any of the numbers being ranked.

The weak side

What it is bad at, and will stay bad at

These are not gaps waiting for a release. They are things this kind of model cannot do, and no amount of engineering inside it changes that.

Predicting a market

Nothing in the software knows what your industry will do next year. Demand growth is an assumption you supplied or a default sitting in the ledger. If the market turns, the model does not turn with it, and every number moves.

One off events

A lawsuit, a fire, a founder leaving, a tariff, an acquisition you did not see coming. The engine draws from distributions of ordinary business variation. It does not have a term for the thing that has never happened to you.

Anything the record does not contain

If your files do not mention a product line, the twin does not have one. If contract end dates are not in the customer list, the model spreads renewals evenly instead of knowing that four of them land in March.

Human decisions with reasons that are not in any file

A customer who stays because their operations director went to school with your plant manager. A competitor who cuts price because their owner wants to sell the business. The model has objectives and personalities, which are a reasonable stand in, and they are not the actual reasons.

Long horizons

Error compounds every month. Twelve months is the default and is defensible. Twenty four is worth looking at for shape. Sixty months is a drawing, not an answer, and you should treat a five year run as a way to see structure rather than a forecast.

Knowing your industry better than you do

The priors are starting points collected to keep a twin from being empty, not findings. Where you know a number, your number is better than the default, and replacing it is the highest value ten minutes you can spend in the software.

What comes back

Two hundred versions of next year, not one

month 0month 12best tenthworst tenth
A standard run is two hundred replications of the scenario and two hundred of the same twin doing nothing, on the same seeds. Each path drew its own parameters from the bands in your own ledger. The median is the line people quote. The tenth and ninetieth are the answer.

Failure mode one

Garbage in, confidently out

A model does not know that the customer list you exported is six months old, that the margin in the management accounts is before a rebate, or that the spreadsheet has a summary row the parser counted as a customer. It reads what is there and gets on with it.

The defences are all in the same place. Every fact carries the source, the chunk and the quoted line, so a wrong number can be traced back and corrected in one click. The readiness score on the overview is the share of the model standing on documents rather than defaults. The parse screen tells you when a file produced nothing readable instead of silently contributing zero.

None of that catches a plausible number that happens to be wrong. Only you can catch that, which is why the quote is next to every figure.

  • Check the ten largest numbers in the ledger against something you trust
  • Look at the customer count the graph found and compare it to the one in your head
  • Open any fact whose quote does not look like the sentence you expected

What to upload

Monthly revenuedocumentGross margindocumentPrice elasticitydefaultMonthly churndocumentCompetitor reactiondefaultContracted revenuedefaultLargest customer sharederivedCost per headderived

Failure mode two

A confident median hiding a very wide band

The median is one number and it is the one that ends up in the board pack. If the tenth percentile is a loss and the ninetieth is a large gain, that median is not an answer, it is the middle of an argument you have not settled yet.

The software will show you a median for a twin built entirely out of industry defaults. It is not lying. It is doing arithmetic on the numbers it has, and the band around it is honestly wide. The risk is a human reading the middle and ignoring the edges.

So the rule is simple. If the band straddles zero, you do not have an answer, you have a question about which assumption to go and find out. The sensitivity view names that assumption, measured by moving each one to its own low and high and re-running, rather than asserted.

640 runs of one scenario

Failure mode three

Elasticity is the single most load bearing default in the model

Price elasticity decides how much demand falls when you raise price. It sits under every pricing question, most competitor questions and a fair number of the growth ones. If your files do not state it, and almost no company does, it arrives as an industry prior with the widest band any origin gets.

A default assumption carries a band multiplier of one, against 0.45 for something read from a document and 0.25 for something you set by hand. That is deliberate: a default should make the answer visibly less certain. It still sits under the whole result.

If you change one thing in the ledger before making a decision, change this one. A quarter of real price history, or even a considered estimate from the person who does your quoting, moves the answer more than any other input in the software.

  • The curve on this page is computed from the same equations the engine uses, not drawn
  • Near the peak the curve is flat, so being on the right side of it matters more than the exact number
  • A sweep across a price range shows you the flat part directly, which is often the real finding

How sweeps work

+60%price down 30%price up 60%profit

Failure mode four

Adaptive management makes the company look smarter than it is

Runs are adaptive by default. That means the executive agents watch their own numbers and act when one crosses a line: a discount authorised when the pipeline falls behind, a hiring freeze when cash gets tight, a marketing push when new logos stall. Every action is gated on authority, on whether that personality moves that early, and on a cooldown so nobody pulls the same lever twice in a month.

This is more realistic than a company that sits still while you change one cell. It is also the most flattering assumption in the engine, because the simulated executives always notice, and they notice on time.

Real companies notice late. They argue about whether the number is real for two months, then act. If you want the unflattering version, turn adaptive management off and run the same scenario again. The gap between the two runs is roughly the value the model is assigning to your management team reacting well, and it is worth knowing how big that is before you rely on the result.

What the engine computes

CusCustomersComCompetitorsSupSuppliersSalSalespeopleExeExecutivesEmpEmployeesInvInvestorsRegRegulatorsParPartnersYou

Failure mode five: a beautiful simulation of the wrong question is still the wrong question.

This is the one that catches people who are good at models. The software will happily run two thousand simulations of a twenty percent price rise while the actual problem is that one account is eighteen percent of revenue and their contract ends in November. The output will be well made, the band will be sensible, the brief will read well, and it will be about the wrong thing. Before you run anything, write the decision down in one sentence. If that sentence is not the one the scenario is testing, close it and set up a different scenario.

The questions it is built for

Calibration

How much weight to put on each kind of answer

The question you are asking itHow much weight it deservesWhy
Is this decision directionally good or badA lotDirection survives large errors in the inputs. If every draw in the band has the same sign, the sign is probably right.
Roughly how big is the effectA fair amountOrder of magnitude holds up when revenue and margin came out of a document. It drifts when they came from a default.
When does it show upA fair amountTiming comes from ramps, contract lengths and reaction lags. Those are defaults unless your files set them, so check those three first.
Which actor causes the damageA lotThis comes from the graph and the objectives, not from a forecast. It is the most reliable thing the software produces.
Which of these six options is bestA lotCompared against one baseline on one set of seeds, so the errors the options share cancel out of the difference.
What is the exact figure in month nineVery littleA single point from a stochastic model is a coincidence. Read the band that month, not the line.
What will this one named customer doVery littleNamed accounts are modelled individually, but one account is one draw. The aggregate is far more reliable than any member of it.
What will my market do next yearNoneThe model does not know. Growth is an assumption in the ledger, and if it is wrong everything downstream is wrong with it.
Will this specific competitor respondSome, for shape onlyCompetitors match part of your move, late, and once. That is a reasonable default behaviour, not intelligence about that firm.

A good rule: the further down this table your question sits, the more the answer should change what you go and find out rather than what you go and do.

How to use it without being misled by it

Four habits. They take about fifteen minutes in total and they are the difference between a model that helps and a model that flatters.

  1. 1

    Check the ledger first, before you read the answer

    Sort by origin and read the defaults at the top. Those are the numbers nobody has checked. If a default is sitting under the mechanism the whole answer depends on, replace it or go and find it out. Reading the ledger after the answer is how you end up arguing for a number you already like.

  2. 2

    Look at the band, not the median

    Quote the tenth and the ninetieth out loud. If the tenth is a result you could not live with, the decision is about whether you can carry that tail, not about the middle. If the band straddles zero, say that plainly rather than reporting the median.

  3. 3

    Take the tripwires with you

    Every brief ends with the earliest observable signals: monthly churn at the quarter mark, new customers per month, where competitor pricing should be two months later, and the profit gap against doing nothing. Each has an expected value and the level that means the model was wrong. Put those four in your monthly pack.

  4. 4

    Re-run when a tripwire fires

    A tripwire firing is not a failure of the model, it is the model doing its job on time. Change the assumption that the real world just contradicted, re-run on the same seed, and see whether the decision actually changes. Often it does not, and knowing that is worth more than the original run.

The questions a sceptical reader asks next

Has this been backtested against companies where the outcome is known?

No. Nobody has run this against a set of real decisions with known outcomes and published the fit. That work has not been done, and until it has, the honest position is that the calibration of this model is untested.

When it is done it will go on the changelog with the numbers, including the cases where it was wrong. Anything else would make this page worthless.

If I replace every default with a real figure, is the answer then right?

Better, not right. Replacing defaults removes parameter error. It does not remove structure error, which is the model having the wrong equation for your business somewhere. If your demand is driven by a tender cycle rather than by price and marketing, a perfectly calibrated price elasticity is still modelling the wrong mechanism.

The way to find structure error is to look at whether the baseline run, with no scenario applied, resembles the last twelve months you actually had. If it does not, fix that before you trust anything built on top of it.

Why not give me a single confident number, like everything else does?

Because it would be false, and because the useful part is the width. A tool that returns one number is either hiding its uncertainty or does not know it has any. The band is the product.

Would a consultant do better?

On knowing your industry, often yes. A good one has seen twenty companies like yours and will tell you things no file contains. This software has read your files and nothing else.

On showing their work, on running the same question a hundred more ways at no extra cost, and on giving you the same answer again in six months, no. The two are not substitutes. The comparison page sets out where each stops.

What happens if I disagree with an assumption?

Change it. Every assumption is editable, an edited one is marked as set by you and carries the narrowest band of any origin, and you can lock it so a later rebuild does not overwrite it. Then re-run on the same seed and look at what moved.

Test it on something you already know the answer to

The fastest way to judge a model is to ask it a question you have already lived through and see whether it produces the shape you remember. Or start on the invented worked example, a contract manufacturer called Harborline Components with one customer far too large and a price it has not moved since 2023.

Build a twinWhat is built and what is not