Comparing scenarios

Six options, one baseline, one set of seeds

Side by side comparison runs every scenario in the set against a single baseline, on the same seeds, with the same horizon and the same replications. The differences you read off the table are then decisions rather than sampling noise.

Build a twinThe scenarios

The baseline is the same twin with an empty lever list.

640 runs of one scenario

Two scenarios run separately are two different worlds, and the gap between them is partly weather.

Each replication draws its own elasticity, churn, margin and the rest from the bands in your ledger. Run scenario A today and scenario B tomorrow on different seeds and you have compared two different sets of draws as well as two different decisions. The result will look precise and it will be partly luck. On a decision with a wide band, luck can be larger than the difference you are trying to measure.

The seed

Same seed, same world, different decision

The seed determines every draw in a run: which elasticity this replication got, which month each named account's contract comes up, whether the competitor committed to a match, whether the launch landed.

Hold the seed and all of that is identical across the scenarios you are comparing. Replication seventeen is the same company having the same year in every one of them, except for the lever you pulled. The difference between two columns is then attributable to the decision, which is the only reason to put them next to each other in the first place.

It is the same argument as a controlled experiment, and it is available here for free because the world is simulated. Refusing to use it would be strange.

  • One seed across the whole set, not one per scenario
  • One horizon and one replication count, so nothing has more or less precision than anything else
  • One baseline run once, not a fresh baseline per scenario that would drift between them

What it does

How a comparison is assembled

  1. 1

    The baseline runs first, once

    The same twin with no levers at all, on the chosen seed, horizon and replication count. Everything else in the set is measured against this.

  2. 2

    Each scenario runs on the same seed

    In fast mode, because there are several of them and the per scenario dialogue log is not what you are here for. Each scenario keeps its own adaptive management setting, so a scenario deliberately run without management reacting stays that way.

  3. 3

    The ending month is lined up

    Every scenario reports the same metrics at the same month, so the columns are genuinely comparable rather than nearly comparable.

  4. 4

    The top events come with it

    A handful of the most significant agent events per scenario, so a column that wins on profit while losing its largest customer does not look like a quiet win.

The failure mode

What goes wrong when people compare two separate runs

It happens constantly with spreadsheets and it happens with simulations too. Somebody runs the price option in the morning, the hiring option in the afternoon, puts the two ending numbers in a slide and presents the difference as the answer.

Three things are wrong with that. The two runs drew different parameters, so part of the gap is the draws rather than the decisions. They may have used different horizons or different replication counts, so one is measured more precisely than the other. And neither is measured against the same do nothing line, so a scenario that simply happened to run in a better world looks better.

The fix is not more replications. More replications narrow the noise around each answer separately, which helps, but they do not remove the fact that the two answers are answers about different worlds. Sharing the seed removes it entirely.

The table

What you get back

ColumnWhat it isHow to read it
Change nothingThe shared baselineThe line every other column is measured against. Look at it first, because sometimes doing nothing is already going somewhere
Revenue at the end monthMedian across the replicationsCompare against the baseline, not against the starting month
Operating profitMedian at the end monthA column that wins here and loses on customers is buying the present from the future
CustomersMedian count at the end monthThe slowest moving number in the table and usually the most informative
CashMedian at the end monthThe lowest point on the path matters more than the ending value, which is why the brief reports it separately
Top eventsThe most significant agent moves in that scenarioA scenario where a regulator opened a review or a partner drifted is carrying a risk the ending number does not show

A comparison is deliberately shallower than a full run. When a column wins, run that scenario properly to get the full events list, the dialogue log and the brief.

Honestly

What a comparison cannot settle

Holding the seed removes one source of error. It does not remove the others, and a table with six tidy columns is very good at making you forget that.

  • Every column is standing on the same ledger, so a wrong elasticity is wrong in all six of them and the ranking can still be wrong
  • A twelve month horizon favours decisions that pay early. Some of these options are genuinely about year two
  • Ranking by one metric hides everything else. Read the customer column and the cash column before you accept the profit column
  • The comparison does not know your strategy, your covenants, your board, or which of these you have the appetite to execute

Which assumption is carrying the ranking

Next

What to do with the winner

Run it properly

A full run gives the events with their frequencies, the dialogue from the first replication, and the brief.

The brief

Sweep its main lever

You compared six discrete options. A sweep asks whether the setting inside the winner is anywhere near right.

Sweeps

Attack the ledger underneath it

If the ranking flips when one assumption moves inside its own band, the ranking is not a finding yet.

The ledger

Put the options on one page

Most decisions are not one question with a yes and a no. They are five things you could do with the same money, and the useful comparison is the one where all five lived through the same year.

Build a twinThe twelve ready made questions