Learn

Why not just ask a language model

It is a fair question and it deserves a specific answer rather than a defensive one. A model will give you a fluent, well organised, entirely reasonable paragraph about your price rise. The problem is what the paragraph is made of.

Where a model is used here

This product uses a language model in three places. They are named on this page and on the AI use page.

A spreadsheetone path, your own assumptions, no reactionAsking a modela plausible paragraph, no mechanism, no repeatA consultanta real answer, six weeks later, onceA twina range, a mechanism, and you can ask again tomorrow

A plausible paragraph and a mechanism are different objects, and only one of them can be checked.

Ask a model what happens if you raise price 20 percent and you get a genuinely good summary of what usually happens when companies raise price 20 percent. Ask it what happens to your company and it will write the same paragraph with your numbers in it, because that is the operation it is performing.

Five things a paragraph cannot do

It has no state

Your customers have contracts that end in particular months. Fourteen of them renew before the change reaches them and ten do not. A model has no ledger of those dates carrying forward month by month. It can mention that contracts matter, which is not the same as knowing that 54 percent of your revenue is locked until the fourth quarter.

It is not repeatable

Ask twice, get two answers. Ask in a different order, get a third. For a brainstorm that is a feature. For a decision that somebody will be held to in a board meeting it is a serious problem, because the analysis cannot be reproduced, only re-performed.

It has no band

A model will happily give you a number, and it can give you a range if you ask for one, but the range is drawn from how ranges are usually written rather than from propagating uncertainty through anything. Nothing was sampled. No replication ever happened. The interval is a rhetorical object.

You cannot ask which assumption mattered

The most useful question about any answer is which input it was standing on. In a simulation that is a measurement: move each assumption to the ends of its band, re-run, rank by how far the answer swings. In a paragraph it is another paragraph, generated by the same process that produced the first one.

It agrees with your framing

This is the one that costs real money. Ask whether the price rise is a good idea and you get a balanced answer. Ask why the price rise is the right move and you get support, fluently argued, with the counterarguments demoted to a closing caveat. The question carries the conclusion and the model is obliging. A simulation does not care how you phrased it, because your phrasing does not reach the arithmetic.

What a simulation gives you instead

The same answer twice, a range, and an audit trail

A run is deterministic given its seed. Ask the same question tomorrow and you get the same answer, so the conversation can be about the model rather than about which version of the model somebody saw.

The range comes from somewhere: each replication draws its parameters from the bands in your own ledger, and the tenth and ninetieth percentiles are taken across those replications. Narrow the band on an input by replacing a default with a real number and the band on the output narrows too. That is a mechanical relationship you can verify by doing it.

And every number is traceable. Not to a plausible sounding source, to a line in a file you uploaded, quoted in the ledger.

  • Same seed, same answer, every time
  • A band produced by sampling rather than by writing the word roughly
  • A ranked list of which assumptions moved the answer, measured rather than asserted
  • A named agent attached to each event, with the share of runs it happened in
month 0month 12best tenthworst tenth

Where a language model is better than any rule, and is used here

A contract is prose written by a lawyer. It says the term shall continue for a period of twenty four months from the commencement date, subject to the provisions of clause 9.2, and a regular expression looking for a number followed by the word months will find the wrong number about a third of the time. A model reads it correctly. That is not a small thing and it is not worth pretending otherwise.

So this product uses one in three places, all optional and all outside the simulation.

  • Reading chunks of prose the rules found nothing in. Anything it returns is written to the ledger as an estimate with a lower confidence, so you can always see which numbers came from a model rather than from your file
  • Rewriting a finished brief into better English. The numbers are already computed and are not touched. Turn it off and the engine brief stands on its own
  • Answering questions in Ask, from the graph, the ledger and the runs, with the rows that produced the answer attached

What it never does is supply a number to the simulation without passing through the ledger first, where it is labelled. With no key configured at all, the engine runs identically and makes no outbound calls.

Which to reach for

The taskAsk a modelRun a simulation
What might I be missing about this decisionYes, this is what it is forNo, it only knows what is in your files
Read this 40 page supply agreement and tell me the termYesNo
What does 20 percent do to my cash in month eightNoYes
How wide is the uncertainty on thatNoYes
Which of my assumptions is this answer standing onNoYes
Draft the memo explaining the decisionYesIt writes the brief, you write the memo
Will my biggest customer actually leaveNo, and neither will the simulationIt gives a frequency across runs, which is not the same as a prediction

The last row applies to both. Nothing in this list knows what your customer will do. One of them tells you how often it happens under stated assumptions you can inspect, which is the most honest thing available.

The real answer to why not just ask a model is that you should ask one, about the things it is good at, and then go and compute the part that needs computing.

Compute the part that needs computing

Build a twin from the record you already have and see the mechanism rather than the summary of one.

Build a twinHow this compares