Monte Carlo and scenarios

The Plan tab answers one question — are we going to be all right? — and it answers it with a number: projected success, a percentage. This page is about where that percentage comes from, what it does and doesn't claim, and how to use the plan to compare one future against another.

For the buttons and fields, see Households and the plan. For what you're actually describing when you fill them in, see Money in, money out. For the formulas themselves, worked through with numbers, see How the plan is simulated.

Why one projection isn't enough

The obvious way to project a retirement is to pick an average return, apply it every year, and read off the ending balance. It's the arithmetic behind every "your money will last until age 91" calculator, and it has one fatal flaw: the order of your returns matters, and an average throws the order away.

Here are two retirees. Both start with $1,000,000, both take out $50,000 at the end of each year, and both live through exactly the same three years of market returns — +25%, 0%, −25% — in different order.

Return Balance after
Retiree A, year 1 +25% $1,200,000
year 2 0% $1,150,000
year 3 −25% $812,500
Retiree B, year 1 −25% $700,000
year 2 0% $650,000
year 3 +25% $762,500

Same three returns. Same average. $50,000 apart after three years, because B had to sell more shares to raise the same $50,000 while the market was down, and those shares weren't there for the recovery. From here the two never converge: whatever happens next happens to both of them equally, so the gap just compounds along with everything else.

Take the withdrawals away and the difference vanishes entirely: 1.25 × 1.00 × 0.75 is $937,500 whichever order you multiply in. Sequence risk is created by withdrawing, which is precisely what a retirement is, and precisely what a single average return cannot show you.

So the plan doesn't use one average. It uses a thousand orders.

What the simulation actually does

Every time something in the plan changes, the app runs 1,000 simulated lifetimes of your plan, month by month, from today to the year your primary household member reaches your "plan through age".

In each of those months:

  • The portfolio earns a random return, drawn from a bell curve whose center and width come from your Assumptions — the expected return and volatility of each asset class, blended by the allocation you actually hold today.
  • Everything else is not random. Your contributions go in, your spending, goals and health insurance come out, your Social Security and pensions come in, and any shortfall is drawn from the portfolio and grossed up for tax (see Taxes in the plan).

That split is the whole design. Only the market is uncertain; the plan is a fixed description of your life against which a thousand different markets are tried. If the answer changes, it's because you changed the plan.

Two conventions worth knowing:

  • Everything is in today's dollars. Rather than inflating fifty years of cash flows forward and making you deflate them back in your head, the app discounts the return by inflation once and simulates in real terms. A balance the chart shows for 2062 means what it would buy today.
  • A year's money is spread over its twelve months. You think in years — a tuition bill, a claiming age — and the portfolio compounds in months, so the year's net flow is divided by twelve.

Projected success

Projected success is the share of those 1,000 runs in which the portfolio was never emptied. 81% means 810 of them made it and 190 didn't.

One rule makes that number stricter than it sounds, and it's deliberate:

Running out is permanent. A run that fails to fund a single real dollar of need has failed, and stays failed — even if a later Social Security check, a pension, or a good decade would have refilled it.

The alternative — asking only whether the ending balance was positive — quietly forgives a household that ran dry at 72 and recovered at 84. Nobody actually lives through that; you sell the house, or you don't spend the money. So the plan counts it as a miss.

The percentage is graded by color on both the Plan screen and the dashboard's plan card, at the same two thresholds: green from 85%, bronze from 70%, and red below 70%. The thresholds are a presentation choice, not a verdict — a household with a large paid-off house or a willingness to spend less in a bad decade is reading a different situation from one without.

What the number is not: a probability that you personally will be fine. It's the share of one model's futures that survived, under your assumptions, your spending, and a market model that draws each month independently from a normal distribution. Treat it as a comparison instrument — the difference between two plans is far more trustworthy than the level of either one.

Reading the fan chart

The chart under the percentage plots calendar years against portfolio value, in today's dollars, and shows five of the 1,000 runs' worth of spread at each year:

  • the solid line is the median — the 50th percentile,
  • the inner band is the 25th to 75th percentile, the middle half of runs,
  • the outer band is the 10th to 90th,
  • a dashed vertical line marks the year the household retires (the year your last working adult stops), and
  • a dot on the baseline marks the first year of each goal you've switched on, colored by its tier.

Two things to keep straight while reading it:

The median is not a prediction. It's the middle of a thousand outcomes, and you will not live the median any more than a family has 1.9 children. The band is the point of the chart; the line is just its center.

The bands are not one future. The 10th percentile in 2050 and the 10th percentile in 2060 are almost certainly different runs. The lower edge traces "how bad did the bad cases look at each moment", not a single unlucky lifetime.

The bottom edge is often the most useful line on the chart, and it moves when the headline doesn't: a plan can hold at 96% success while its 10th percentile ending balance halves. If you're comparing two plans, look there as well as at the percentage.

Why the number moves when nothing changed

Change something, change it back, and the percentage may land a point away from where it started. That's expected:

The simulation draws fresh randomness on every run. A point of wobble is the sampling, not your plan.

This is a deliberate choice rather than an oversight. A fixed seed would make the screen perfectly repeatable and would also make it falsely precise — you'd read "84% versus 85%" as a real difference between two plans when it's noise from a thousand samples. Seeing the wobble teaches you the resolution of the instrument: differences of a point or two are not findings.

Two numbers you compare against hold still — but for different reasons, and the difference is worth knowing:

  • A scenario's score is genuinely frozen. It's stored alongside the snapshot, exactly as it was measured, and changes only if you save over that scenario. What gets frozen is the number that was on screen — an app run, random seed and all — so a scenario chip and the Base plan chip can sit a point apart on identical plans. That difference is the seed, and it's the reason a scenario stores its answer instead of re-deriving one that would wobble on every render.
  • The saved plan's score is re-scored, not stored. The one on the chips and on the dashboard's plan card is computed by the server on every request, always from the same fixed starting point — so two requests a minute apart agree. But it reads your live accounts, their current valuations, and today's date. It therefore moves when your data moves, with nobody touching the plan: a price backfill finishing overnight, another member importing trades, or simply the calendar rolling into a new month can all shift it.

So the big number is jittery because it re-rolls the randomness; the base plan's chip is steady minute to minute but follows your portfolio; and a scenario's score is the only figure that is truly a record of a past answer.

Scenarios

A scenario is a frozen copy of a whole plan — every field, exactly as it was — plus the score it earned when you froze it. It's how you keep an answer instead of a memory of one.

They live in the row of chips beside the success percentage:

● Base plan 81% ● Scenario B 74% ● Scenario C 88% + Snapshot Compare

  • Base plan is first, and it isn't a scenario — it's your live plan document, the one the dashboard reads and the one Save writes to. There is never a second copy of it to disagree with it.
  • Snapshot freezes whatever you're currently tweaking as a new scenario. The first one is called Scenario B (the base plan is A by implication), then C through G, then numbers. You can keep 12.
  • Deleting one is offered on the active scenario chip only, and only to an editor — select it, then delete, which puts a deliberate pause before a destructive click. The base plan has no delete: it's your plan, not a copy.
  • Clicking a chip loads that parameter set into the editor. It drops whatever unsaved tweaks you had — that's the one action besides Save and Discard that lets go of a draft.
  • Save plan writes to whichever chip is active. Its tooltip names the target, and the status line beside it says so in words: Saved to Scenario B at 3:42 PM.
  • Compare overlays every other chip's stored median on the fan chart as a dashed line in that chip's color — including the base plan's when a scenario is active, because "what does this what-if do to the plan we actually have?" is usually the comparison you're making.

Beside the live percentage, a small badge gives the difference against the active chip:

74% projected success −7 pts vs Base plan

The chips show stored scores and the big number is the live recompute, which is what makes that badge meaningful. If every chip tracked the live run they'd all say the same thing and there'd be nothing to compare against.

Changing one thing at a time

The plan is built for a probe loop rather than a single sitting: change one number, watch the percentage move, change it back. Two features exist to support that specifically.

The goals section prices each goal in points. Beside every goal you're funding is a badge like −6 pts: what the plan scores without it, subtracted from what it scores with it. A goal you've switched off reads would cost 6 pts instead — same arithmetic, different question. A goal's dollar cost is something you already know; what it costs the plan is only knowable by running the simulation again without it. Those badges come from a cheaper 200-run probe at a fixed starting point rather than the full 1,000 — a badge is a difference, and a difference wants a steady reference much more than it wants the headline's precision. They can therefore sit a point or two away from the big number, and they hold still while you type.

Almost everything can be switched off instead of deleted. Goals, pensions, income streams, an adult's health insurance and the whole Social Security program each carry a checkbox that parks them. Pensions, health insurance cards and the Social Security header label theirs In the plan; on a goal or an income row it's the bare checkbox at the start of the row, which a screen reader announces as Fund . Unticking one keeps every figure exactly where it was, so you can ask what if this weren't there? and get the numbers back with one click. Deleting is for figures that were never right.

The discipline that makes the loop work: change one thing per question. Two edits give you one number and no way to attribute it, and the interesting result is almost always that one of the two did nearly all the work.

Letting an assistant do the probing

If you've connected an AI assistant to your household, it can run that loop for you. There's a read-only simulate operation it can call: it hands over a whole what-if plan document and gets back the score, measured against your real accounts and allocation, without writing anything — not your plan, not a scenario, nothing. So an assistant can sweep a dozen variations (each adult's retirement year, a 10% spending miss, a Social Security haircut, every goal switched off in turn) and report which single change is worth the most points, in your terms — "one more year of work", "$400 a month less" — rather than handing you a list of numbers.

A ready-made version of that sweep ships with the connector as a plan-review prompt; in Claude Code it appears as /inv:plan-review. A finding worth keeping becomes a scenario, saved beside your plan rather than over it. Setting the connection up is covered in Connect your AI assistant.