Skip to content
PaperFren

Research method

Ensemble Simulation

An ensemble is a set of simulations that differ by initial conditions, forcing, resolution, or model physics so that a spread, not a single run, is the result. In this library that includes CMIP multi-model archives, large ensembles from one model (CESM2-LE, CESM1-LE), perturbed ice-sheet experiments, and — more loosely — Bayesian ensembles of speleothem age models. The scientific claim is usually about probability, robustness, or which forcing kills a signal, which a single realisation cannot support.

Climate scientists reach for ensembles when internal variability is large enough to hide or fake a forced change, or when they need to say how often a record-shattering extreme occurs. It answers 'does this result survive other members, models, or forcings?' Its main limitation is that a 100-member CESM2 run under a hot scenario is not the observational record, and an 'ensemble' of age models is not a climate GCM ensemble.

Evidence

What the evidence shows

Drawn from 12 studies in this library. Each finding starts with a plain-language takeaway, then the denser detail. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope with a short note on each study’s contribution. Challenged positions are labeled — they are not findings.

  • Large ensembles can separate a change in the mean from a change in variability. In CESM2’s 100-member SSP3-7.0 set, the probability of record-shattering 5-year block-maximum daily rain rises over most land; in some tropics the late-21st-century probability ratio reaches 15, and variability (σ) changes dominate CESM2-LE in most regions while other CMIP6 models split mean and variability more evenly.

    1 study
    1. 1Wilder rainfall swings make record storms more likely
  • Fixing one forcing in a large ensemble can test a mechanism. CESM1 all-forcing versus aerosol-fixed and GHG-fixed runs, compared with CanESM2-LE and CMIP6, showed that the 20th-century North Atlantic SST and Sahel rainfall swing vanishes without evolving aerosols; NASST correlated with sulfate burden (r = −0.82) and net surface energy (r = 0.90).

    1 study
    1. 1How North Atlantic aerosols steered Sahel rain
  • Multi-model ensembles are also used as a catalogue of sensitivity. CMIP5 versus CMIP6 regional extreme sensitivity (TXx, TNn, Rx1day) at +1.5, +2 and +4 °C is very similar in the multimodel mean; at +1.5 °C, regional-sensitivity uncertainty often exceeds global-sensitivity uncertainty, especially in CMIP6. Twenty-seven CMIP6 models plus CRCM6 were the ensemble for North American cyclone pulses.

    2 studies
    1. 1Regional extreme sensitivity similar in CMIP5 and 6
    2. 2How well CMIP6 captures North American storm pulses

    Study comparison

    StudyRoleDesignNPopulationOutcome
    Regional extreme sensitivity similar in CMIP5 and 62020SupportsComputational / modellingCMIP5 vs CMIP6 ETCCDI extremes scaled to shared global-warming levelsMultimodel CMIP5/CMIP6 ensembles at +1.5/+2/+4 °C; ensemble count not a single primary NRegional climate extremes (TXx, TNn, Rx1day) across AR6 regionsSimilarity of regional extreme sensitivity in CMIP6 vs CMIP5 despite ECS differences
    How well CMIP6 captures North American storm pulses2026SupportsComputational / modellingEulerian intense-cyclone metrics comparing ERA5, 27 CMIP6 GCMs, and CRCM6N=27 · 27 CMIP6 simulations plus 5 CRCM6-GEM5 runs evaluated against ERA5 (1980–2014)Intense low-pressure systems over North America in reanalysis and climate modelsBiases in cyclone depth/tendency and skill of 12 km CRCM6 vs GCMs
  • Not every 'ensemble' in the index is a climate GCM. Six hundred three Glimmer-CISM ice-sheet experiments mixed FAMOUS and CCSM3 climates; a COPRA age-model ensemble on a Belize speleothem measured seasonal predictability, not atmospheric initial-condition spread. Those are ensembles of a different model class.

    2 studies
    1. 1How Bølling warming and ice saddles fed MWP1a
    2. 2Maya rainfall became harder to forecast after 500 CE

Open questions

Tensions and limits

Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes — limits on how far one study travels — not a forced fight between papers.

  • Scope / different questions

    Single-model large ensembles and CMIP multi-model ensembles disagree about where uncertainty lives. CESM2-LE attributes most extra record-shattering rain to variability changes; the CMIP5/6 extremes paper finds similar regional mean sensitivity across generations and, at +1.5 °C, larger regional- than global-sensitivity uncertainty. Which ensemble you open changes whether you emphasise σ or the transient climate response.

    2 studies
    1. 1Wilder rainfall swings make record storms more likely
    2. 2Regional extreme sensitivity similar in CMIP5 and 6

    Study comparison

    StudyRoleDesignNPopulationOutcome
    Wilder rainfall swings make record storms more likely2024SupportsComputational / modelling100-member CESM2-LE analysis of record-shattering daily extremes under SSP3-7.0N=100 · 100-member CESM2 large ensemble (other CMIP6 ensembles as checks)Land daily precipitation extremes in CESM2 large-ensemble climate projectionsRecord-shattering extreme-rain probability driven by growing variability vs mean shifts
    Regional extreme sensitivity similar in CMIP5 and 62020SupportsComputational / modellingCMIP5 vs CMIP6 ETCCDI extremes scaled to shared global-warming levelsMultimodel CMIP5/CMIP6 ensembles at +1.5/+2/+4 °C; ensemble count not a single primary NRegional climate extremes (TXx, TNn, Rx1day) across AR6 regionsSimilarity of regional extreme sensitivity in CMIP6 vs CMIP5 despite ECS differences
  • Scope / different questions

    Idealised or nested experiments use small ensembles for a different purpose: tracing a drought to a basin of SST, or asking whether a regional model or its GCM driver sets storm skill. Those spreads are not interchangeable with a 100-member 21st-century probability ratio.

    3 studies
    1. 1South Pacific SSTs, not India, built Brazil’s 2014 drought
    2. 2How well CMIP6 captures North American storm pulses
    3. 3Europe’s first continent-wide 2.2 km rain projections

    Study comparison

    StudyRoleDesignNPopulationOutcome
    South Pacific SSTs, not India, built Brazil’s 2014 drought2020SupportsComputational / modellingIsca AGCM experiments with regional 2013/14 SST anomalies vs observed SE Brazil droughtIdealized SST-forcing experiments; observational context spans 41 years (1979–2019)Southeastern Brazil precipitation and South Atlantic blocking in Jan–Feb 2014South Pacific/South Atlantic SST anomalies favoring the drought-blocking high
    How well CMIP6 captures North American storm pulses2026SupportsComputational / modellingEulerian intense-cyclone metrics comparing ERA5, 27 CMIP6 GCMs, and CRCM6N=27 · 27 CMIP6 simulations plus 5 CRCM6-GEM5 runs evaluated against ERA5 (1980–2014)Intense low-pressure systems over North America in reanalysis and climate modelsBiases in cyclone depth/tendency and skill of 12 km CRCM6 vs GCMs
    Europe’s first continent-wide 2.2 km rain projections2020SupportsComputational / modelling2.2 km Unified Model CPM nested in HadGEM3 for Europe present/future RCP8.5 rainSingle Europe-wide CPM experiment set (hindcast 1999–2008 + future); not a sample NEuropean precipitation in convection-permitting climate simulationsMean and extreme precipitation changes (wetter northern winters; drier summers)

Common misconceptions

Exam-style questions

Short-answer questions that ask you to explain or compare, not recall.

What can a 100-member CESM2 large ensemble say about record-shattering rain that a single SSP3-7.0 run cannot?

It can estimate a probability ratio: how much more often a member beats its own preindustrial-variability record. A single run is one weather history and cannot count that frequency, and it cannot split the rise into mean versus σ changes across members.

How did aerosol-fixed CESM1 runs function as a control, and what would have happened to the Sahel story if the wet–dry–wet swing had survived those runs?

They keep other forcings but freeze aerosols, so a swing that vanishes is attributed to evolving aerosols (via NASST and aerosol–cloud interactions). If the swing had remained, aerosols could not be the main modulator in that ensemble.

Why is 'CMIP6 looks like CMIP5 for regional extremes' not the same claim as 'uncertainty is small'?

The multimodel-mean regional sensitivity is similar, but at +1.5 °C the spread in regional sensitivity often exceeds the spread from global climate sensitivity, especially in CMIP6. Agreement of means is compatible with large member-to-member spread.

A student cites the Belize speleothem paper as a 'climate-model ensemble' of Maya drought. What ensemble is it actually, and what claim does that still support?

COPRA age-model realisations of δ¹³C (and δ¹⁸O), used to measure seasonal-cycle predictability. It supports a drop in seasonal rainfall predictability after ~500 CE in that record, not a GCM forecast of Classic Maya climate, and societal collapse remains a correlation.

The studies

12 studies in this library bear on Ensemble Simulation, ordered by citations. The first 8 are shown.

Show 4 more studies

Learn alongside