Research method
Ensemble Simulation
An ensemble is a set of simulations that differ by initial conditions, forcing, resolution, or model physics so that a spread, not a single run, is the result. In this library that includes CMIP multi-model archives, large ensembles from one model (CESM2-LE, CESM1-LE), perturbed ice-sheet experiments, and — more loosely — Bayesian ensembles of speleothem age models. The scientific claim is usually about probability, robustness, or which forcing kills a signal, which a single realisation cannot support.
Climate scientists reach for ensembles when internal variability is large enough to hide or fake a forced change, or when they need to say how often a record-shattering extreme occurs. It answers 'does this result survive other members, models, or forcings?' Its main limitation is that a 100-member CESM2 run under a hot scenario is not the observational record, and an 'ensemble' of age models is not a climate GCM ensemble.
Evidence
What the evidence shows
Drawn from 12 studies in this library. Each finding starts with a plain-language takeaway, then the denser detail. Supports means evidence for a finding; Challenges means evidence against a stated position; Qualifies marks scope with a short note on each study’s contribution. Challenged positions are labeled — they are not findings.
Large ensembles can separate a change in the mean from a change in variability. In CESM2’s 100-member SSP3-7.0 set, the probability of record-shattering 5-year block-maximum daily rain rises over most land; in some tropics the late-21st-century probability ratio reaches 15, and variability (σ) changes dominate CESM2-LE in most regions while other CMIP6 models split mean and variability more evenly.
Fixing one forcing in a large ensemble can test a mechanism. CESM1 all-forcing versus aerosol-fixed and GHG-fixed runs, compared with CanESM2-LE and CMIP6, showed that the 20th-century North Atlantic SST and Sahel rainfall swing vanishes without evolving aerosols; NASST correlated with sulfate burden (r = −0.82) and net surface energy (r = 0.90).
Multi-model ensembles are also used as a catalogue of sensitivity. CMIP5 versus CMIP6 regional extreme sensitivity (TXx, TNn, Rx1day) at +1.5, +2 and +4 °C is very similar in the multimodel mean; at +1.5 °C, regional-sensitivity uncertainty often exceeds global-sensitivity uncertainty, especially in CMIP6. Twenty-seven CMIP6 models plus CRCM6 were the ensemble for North American cyclone pulses.
- Regional extreme sensitivity similar in CMIP5 and 6
- How well CMIP6 captures North American storm pulses
Study Role Design N Population Outcome Regional extreme sensitivity similar in CMIP5 and 6 Supports Computational / modellingCMIP5 vs CMIP6 ETCCDI extremes scaled to shared global-warming levels Multimodel CMIP5/CMIP6 ensembles at +1.5/+2/+4 °C; ensemble count not a single primary N Regional climate extremes (TXx, TNn, Rx1day) across AR6 regions Similarity of regional extreme sensitivity in CMIP6 vs CMIP5 despite ECS differences How well CMIP6 captures North American storm pulses Supports Computational / modellingEulerian intense-cyclone metrics comparing ERA5, 27 CMIP6 GCMs, and CRCM6 N=27 · 27 CMIP6 simulations plus 5 CRCM6-GEM5 runs evaluated against ERA5 (1980–2014) Intense low-pressure systems over North America in reanalysis and climate models Biases in cyclone depth/tendency and skill of 12 km CRCM6 vs GCMs Not every 'ensemble' in the index is a climate GCM. Six hundred three Glimmer-CISM ice-sheet experiments mixed FAMOUS and CCSM3 climates; a COPRA age-model ensemble on a Belize speleothem measured seasonal predictability, not atmospheric initial-condition spread. Those are ensembles of a different model class.
Open questions
Tensions and limits
Some items are genuine disagreements on the same question. Others mark different assays, populations, or outcomes — limits on how far one study travels — not a forced fight between papers.
Single-model large ensembles and CMIP multi-model ensembles disagree about where uncertainty lives. CESM2-LE attributes most extra record-shattering rain to variability changes; the CMIP5/6 extremes paper finds similar regional mean sensitivity across generations and, at +1.5 °C, larger regional- than global-sensitivity uncertainty. Which ensemble you open changes whether you emphasise σ or the transient climate response.
- Wilder rainfall swings make record storms more likely
- Regional extreme sensitivity similar in CMIP5 and 6
Study Role Design N Population Outcome Wilder rainfall swings make record storms more likely Supports Computational / modelling100-member CESM2-LE analysis of record-shattering daily extremes under SSP3-7.0 N=100 · 100-member CESM2 large ensemble (other CMIP6 ensembles as checks) Land daily precipitation extremes in CESM2 large-ensemble climate projections Record-shattering extreme-rain probability driven by growing variability vs mean shifts Regional extreme sensitivity similar in CMIP5 and 6 Supports Computational / modellingCMIP5 vs CMIP6 ETCCDI extremes scaled to shared global-warming levels Multimodel CMIP5/CMIP6 ensembles at +1.5/+2/+4 °C; ensemble count not a single primary N Regional climate extremes (TXx, TNn, Rx1day) across AR6 regions Similarity of regional extreme sensitivity in CMIP6 vs CMIP5 despite ECS differences Idealised or nested experiments use small ensembles for a different purpose: tracing a drought to a basin of SST, or asking whether a regional model or its GCM driver sets storm skill. Those spreads are not interchangeable with a 100-member 21st-century probability ratio.
- South Pacific SSTs, not India, built Brazil’s 2014 drought
- How well CMIP6 captures North American storm pulses
- Europe’s first continent-wide 2.2 km rain projections
Study Role Design N Population Outcome South Pacific SSTs, not India, built Brazil’s 2014 drought Supports Computational / modellingIsca AGCM experiments with regional 2013/14 SST anomalies vs observed SE Brazil drought Idealized SST-forcing experiments; observational context spans 41 years (1979–2019) Southeastern Brazil precipitation and South Atlantic blocking in Jan–Feb 2014 South Pacific/South Atlantic SST anomalies favoring the drought-blocking high How well CMIP6 captures North American storm pulses Supports Computational / modellingEulerian intense-cyclone metrics comparing ERA5, 27 CMIP6 GCMs, and CRCM6 N=27 · 27 CMIP6 simulations plus 5 CRCM6-GEM5 runs evaluated against ERA5 (1980–2014) Intense low-pressure systems over North America in reanalysis and climate models Biases in cyclone depth/tendency and skill of 12 km CRCM6 vs GCMs Europe’s first continent-wide 2.2 km rain projections Supports Computational / modelling2.2 km Unified Model CPM nested in HadGEM3 for Europe present/future RCP8.5 rain Single Europe-wide CPM experiment set (hindcast 1999–2008 + future); not a sample N European precipitation in convection-permitting climate simulations Mean and extreme precipitation changes (wetter northern winters; drier summers)
Common misconceptions
If 100 CESM2 members agree, the result is an observation.
CESM2 has high climate sensitivity and SSP3-7.0 is a high-warming path, so probability ratios can be upper-end. Other CMIP6 models do not all attribute the same share to variability.
An ensemble mean is the most likely weather in a given year.
The mean averages internal variability away. The Sahel paper needs the time-evolving aerosol forcing to recover the 1925–1985 swing; an ensemble mean of SST-forced Isca members can mix many members and still miss mesoscale convection.
Every paper that says 'ensemble' is running many climate models.
The Maya paper ensembles age models of a speleothem; the MWP1a paper ensembles ice-sheet experiments. Those are uncertainty tools, not CMIP.
Exam-style questions
Short-answer questions that ask you to explain or compare, not recall.
What can a 100-member CESM2 large ensemble say about record-shattering rain that a single SSP3-7.0 run cannot?
It can estimate a probability ratio: how much more often a member beats its own preindustrial-variability record. A single run is one weather history and cannot count that frequency, and it cannot split the rise into mean versus σ changes across members.
How did aerosol-fixed CESM1 runs function as a control, and what would have happened to the Sahel story if the wet–dry–wet swing had survived those runs?
They keep other forcings but freeze aerosols, so a swing that vanishes is attributed to evolving aerosols (via NASST and aerosol–cloud interactions). If the swing had remained, aerosols could not be the main modulator in that ensemble.
Why is 'CMIP6 looks like CMIP5 for regional extremes' not the same claim as 'uncertainty is small'?
The multimodel-mean regional sensitivity is similar, but at +1.5 °C the spread in regional sensitivity often exceeds the spread from global climate sensitivity, especially in CMIP6. Agreement of means is compatible with large member-to-member spread.
A student cites the Belize speleothem paper as a 'climate-model ensemble' of Maya drought. What ensemble is it actually, and what claim does that still support?
COPRA age-model realisations of δ¹³C (and δ¹⁸O), used to measure seasonal-cycle predictability. It supports a drop in seasonal rainfall predictability after ~500 CE in that record, not a GCM forecast of Classic Maya climate, and societal collapse remains a correlation.
The studies
12 studies in this library bear on Ensemble Simulation, ordered by citations. The first 8 are shown.
- Regional extreme sensitivity similar in CMIP5 and 6
CMIP6 and CMIP5 agree on how hot days, cold nights, and heavy rain scale with global warming, even though global climate sensitivity differs.
- Europe’s first continent-wide 2.2 km rain projections
A 2.2 km convection-permitting model projects wetter northern winters and drier summers, with summer hourly extremes unlike coarser GCMs.
- Wilder rainfall swings make record storms more likely
Future record-shattering daily downpours become much more probable because extreme-rain variability grows, not only because the mean inches up.
- South Pacific SSTs, not India, built Brazil’s 2014 drought
Isca experiments show 2013/14 South Pacific and South Atlantic SST anomalies favored the blocking high that dried SE Brazil; Indian Ocean SSTs did not.
- Future Floyd–Florence rains could jump sharply
A design-rainfall method applied to three historic tropical cyclones projects large late-century increases in eastern North Carolina extreme rain.
- Maya rainfall became harder to forecast after 500 CE
A Belize speleothem shows seasonal rainfall predictability collapsing after 500 CE and bottoming in the Terminal Classic, alongside drought.
- How Bølling warming and ice saddles fed MWP1a
North American ice-sheet ensembles show saddle collapse and Bølling warming can supply several meters of Meltwater Pulse 1a in 340 years.
- How North Atlantic aerosols steered Sahel rain
Fixing aerosols at 1920 levels removes the observed multidecadal Sahel drought–recovery cycle that tracks North Atlantic SST.
Show 4 more studiesShow fewer studies
- US 1-in-1000-year rainstorms become much more common
Today’s thousand-year three-day downpours over the United States could strike several times more often at 2–4 °C of global warming.
- CMIP6 sea-level sensitivity lags observations
CMIP6/ISMIP6 transient sea-level sensitivity is ~30% below the historical “all but glaciers” rate, mainly because Antarctic ice-sheet models barely respond to warming.
- How well CMIP6 captures North American storm pulses
An Eulerian three-parameter storm metric shows CMIP6 GCMs share cyclone biases versus ERA5, while 12 km CRCM6 is usually closest.
- Finer climate models capture South Asian monsoon rain better
High-resolution HighResMIP runs reduce dry bias and better match monsoon timing and intensity in the Ganga–Brahmaputra–Meghna basin.
Learn alongside