Methodology
Hamilton-Perry cohort change ratio model
UK Demographics uses the Hamilton-Perry (HP) method to project ethnic composition at local authority level. The HP method computes cohort change ratios (CCRs) from two Census observations (2011 and 2021), then applies these ratios forward to project future populations by age, sex, and ethnic group.
The model works in 20 ethnic groups, single year of age from 0 to 90 (with 90 as a closing group), and both sexes. Projections extend to 2061, with 2051 as the primary horizon and 2061 as illustrative.
Coverage, stated precisely. The dataset holds 318 distinct local authorities. 22 of them are Welsh, so "English local authorities", as earlier versions of this page put it, was wrong. The Hamilton-Perry model itself projects 269 of the 318: it needs an area to appear in both the 2011 and the 2021 Census under one code, which excludes Welsh authorities from the England-only 2011 extract and any authority created after 2011. The remaining 49 carry projections from an earlier model run that the current code cannot reproduce, and their uncertainty bands have been removed rather than left describing a model that did not produce them.
Data sources
- Census 2021: ONS custom dataset, single year of age by ethnic group by sex, 292 LAs (direct observations, no IPF dependency)
- Census 2011: 18-group age-sex-ethnicity dataset via NOMIS API
- ONS SNPP 2022: Subnational Population Projections used as total population envelope constraint
- DfE School Census 2024/25: Independent validation, 10 years of ethnic composition data by school and LA
- ONS Births Table 8: Age-specific fertility rates by ethnicity and IMD decile
Migration data sources (YE March 2026)
For migration-specific surfaces, UK Demographics tracks the Home Office Immigration Statistics quarterly release and the ONS Long-Term International Migration provisional estimates. Both were published 21 May 2026.
- ONS LTIM YE Dec 2025: Net migration 171,000 (95 percent CI 145,000 to 197,000). Peak 944,000 YE Mar 2023. Methodology shifted from International Passenger Survey (IPS) to RAPID + HOBID admin data from YE June 2021 onwards, so pre-2021 estimates are NOT directly comparable. See the net migration page for the full revisions tracker and the LTIM vs Home Office vs DWP NINo triangulation.
- Home Office YE Mar 2026: 778,625 non-visit visas granted (study 412,825, work 252,775, family 62,470, other 50,555). Skilled Worker visas 111,000 (-76 percent from YE Dec 2023 peak), driven by the Health and Care Worker route collapse (Caring Personal Service grants fell from 108,000 to 1,400). 94,000 asylum claims (-12 percent YoY), 128,000 decisions (+32 percent), grant rate 39 percent (down from 49 percent). Backlog 49,000 (-55 percent). 39,000 returns (+7 percent). 23,000 entered detention.
- EU Settlement Scheme: 371,000 settled-status grants in YE Mar 2026, up 12 percent. Not counted as immigration in LTIM because it confirms status for people already in the UK.
Caveats: ONS LTIM provisional figures are routinely revised. YE December 2024 was revised down by 100,000 (-23 percent) versus its initial estimate, almost entirely because of methodology change rather than underlying trend. Visa overstayers without asylum claims are assumed emigrated under the current method; irregular migrants who do not claim asylum are not counted at all. See the net migration page for the full discussion.
Validation
The primary test fits on 2001 to 2011 and forecasts 2021. Cohort change ratios are built from the Census 2001 and Census 2011 populations, projected one full decade forward, and scored against the actual Census 2021. The fitting window never touches the target, so what comes out is a real forecast error. It is also the exact analogue of what the published model does, which is fit on 2011 to 2021 and project to 2031 and beyond, so the error it reports is the error the published projections carry. 285 local authorities have an unchanged code across all three censuses and carry a Census 2021 observation; they are the scored set. Ratios are fitted at the 16 groups common to all three censuses and scored on the six broad groups, which are the only ones stable from 2001 through 2021.
Result on the current settings: MAE 1.56pp on the White British share, with a mean bias of +0.05pp. The forecast is close to unbiased, which matters more than the MAE because bias compounds across projection steps and noise does not.
Two details that change the number, so they are stated rather than buried. First, the ratios in this test are fitted at 16 ethnic groups, the finest classification common to 2001, 2011 and 2021. An earlier version fitted at six broad groups and reported 1.53pp with a bias of +0.03pp. Six groups is a coarser model than the 20 this site publishes, and the shrinkage constant is a cell count rather than a proportion, so a setting selected on the coarse fit shrinks the fine model considerably harder. Refitting at 16 groups moved the unbiased ceiling from 1.60 to 1.65 and the honest error from 1.53pp to 1.56pp. Second, the settings were selected on this same 285-area set. A split-half check, selecting the ceiling on half the areas and scoring on the other half, put the optimism from that at 0.02pp, which is small because the optimum is a wide plateau rather than a spike.
| Group | MAE | Bias | RMSE |
|---|---|---|---|
| White British | 1.56 | +0.05 | 2.35 |
| White Other | 1.04 | -0.87 | 1.35 |
| Asian | 0.80 | +0.45 | 1.31 |
| Black | 0.66 | +0.20 | 1.32 |
| Mixed | 0.69 | +0.68 | 1.04 |
| Other | 0.59 | -0.51 | 0.86 |
This test reversed what we previously told readers. An earlier version of this page reported a backcast score of 1.71pp and a bias of +1.70pp, and concluded that the model over-predicted the White British share and therefore understated the pace of change. Both the number and the direction were wrong. The backcast fitted its ratios on the same two Censuses it was tested against, so with the guardrails removed it reproduced the target to 0.14pp by construction: it was measuring how far the model's own guardrails pulled a circular fit away from an answer it already contained, not predictive skill. On the genuine out-of-sample test the previous settings scored MAE 2.82pp with a bias of -2.13pp, under-predicting White British in 192 of 285 areas. The model was projecting change too fast, not too slow.
The guardrails were the problem, and they were chosen on this evidence. Two rules bound each ratio. The old pair, a ceiling of 5.0 and a freeze to no-change for any cell whose 2011 base held five people or fewer, is what produced both the bias and the runaway long-horizon projections: a ceiling of 5.0 lets a group quintuple in a decade, which is 625 times over four steps. The current settings shrink each local ratio toward the national ratio for its group, age and sex in proportion to how much data it rests on, and cap growth at 1.65 per decade.
| Setting | MAE | Bias |
|---|---|---|
| Ceiling 5.0, freeze at 5 (previous) | 2.82 | -2.13 |
| Ceiling 3.0, no freeze | 2.61 | -2.35 |
| Ceiling 2.0, no freeze | 1.69 | -0.96 |
| Shrinkage K=25, ceiling 1.60 | 1.58 | +0.19 |
| Shrinkage K=25, ceiling 1.65 (current) | 1.56 | +0.05 |
| Shrinkage K=25, ceiling 1.80 | 1.55 | -0.33 |
The optimum is a plateau rather than a knife edge: every ceiling between 1.6 and 2.0 scores within about a tenth of a percentage point. 1.65 was chosen as the point where the forecast is unbiased, not the point that minimises MAE by a hundredth.
The largest thing this validation does not establish is the horizon. It tests a single ten-year step, which is what the model's first step does. The published 2051 figure runs that step three times and 2061 runs it four, and no data available here can test whether a calibration chosen on one step still holds over four. A growth ceiling is a per-step correction; applying the same one at every step assumes the tendency it corrects for does not itself change with distance. That assumption is untested. Read 2031 as the best-evidenced year on this site, 2051 as materially more uncertain than its error bar suggests, and 2061 as illustrative.
Comparison with NEWETHPOP. NEWETHPOP (Rees, Wohland et al., University of Leeds) projected 2021 from a 2011 base and scored MAE 3.95pp on the White British share across 296 areas, over-predicting in 282 of them. That is a genuine out-of-sample forecast error and is comparable with the 1.56pp above, which is also out-of-sample. The previously published head-to-head, which set NEWETHPOP against this model's circular backcast and claimed a 33% win, was not comparable and has been withdrawn.
School Census check, and why it is not independent. DfE School Census data gives a check for ages 4-15: MAE 2.36pp across 126 areas. That headline averages all six groups and flatters the one this site reports on; White British specifically is 6.20pp, the worst of any group. It is also not out-of-sample, because the forward model consumes the school census as a calibration input for ages 0 to 5. Treat it as a consistency check, not a test.
Why we publish two models
Hamilton-Perry is the central published projection. Alongside it, every place page shows the 2051 endpoint of a second, independent model: a classical cohort-component projection with births by ethnicity-specific total fertility rate and a half-convergence assumption (ethnic TFRs move halfway to the national mean by 2061).
The two models share the same Census 2011 and Census 2021 base. They differ in what they assume about the future:
- Hamilton-Perry (HP) extrapolates the 2011 to 2021 cohort change ratios forward. It is descriptive: it says "the demographic dynamics of the last decade continue." It captures observed migration, fertility, and mortality jointly through CCRs but cannot decompose them.
- Cohort-component (CC) models births, deaths, and migration separately, with explicit ethnicity-specific fertility rates. The half-convergence variant assumes higher-fertility groups gradually converge toward the national mean. This pulls projections toward a slower pace of change.
Because CC has explicit fertility convergence and HP does not, CC typically projects a higher White British share by 2051. The two-model spread for an area is the most honest single-number measure of structural model uncertainty: bigger than the HP-internal Monte Carlo confidence band, because it captures disagreement between methods rather than noise within one method.
Across 318 English LAs, the median 2051 spread is approximately 8 percentage points; the largest spreads exceed 20pp (typically high-diversity urban areas where the cohort-component fertility-convergence assumption diverges most sharply from observed CCR dynamics).
HP is treated as central because it requires fewer assumptions (it does not impose convergence) and validates better than NEWETHPOP, an established cohort-component model trained on the same Census data, on the only test that is genuinely out-of-sample for both: MAE 1.56pp across 285 areas against NEWETHPOP's 3.95pp across 296. See the validation section above, including what that test does not establish about the horizon.
Uncertainty quantification
1,000 Monte Carlo simulations with stochastic perturbation (sigma = 0.02) generate 80% and 95% confidence intervals for all projections. This captures the range of plausible outcomes within the Hamilton-Perry method, not a single point estimate.
The bands and the projections are produced by two separate jobs, and keeping them in step has taken work. They were not two runs of one model: the stochastic script carried its own copy of the ratio construction, its own shrinkage constant, and its own handling of the population envelope beyond 2047. The published estimate used to fall outside its own 80% band in 71% of area-years. Both scripts now read the same settings and treat the envelope the same way. They were also built on different 2011 Census populations, which is the largest reason the bands disagreed: cohort change ratios are the 2021 population over the 2011 population, so a different 2011 base is a different model. Both now load that base from one shared module. The projection falls outside its own band in 8% of area-years, down from 71%, and those cases are small: 45 of the 68 are under a quarter of a percentage point.
Where a band does not contain the projection it is drawn around, the band is not shown. That is a display guard, not a repair, and it means an absent band on a place page indicates the two runs disagree for that area rather than that uncertainty is unknown.
For structural uncertainty, the disagreement between modelling approaches, see the two-model comparison above. The two-model spread is generally the larger of the two uncertainty signals and the one to weight more heavily when reading any single projection.
Hand adjustments applied to the raw ratios
The projection is not a pure extrapolation. Three adjustments sit on top of the observed cohort change ratios:
- Shrinkage toward the national ratio. Each local ratio is pulled toward the national ratio for its group, age and sex, weighted by the size of the 2011 cell it rests on. A cell holding thousands of people keeps its own ratio; a cell holding two borrows the national one. This replaced an older rule that froze any cell of five people or fewer at no change, which held thin minority populations flat while the majority moved.
- A growth ceiling of 1.65 per decade. Selected on the out-of-sample test above. The previous ceiling of 5.0 permitted a group to quintuple in a decade and was the direct cause of the runaway long-horizon projections.
- Brexit White Other damp. White Other growth above 1.0 is cut by 15% at ages 10 to 34. This is a judgement about EU migration after free movement ended, not an observation, and it is the one adjustment the out-of-sample test cannot referee: that test forecasts 2021, so the post-Brexit period sits inside its target window rather than beyond it. Its effect is not small. Removing it would raise the projected White Other share in 2051 by 0.42 percentage points on average and up to 1.48, and lower White British by 0.26 on average and up to 1.23. It is kept because the end of free movement is a real discontinuity that cohort change ratios drawn from 2011 to 2021 cannot see. But it is a hand adjustment that moves the headline figure on this site in one direction, and it pushes the same way as an error the method already makes: on the out-of-sample test the undamped model under-projects White Other by 0.87pp. Run the model with
BREXIT_DAMP=0to see the difference. - DfE School Census calibration. Ratios for ages 0 to 5 are nudged toward the observed school-census ethnic mix, damped to 20% of the implied correction. This is why the school census cannot also serve as an independent validation of the model.
The first two are set by evidence. The last two are judgements, and neither can be tested with the data available here.
Reproducibility
The model inputs are not distributed with this site. The Census 2021 custom dataset base, the Census 2011 DC2101EW extract, the NEWETHPOP archive and the ONS SNPP file all sit outside version control, so a fresh checkout cannot regenerate the published projections. The model code is in the repository and the outputs are in the repository; the bridge between them is not. Until that is closed, the projections on this site cannot be independently reproduced by a reader and should be read as a published result rather than a verifiable one.
What can be checked without the inputs is internal consistency, and a guard script does that on every published release: group shares summing to 100, projections that run away from their own 2021 base, confidence bands that fail to contain the estimate they annotate, and agreement between the headline model and the scenario fields.
Known limitations
- No explicit international migration component. Cohort change ratios absorb migration jointly with fertility and mortality and cannot separate them. The 2011 to 2021 window carries roughly 300K net migration a year while post-2021 net migration peaked at 944,000 in the year ending March 2023, so the level the model extrapolates is not the level actually observed since. Which way that pushes the answer is genuinely unclear, and the out-of-sample test is the only evidence here: on it the model with the previous settings ran too fast, not too slow.
- Long-horizon divergence, largely fixed at source. With the old ceiling of 5.0 the ratios compounded without any effective bound and produced projections that were arithmetic rather than demographic: 177 area-years put a group above a quarter of the population and at least three times its 2021 share, Enfield's Other category reaching 67% by 2051 and 82% by 2061, and Barnsley's White Other going from 4.3% to 45.3%. The ceiling selected on the out-of-sample test cuts that to a handful of area-years. Those remaining years are still withheld rather than published, and a place page truncates at the last year the model was projecting rather than diverging. The guard is a backstop now, not the main defence.
- Per-group confidence intervals are not computed; the stochastic run covers the White British share only.
- 2061 is projected for 269 of 318 areas. National 2061 aggregates cover that subset only.
- No out-of-sample validation beyond DfE school data, and that validation uses the 2024/25 school census. The 2025/26 release is now available and has not yet been incorporated.
- Not peer-reviewed (submission to Demographic Research planned).
- COVID period (2020-21) may distort the 2021 Census baseline for some groups and areas.
Evidence standard
All figures on this site are categorised by evidence quality:
- Official: Published ONS, Census, DfE, or parliamentary data with stable provenance.
- Derived: Computed from official sources using documented methods (e.g., CCRs from two Census observations).
- Modelled: Output of the projection model. Inherently uncertain, always shown with confidence intervals where available.
- Estimated: Best available approximation where no direct data exists. Always flagged.