This is a technical discussion draft relating to the Population Replacement Evidence post. https://www.sarsen.org/2026/08/population-replacement-evidence.html
Summary
Two studies appear to disagree about whether ancestry from
Neolithic Britain persisted after the Bell Beaker migrations of the mid-third
millennium BC. Booth et al. (2021) reported a small but rising
Neolithic-related component through the Chalcolithic and Early Bronze Age.
Olalde et al. (2026), using a source population from the Lower
Rhine–Meuse that was unavailable in 2021, found the main English Beaker group
indistinguishable from that source, with any surviving British Neolithic
ancestry bounded between zero and eight per cent.
Re-analysing the supplementary data of both papers shows
that they do not disagree about how much ancestry survived. Placing Booth's
individual-level estimates inside Olalde's group definitions returns 2.52 per
cent for the Beaker group and 8.55 per cent for the later Chalcolithic–Early
Bronze Age group, against Olalde's zero and 7.3–7.9 per cent. Both papers also
find the same rise over time, and the rise turns out to be a continuous
gradient rather than a step between burial categories.
What the 2026 data cannot establish is where that ancestry
came from. Seven candidate source populations — from Yorkshire to Andalusia —
return estimates spanning only 7.3 to 9.0 per cent and fit the data comparably
well. The English Neolithic source ranks third of seven. The reason is
substantive rather than technical: English Neolithic populations and
continental Middle/Late Neolithic populations are genuinely alike in their
hunter-gatherer content, which is the only axis on which this design can separate
them.
The quantity is well measured. The address is unknown, and
no amount of additional sampling in the candidate source regions will change
that. One extension of an analysis already performed in the 2026 paper —
identity-by-descent between the English Chalcolithic–Early Bronze Age
individuals and the English Neolithic ones — could settle it.
1. Why this matters, and why it is hard
The genetic transformation of Britain between roughly 2450
and 2000 BC is one of the most dramatic events in European prehistory. People
associated with the Bell Beaker complex arrived from the continent, and within
a few centuries the ancestry of the population buried in Britain had changed by
something like ninety per cent. What is disputed is not that this happened but
what it means: whether a resident population was displaced, absorbed, or had
already dwindled to the point where "replacement" is the wrong word.
The residual matters more than its size suggests. If a few
per cent of Neolithic British ancestry survived into the Bronze Age, some
Neolithic families had descendants and the transition involved intermarriage
rather than pure displacement. If that residual rose over time, as Booth and
colleagues argued, then descendants of the Neolithic population were not merely
surviving but were increasingly represented — which is very hard to reconcile
with rapid violent replacement. And if the residual is actually zero, the
record is silent on all of this, because you cannot infer anything from a
signal you cannot see.
Two obstacles make the question hard, and both are
statistical rather than archaeological.
The first is that these ancestry proportions are estimated
by a method — qpAdm — that answers a narrow question. It asks whether a
proposed mixture of source populations is rejected by the data, given a
set of reference populations held out for comparison. A model that is not
rejected is not thereby confirmed. A proportion estimated at zero may mean the
ancestry is absent or may mean the method cannot see it. Distinguishing these requires
knowing how small a contribution the design could have detected, which nobody
has calculated for this case.
The second obstacle is that ancestry proportions are
estimated for a named source. When you ask how much English Neolithic ancestry
a population carries, you get an answer conditioned on English Neolithic being
the right source. If a different but genetically similar population were the
true contributor, the estimate would come back much the same. Booth and
colleagues flagged exactly this, writing that it was difficult to distinguish
ancestry from Neolithic Britain from ancestry from more western parts of continental
Europe — northern France in particular — where archaeogenetic data were then
almost absent. They judged that this was unlikely to account for much.
That judgement was the weakest joint in their argument. The
2026 paper supplies part of the missing data. What follows tests both.
2. Data and method
Three published supplementary datasets are used. All figures
below can be regenerated from them by the accompanying script.
Booth et al. 2021, Supplementary Table 1
gives, for each of 98 individuals from Chalcolithic to Late Bronze Age Britain,
a point estimate of ancestry related to the Neolithic populations of Britain, a
standard error, and a calibrated radiocarbon median. These estimates originate
in the qpAdm models of Olalde et al. 2018.
Olalde et al. 2026, Supplementary Tables 1, 3, 7,
9, 10, 12 and 14 give group assignments by genetic identifier, qpAdm models
for the English groups with the Lower Rhine–Meuse source in place, distal
three-way models decomposing ancestry into Balkan Neolithic, western
hunter-gatherer and Corded Ware components, and identity-by-descent
connections.
Olalde et al. 2018, Supplementary Table 1
gives autosomal SNP counts.
Merging on individual identifiers places Booth's
per-individual estimates inside Olalde's 2026 group boundaries. Coverage is 26
of the 28 individuals in England_BB, 28 of the 39 in England_CA_EBA, and 3 of
the 4 in England_BB_highEEF — I14200 is not in Booth's table. The remaining 44
Booth individuals are Scottish, Welsh, northern English or Middle-to-Late
Bronze Age and fall outside the English groups; they are used only where the
temporal analysis is stated as covering all dated individuals.
Group summaries are inverse-variance weighted means.
Heterogeneity is assessed by Cochran's Q, I², and the
random-effects between-individual standard deviation τ. Temporal models are
weighted least squares of ancestry proportion on calibrated median date, with
the residual scale factor φ = χ²/df used to correct slope standard errors for
over- or under-dispersion.
3. The two papers agree about magnitude
|
Group |
Booth per-individual, pooled |
Olalde 2026, group-level qpAdm |
|
England_BB (26 of 28) |
2.52 % ± 0.84 |
0 % — cladal with RhineMeuse_LNB_BB, P = 0.61 |
|
England_CA_EBA (28 of 39) |
8.55 % ± 0.85 |
7.3–7.9 % (setup 1); 7.5–9.0 % (setups 2–3) |
|
England_BB_highEEF (3 of 4) |
26.98 % ± 2.44 |
27–39 % in three-way models |
The estimates are not in conflict. They are the same
numbers, reached through different model structures — one without a Rhine–Meuse
proxy and one with it — and the concordance holds across an order-of-magnitude
range. Whatever else the 2026 paper achieved, it did not overturn Booth's
measurements.
Two qualifications, both of which tighten the agreement
rather than loosening it. Booth's per-individual standard errors derive from a
common set of f₄ statistics and
are correlated; treating them as independent makes the pooled standard errors
above optimistic, plausibly by a factor of order √2. And 19 of the 98 estimates
sit exactly on the zero boundary, which biases the pooled means upward.
4. The zero boundary is diagnostic
qpAdm can return negative admixture proportions. Olalde 2018
reported them censored at zero. This is normally a nuisance; here it carries
information.
If a group's true British Neolithic ancestry were exactly
zero, half its individual estimates would fall below zero by chance and be
reported as zero. In England_BB, 13 of 26 estimates sit on the boundary
— a fraction of 0.500 against an expectation of 0.500 (binomial p =
1.00). In England_CA_EBA, 4 of 28; under a true value of zero the
expectation is 14 (p = 1.8 × 10⁻⁴).
Simulating the censoring directly — drawing each individual
from a normal centred on an assumed common true value with that individual's
own reported standard error, censoring at zero, re-pooling, 200,000 times:
|
Assumed true value |
England_BB pooled, median (95 %) |
Expected zeros |
P(pooled
≥ 2.52 %) |
|
0 % |
1.66 % (0.81–2.72) |
13.0 |
0.053 |
|
1 % |
2.21 % (1.21–3.39) |
10.6 |
0.294 |
|
2 % |
2.85 % (1.72–4.15) |
8.4 |
0.706 |
|
3 % |
3.58 % (2.33–4.98) |
6.4 |
0.949 |
Inverting the simulation gives a consistency interval of 0
to 3.25 per cent for England_BB's true common value. Truncation accounts
for most but not all of the observed 2.52 per cent: under a true value of
exactly zero the expected pooled estimate is 1.66 per cent, and the observation
sits at the 95th percentile. Zero survives, but only just.
For England_CA_EBA the same simulation excludes every true
value below about 7 per cent (P < 10⁻³ throughout); at 7.15 per cent
the observed pooled estimate sits at p = 0.058. The later group's
residue is real and is not an artefact of censoring.
So Booth's own data, read through nothing but their boundary
counts, say that the Beaker group's Neolithic ancestry is indistinguishable
from zero and the Chalcolithic–Early Bronze Age group's is not. That is the
2026 result, obtainable five years earlier from data already in hand.
5. The rise is real, and it is a gradient
England_BB has a median calibrated date of 2146 cal BC;
England_CA_EBA, 1865 cal BC (Mann–Whitney p = 1.8 × 10⁻⁵). Olalde 2026
assigns the earlier group no additional Middle/Late Neolithic ancestry and the
later group 7.3–7.9 per cent. That is a rise, in the same direction and of
similar magnitude to Booth's reported increase from about 6 to about 12 per
cent, obtained from independent model specifications with the correct proxy in
place.
Booth's trend also survives being tested more stringently
than they tested it. A rank-sum comparison across a split chosen partly by
inspection is weak evidence; weighted regression on the continuous date
variable is not.
|
Sample |
n |
Slope (pp/century) |
Z (naive) |
Z (corrected) |
φ |
|
All dated individuals |
87 |
+0.62 |
5.84 |
3.35 (p = 0.0008) |
3.03 |
|
Excluding the highEEF outliers |
84 |
+0.82 |
7.52 |
5.20 (p < 10⁻⁶) |
2.09 |
|
England_BB + England_CA_EBA only |
48 |
+1.83 |
5.69 |
4.15 (p = 3 × 10⁻⁵) |
1.88 |
|
…excluding I2462 as well |
47 |
+1.73 |
5.39 |
6.85 (p < 10⁻⁵) |
0.62 |
Spearman's ρ is 0.509 (p = 4.9 × 10⁻⁷) across all
dated individuals and 0.592 (p = 9.5 × 10⁻⁶) on the restricted set.
Restricting to the 48 individuals inside the two English groups removes any
confounding from Scottish, Welsh or Middle Bronze Age individuals, and
strengthens the slope roughly threefold — the full-sample figure is depressed
by the later Bronze Age tail, where the trajectory plateaus.
The rise is a gradient, not a step. Within England_BB
alone the slope is +1.26 pp/century (corrected Z = 1.77); within
England_CA_EBA alone, +1.38 (corrected Z = 1.92). Fitting date and a
group indicator together, the date term survives (Z = 5.02 excluding
I2462) and the group step does not (+1.5 pp, Z = 1.24). Olalde's two
groups are not two states of the population but two windows onto one continuous
trajectory. This is the strongest available vindication of Booth's actual
thesis, which was gradualism: a group-level model with two categories cannot represent
a gradient, and the gradient is in the data.
6. What cannot be identified is the address
Olalde 2026 tested nine candidate second sources for
England_CA_EBA across three SNP-filtering setups. The two Lower Rhine–Meuse
sources fail decisively. The remaining seven:
|
Second source |
n |
P (setup 1) |
P (setup 2) |
P (setup 3) |
Proportion range |
|
TRB_N (Funnel Beaker) |
6 |
7.4 × 10⁻³ |
0.514 |
0.488 |
7.3–8.2 % |
|
S_France_LN |
18 |
1.8 × 10⁻³ |
0.220 |
0.351 |
7.3–8.6 % |
|
England_N |
27 |
2.2 × 10⁻³ |
0.205 |
0.334 |
7.6–8.8 % |
|
Germany_Baalberge_MN |
3 |
1.4 × 10⁻³ |
0.115 |
0.244 |
7.8–8.9 % |
|
NE_France_LN |
14 |
1.2 × 10⁻³ |
0.083 |
0.199 |
7.4–8.6 % |
|
GlobularAmphora_LN |
10 |
6.3 × 10⁻⁴ |
0.051 |
0.099 |
7.9–9.0 % |
|
N_Spain_LNCA |
45 |
3.1 × 10⁻⁴ |
0.037 |
0.100 |
7.6–8.7 % |
Across all 21 fits the point estimate ranges from 7.3 to 9.0
per cent — a spread of 1.7 percentage points — while the candidate source
ranges geographically from Yorkshire to Andalusia. The quantity is estimated to
within about two points; the source is not estimated at all. England_N ranks
third of seven where models pass. There is no statistical basis for preferring
the British source and none for excluding it.
This is why Booth's caveat matters more, not less, after
2026. They worried that western continental Neolithic populations could not be
told apart from British ones. NE_France_LN returns 7.4–8.6 per cent with fits
comparable to England_N's. Their caveat has been tested and confirmed as
unresolvable with this design, not refuted.
The 2026 bound of 0 to 8 per cent is a statement of that
non-identifiability. It is not a confidence interval and should not be
collapsed into one. Sampling uncertainty on the quantity is roughly ±2.4
points; the 0-to-8 range is source-attribution ambiguity, a different kind of
ignorance that does not shrink with more British individuals.
7. What the cladality result does and does not
bound
England_BB's cladality with RhineMeuse_LNB_BB at P =
0.61 licenses a bound, not an absence. The bound can be calculated (Appendix
A).
With six outgroups the one-way model carries 5 degrees of
freedom, and P = 0.61 corresponds
to an observed χ² of 3.61. Inverting the non-central χ² gives a one-sided 95
per cent upper bound on the non-centrality of λ ≤ 6.88. The mapping λ ≈
κ(α/SE)² is calibrated against the two groups that demonstrably require a
second source, giving κ = 0.96 across six checks. The standard error for a
two-source England_BB model, obtained by cross-target calibration against
Supplementary Table 10, is 0.0130–0.0140 — about 8 per cent larger
than England_CA_EBA's, since England_BB's higher per-individual coverage does
not compensate for eleven fewer individuals. Hence:
England_BB can conceal up to
3.5 per cent local Neolithic ancestry (95 per cent upper bound; 3.46–3.75 per
cent across the three SNP setups).
The power curve matters as much as the bound. The design
reaches 50 per cent power only at 3.5 per cent, 80 per cent power at 4.7–5.1
per cent, and 95 per cent power at 5.9–6.4 per cent. A genuine residue of four
per cent would have been missed more often than not.
Booth's pooled England_BB estimate of 2.52 per cent, and the
0–3.25 per cent consistency interval from the censoring simulation, sit
entirely inside that window. Two independent routes converge on the same
ceiling. The two papers do not disagree here either; the 2026 design simply
cannot see a residue this small.
8. Structure inside the groups
Cochran's Q on Booth's estimates within each Olalde
group:
|
Group |
Q |
df |
p |
I² |
τ |
|
England_BB |
20.9 |
25 |
0.70 |
0 % |
0.00 pp |
|
England_CA_EBA |
81.0 |
27 |
2.6 × 10⁻⁷ |
66.7 % |
6.36 pp |
England_BB, once Olalde's four outliers are removed, is
statistically homogeneous. Olalde's decision to excise England_BB_highEEF is
therefore vindicated on Booth's own numbers: those four individuals were the
entirety of the Beaker-period signal, and they survive independent testing with
a proper Rhine–Meuse proxy.
England_CA_EBA is not homogeneous, and its heterogeneity has
a single cause. Removing one individual — I2462, Sk 220053 from East Kent
Access Road, 35.3 ± 3.8 per cent, z = +7.04, dated 2131–1890 cal BC
— leaves Q = 28.9 on 26 df (p = 0.32), τ = 1.51 pp, and a group
mean of 7.15 per cent in close agreement with Olalde's 7.3–7.9. Booth
identified her as the latest individual carrying substantial British Neolithic
ancestry. Olalde 2026 separated four outliers from the Beaker group and
retained this statistically comparable fifth inside a 39-individual pool where
she contributes about 0.7 points to the group mean. That is not an error — a
group-level analysis is entitled to pool — but it means the count of
individually demonstrable Neolithic-descended people in the record is five, not
four.
The rise, however, is not carried by the outliers.
Decomposing the 4.72-point gap between the two groups' unweighted means:
|
Individuals removed from England_CA_EBA |
Group mean |
Gap retained |
|
none |
7.40 % |
4.72 pp (100 %) |
|
I2462 |
6.36 % |
3.68 pp (78 %) |
|
top 2 |
6.06 % |
3.38 pp (72 %) |
|
top 3 |
5.76 % |
3.08 pp (65 %) |
|
top 5 |
5.13 % |
2.45 pp (52 %) |
The Hodges–Lehmann shift between the groups is 3.90 points,
and a quarter of England_CA_EBA individuals exceed England_BB's 95th
percentile. This is a broad distributional shift, not a handful of surviving
families visible against an unchanged background — a distinction that matters,
because a broad shift is hard to explain by lineage survival and easy to
explain by continuing gene flow from wherever the source turns out to be.
Two further features of England_CA_EBA, both invisible at
group level. The within-group residual correlates with date even after I2462 is
removed (r = +0.50, p = 0.009), so the gradient of §5 runs inside
the group as well as between the groups. And there is a hint of regional
structure: Amesbury Down (n = 5) averages 2.10 per cent against East
Kent Access (n = 3) at 10.43 per cent. At these sample sizes that is a
hypothesis rather than a finding, but it is precisely the kind of pattern a
pooled model is guaranteed to erase.
9. The reconciliation
Both papers are correct. They are describing different
phases of one sequence, and the appearance of conflict comes from a difference
in what each was able to name.
The Beaker phase (England_BB, median 2146 cal BC).
Local Neolithic ancestry is indistinguishable from zero, at a detection ceiling
of about 3.5 per cent. Booth's boundary structure and Olalde's cladality test
say the same thing. Individual survivals exist but are confined to four
outliers, which are real.
The Chalcolithic–Early Bronze Age phase (England_CA_EBA,
median 1865 cal BC). A second Middle/Late Neolithic component of 7 to 9 per
cent is required, robustly, in both analyses. The increase from the earlier
phase is real, is reproduced by Olalde's own models, is a continuous gradient
rather than a step, and is a broad shift rather than an outlier effect.
The source of that component is unidentified and, with
this design, unidentifiable. It lies somewhere between 0 and 8 per cent
British. Booth's inference that it represents a resurgence of specifically British
Neolithic ancestry is unsupported; so is the inference that it represents
continental ancestry arriving pre-mixed. The honest statement is that the
ancestry rose and we do not know whose it was.
The consequence for the wider argument is uncomfortable for
both sides. A rising Neolithic-related component after about 2100 BC no longer
functions as evidence against rapid displacement, because a continental origin
for that component is fully consistent with the data and would say nothing
about British survival. It also does not function as evidence for displacement.
The quantity is well measured and interpretively inert until the source is
pinned down.
10. What would resolve it
Four interventions suggest themselves. Three can be tested
against the supplementary data, and the results (Appendix D) reorder them
substantially.
1. Identity-by-descent between England_CA_EBA and
England_N. The only proposal that can work, it is a one-line extension of an
analysis already performed, and it would be decisive. Supplementary Table
14 reports 5,846 IBD pairs. Checked against the table's own population labels, England-by-England
pairs number zero — not because none exist, but because the table's scope
is pairs featuring a Lower Rhine–Meuse individual, so the comparison was never
attempted. Shared segments carry information about specific recent common
ancestors that no f₄ ratio supplies, and this is the one axis on which a
British source and a continental one must differ.
Appendix F works out what the test would yield. In outline:
the method has ample resolution at the required time depth (nineteen detected
pairs in the table span date gaps above 1,600 years), the near-contemporary
full-ancestry detection rate is 17.0 per cent of possible pairs, and the
expected yield for a 7.6 per cent ancestry component runs from 3.4 to 13.6
detections at full group coverage depending on how severely one assumes segment
detection decays across 1,600 years. Against a continental-source null that
predicts no England_N-specific excess, that is a decisive test. At the coverage
actually used in Supplementary Table 14 — 19 English Neolithic and 11 English
Chalcolithic–EBA individuals — the expected yield is 0.7 to 2.7 and the test is
underpowered.
Three design points. Relatedness among the English Neolithic
individuals is substantial: at least eight of the nineteen belong to one
annotated pedigree, so the effective number of independent pairs is well below
the nominal count and the sample should be pruned. The coverage threshold
admitting only 12 of 39 England_CA_EBA individuals will need relaxing. And the
arm against England_N should be run simultaneously against date-matched
continental candidates so that time decay, the largest nuisance parameter,
cancels in the ratio — Ireland_MN (n = 8, median 5410 BP) and
MN_Wartberg (n = 26, median 5140 BP) are the best-matched to England_N's
median of about 5500 BP and are already in the table.
2. Individual-level re-analysis of England_CA_EBA with
the Rhine–Meuse proxy. The group-level result conceals one demonstrable
survival, a within-group temporal gradient, and possible regional structure.
Section 8 does what can be done from Booth's estimates; doing it properly
requires running the 2026 models per individual.
3. A rotating qpAdm competition with a right set chosen
to separate the candidate sources. Less promising than it looks, and Appendix D
explains why. The 2026 paper contains an orthogonal discriminator it does
not apply here: the ratio WHG/(WHG + Balkan_N) from the distal three-way
models. That statistic separates the Rhine–Meuse sources (0.397–0.492) from
everything else very cleanly, and the paper uses it to good effect. But
England_N's own value, computed from the 14 England_N individuals in
Supplementary Table 7, is 0.207 with an individual range of 0.180 to
0.230 — sitting inside the continental Middle/Late Neolithic cluster at
0.217 to 0.246. English Neolithic and continental Middle/Late Neolithic
populations are genuinely alike in hunter-gatherer content, to within about the
individual scatter. The seven candidates behave alike in the qpAdm models
because they are alike on the axis this design resolves. A better right
set helps only if it resolves something other than hunter-gatherer level, and
it is not obvious what that would be. The non-identifiability is substantive,
not a defect of outgroup choice.
It is worth noting separately that England_CA_EBA does not
appear in Supplementary Table 9 at all. The one orthogonal statistic in the
paper was never computed for the group that carries the residue.
4. More northern French Middle/Late Neolithic genomes.
This will not help. Across the seven candidate sources, group size ranges
from 3 to 45 — a fifteen-fold span — and the standard error on the
second-source proportion is 0.011 to 0.013 throughout. There is no relationship
between how many individuals a source group contains and how well it fits:
N_Spain_LNCA at n = 45 ranks sixth or seventh, Germany_Baalberge_MN at n
= 3 ranks fourth, TRB_N at n = 6 ranks first (Spearman ρ between source n
and log P = −0.32, p = 0.48 in setup 2; −0.18, p
= 0.70 in setup 3). Sampling density in the source is not the binding
constraint. More French genomes are worth having for other reasons — a northern
French population genuinely distinct from those sampled would change the
picture — but they will not tighten this estimate.
11. Conclusion
The disagreement between these two papers was never about
arithmetic. Placed side by side under a common set of group definitions, they
return the same numbers to within their standard errors: near-zero surviving
Neolithic ancestry among the Beaker-associated dead, seven to nine per cent
among their Chalcolithic and Early Bronze Age successors, and a continuous rise
between the two that neither paper's model structure was designed to represent
but which both papers' data contain.
What separates them is a question about naming rather than
counting. Booth and colleagues measured a quantity and called it British;
Olalde and colleagues measured the same quantity and declined to call it
anything, because with a proper Rhine–Meuse proxy in the model, seven Neolithic
populations spread across western Europe fit it equally well. The 2026 bound of
nought to eight per cent is not a measurement with wide error bars. It is a
statement that the question has two answers and the data cannot choose.
That is a genuinely awkward place for the field to be, and
it is worth being clear about how awkward. The rising Neolithic-related
component after about 2100 BC has done a lot of interpretive work in recent
discussion of the Beaker transition — as evidence that displacement was
incomplete, that intermarriage was common, that the population of Neolithic
Britain left descendants. None of that is refuted here. But none of it is
supported either, because a component whose source is unknown cannot tell us whether
anyone survived. The number is solid and the inference is suspended.
The most useful thing to take from the exercise is that the
non-identifiability turns out to be real rather than fixable by more of the
same. Neolithic Britain and Neolithic northern France, Germany and Iberia are
alike in the one respect that this method resolves. Sequencing more of them
will not separate them.
What might is a technique that looks at shared segments of
chromosome rather than aggregate ancestry proportions — and here the news is
better than the rest of this paper. The relevant data already exist, in a table
in the 2026 paper's own supplement. The comparison it needs has never been run,
because that table was built to answer a different question and by construction
contains no English-by-English pairs at all. But the same table contains
everything required to work out what the comparison would yield: how well
identity-by-descent survives the sixteen centuries between Neolithic and Early
Bronze Age Britain, and how often a seven-per-cent ancestry component leaves a
detectable segment. The answer is that at the coverage used in the published
table the test would be underpowered, and at full group coverage it would be
decisive — somewhere between three and fourteen detections where a continental
source predicts essentially none.
That is an unusually cheap way to settle a question this
old. No new excavation, no new sequencing, no new samples: one comparison, on
data already published, between two sets of people who lived in the same
country sixteen hundred years apart.
Caveats
-
Booth's per-individual standard errors derive
from a common set of f₄
statistics and are correlated. Inverse-variance pooling assumes independence
and therefore understates the pooled standard errors, plausibly by around √2.
All pooled intervals should be read as wider than quoted.
-
Booth's estimates come from Olalde 2018's
models, which lacked a Rhine–Meuse source. They are not independent
measurements of the same parameter as Olalde 2026's; the agreement in §3 is
consistency under a change of specification, not replication.
-
Cochran's Q has low power at these sample
sizes. I² = 0 for England_BB is a failure to detect heterogeneity, not a
demonstration of its absence.
-
Residual z-scores fail normality tests in
both groups (Shapiro–Wilk W = 0.70 and 0.76), driven by the zero floor
in England_BB and by I2462 in England_CA_EBA. Normal-theory intervals
throughout are indicative; the permutation and simulation results are more
reliable.
-
11 of the 39 England_CA_EBA individuals and 2 of
the 28 England_BB individuals are newly reported in 2026 and absent from
Booth's table.
-
The
detection floor in §7 rests on the empirically calibrated mapping λ ≈ κ(α/SE)².
The six calibration checks scatter between κ = 0.83 and 1.21. A direct
recomputation would require the 1240k genotype data.
-
The censoring simulation in §4 assumes a single
common true value per group and independent errors. The first assumption is
falsified for England_CA_EBA (§8), so that interval applies to the group's
central tendency rather than to individuals.
-
qpAdm P-values across Supplementary
Tables 9–12 come from a large number of tested models. None quoted here has
been corrected for multiple testing, and none should be read as a probability
that a model is true.
-
The regional contrast in §8 rests on 5 and 3
individuals respectively and should not be relied on.
Appendix A — Detection floor for England_BB
Degrees of freedom. With n_right outgroups and n_left populations on the left, qpAdm's rank test has df
= 1 × (n_right − n_left + 1). Supplementary Tables 9–12 use six outgroups
(OldAfrica, IronGates_HG, Turkey_Neo, WSHG, CHG_Iran_N, Russia_Afanasievo), so
the one-way model has 5 df and the two-way model 4 df.
Table A1 — cross-target standard-error calibration.
Supplementary Table 10 runs one model family against several targets, allowing
standard errors to be compared with model structure held fixed. Values are
medians across the ten source pairings.
|
Target |
n |
Setup 1 |
Setup 2 |
Setup 3 |
|
England_BB |
28 |
0.0150 |
0.0160 |
0.0160 |
|
England_CA_EBA |
39 |
0.0140 |
0.0145 |
0.0150 |
|
England_BB_highEEF |
4 |
0.0220 |
0.0230 |
0.0230 |
|
RhineMeuse_LNB_BB |
13 |
0.0155 |
0.0160 |
0.0170 |
Ratio SE(England_BB)/SE(England_CA_EBA) = 1.071, 1.103,
1.067; mean 1.081. Applied to the Supplementary Table 12 standard errors for
England_CA_EBA (0.012, 0.012, 0.013), this gives 0.0130–0.0140 for a two-source
England_BB model.
Table A2 —
calibrating λ ≈ κ(α/SE)². For the two groups that demonstrably
require a second source, the one-way model's failure should have non-centrality
proportional to the squared Wald statistic of the omitted component.
|
Target |
one-way P |
implied χ²(5) |
α |
SE |
df + (α/SE)² |
κ |
|
England_CA_EBA (s1) |
2.79 × 10⁻¹⁰ |
53.4 |
0.076 |
0.012 |
45.1 |
1.21 |
|
England_CA_EBA (s2) |
1.49 × 10⁻⁹ |
49.8 |
0.088 |
0.012 |
58.8 |
0.83 |
|
England_CA_EBA (s3) |
3.82 × 10⁻⁷ |
38.0 |
0.079 |
0.013 |
41.9 |
0.89 |
|
England_BB_highEEF (s1) |
1.64 × 10⁻²¹ |
107.1 |
0.243 |
0.024 |
107.5 |
1.00 |
|
England_BB_highEEF (s2) |
1.01 × 10⁻²² |
112.9 |
0.260 |
0.025 |
113.2 |
1.00 |
|
England_BB_highEEF (s3) |
8.23 × 10⁻²⁰ |
99.1 |
0.241 |
0.024 |
105.8 |
0.93 |
Median κ = 0.96. The fit is near-exact for the high-signal
outlier group and scatters by roughly ±20 per cent for England_CA_EBA.
Table A3 — the bound and the power curve.
England_BB's observed one-way statistic is χ² = 3.61 on 5 df (P =
0.607); the critical value at α = 0.05 is 11.07.
|
Quantity |
λ |
Implied England_N ancestry |
|
One-sided 95 % upper bound on λ given χ² = 3.61 |
≤ 6.88 |
≤ 3.46–3.75 % |
|
λ giving 50 % power |
6.99 |
3.49–3.78 % |
|
λ giving 80 % power |
12.83 |
4.73–5.12 % |
|
λ giving 95 % power |
19.78 |
5.87–6.36 % |
Autosomal SNP coverage (Olalde 2018 Supplementary
Table 1): England_BB, 26 of 28 individuals with counts, median 663,686, range
14,794–913,255. England_CA_EBA, 28 of 39, median 491,782, range 17,178–891,333.
England_BB_highEEF, 3 of 4, median 700,532, range 136,956–729,987.
Appendix B — Censoring simulation
200,000 draws per row. Each individual is drawn from N(true,
SE_i²) using its own reported standard error, censored at zero, and re-pooled
by inverse variance.
|
True value |
England_BB pooled, median (95 %) |
E[zeros] |
P(≥
2.52 %) |
England_CA_EBA pooled, median |
E[zeros] |
P(≥
8.55 %) |
|
0 % |
1.66 (0.81–2.72) |
13.0 |
0.053 |
1.71 |
14.0 |
< 0.001 |
|
1 % |
2.21 (1.21–3.39) |
10.6 |
0.294 |
2.26 |
11.6 |
< 0.001 |
|
2 % |
2.85 (1.72–4.15) |
8.4 |
0.706 |
2.90 |
9.3 |
< 0.001 |
|
3 % |
3.58 (2.33–4.98) |
6.4 |
0.949 |
3.62 |
7.3 |
< 0.001 |
|
5 % |
5.25 (3.80–6.77) |
3.4 |
≈ 1 |
5.29 |
4.2 |
< 0.001 |
|
7.15 % |
7.24 (5.67–8.84) |
1.5 |
≈ 1 |
7.27 |
2.1 |
0.058 |
Observed: England_BB, 13 zeros of 26, pooled 2.52 %.
England_CA_EBA, 4 zeros of 28, pooled 8.55 %.
Inverted consistency interval for England_BB's true common
value: 0.00 % to 3.25 %.
Appendix C — Temporal analysis
Group chronology (Booth calibrated medians, cal BC):
|
Group |
n dated |
Median |
IQR |
|
England_BB_highEEF |
3 |
2289 |
2340–2190 |
|
England_BB |
21 |
2146 |
2261–2091 |
|
England_CA_EBA |
27 |
1865 |
2064–1814 |
England_BB earlier than England_CA_EBA: Mann–Whitney p
= 1.8 × 10⁻⁵.
Weighted regression on date. Slope in percentage
points per century; corrected Z uses the residual scale factor φ.
|
Sample |
n |
Slope |
Z naive |
Z corrected |
p |
φ |
|
All dated |
87 |
+0.62 |
5.84 |
3.35 |
8 × 10⁻⁴ |
3.03 |
|
Excluding highEEF |
84 |
+0.82 |
7.52 |
5.20 |
< 10⁻⁶ |
2.09 |
|
England_BB + CA_EBA |
48 |
+1.83 |
5.69 |
4.15 |
3 × 10⁻⁵ |
1.88 |
|
…excluding I2462 |
47 |
+1.73 |
5.39 |
6.85 |
< 10⁻⁵ |
0.62 |
|
England_BB alone |
21 |
+1.26 |
1.44 |
1.77 |
0.077 |
0.66 |
|
England_CA_EBA alone |
27 |
+1.38 |
3.19 |
1.92 |
0.055 |
2.75 |
Note that φ < 1 in two rows: the reported standard errors
there slightly overstate the scatter, so the corrected Z is
conservative.
Gradient versus step (date and a group indicator
fitted jointly):
|
Model |
Date slope |
Group step |
φ |
|
With I2462 |
+1.35 pp/century (Z = 2.60) |
+3.40 pp (Z = 1.63) |
1.81 |
|
Without I2462 |
+1.52 pp/century (Z = 5.02) |
+1.52 pp (Z = 1.24) |
0.61 |
Distribution-free checks. Spearman ρ = 0.509 (p
= 4.9 × 10⁻⁷) on all dated individuals; 0.592 (p = 9.5 × 10⁻⁶) on the
restricted set. Split at 2000 cal BC: earlier n = 39, mean 5.9 %; later n
= 48, mean 11.5 %; Mann–Whitney one-tailed p = 3.9 × 10⁻⁶. Restricted
Mann–Whitney, England_CA_EBA > England_BB: p = 0.0011. Permutation
test on the restricted slope excluding I2462, 20,000 draws: p < 5 ×
10⁻⁵.
Appendix D — Testing the four proposals
D.1 Source sample size does not predict fit
England_CA_EBA, two-way models with RhineMeuse_LNB_BB
(Supplementary Table 12), against source group sizes from Supplementary Tables
1 and 3:
|
Source |
n |
P s1 |
P s2 |
P s3 |
α s1 |
α s2 |
α s3 |
SE s1 |
|
MLN_Belgium |
18 |
1.3 × 10⁻⁹ |
1.4 × 10⁻⁸ |
2.2 × 10⁻⁶ |
3.2 % |
3.6 % |
3.7 % |
0.013 |
|
MN_Wartberg |
40 |
6.5 × 10⁻⁷ |
1.7 × 10⁻⁵ |
5.7 × 10⁻⁴ |
5.9 % |
6.8 % |
6.4 % |
0.013 |
|
Germany_Baalberge_MN |
3 |
1.4 × 10⁻³ |
0.115 |
0.244 |
7.8 % |
8.9 % |
7.9 % |
0.012 |
|
TRB_N |
6 |
7.4 × 10⁻³ |
0.514 |
0.488 |
7.3 % |
8.2 % |
7.5 % |
0.011 |
|
GlobularAmphora_LN |
10 |
6.3 × 10⁻⁴ |
0.051 |
0.099 |
7.9 % |
9.0 % |
8.3 % |
0.013 |
|
England_N |
27 |
2.2 × 10⁻³ |
0.205 |
0.334 |
7.6 % |
8.8 % |
7.9 % |
0.012 |
|
S_France_LN |
18 |
1.8 × 10⁻³ |
0.220 |
0.351 |
7.3 % |
8.6 % |
7.7 % |
0.012 |
|
NE_France_LN |
14 |
1.2 × 10⁻³ |
0.083 |
0.199 |
7.4 % |
8.6 % |
7.7 % |
0.012 |
|
N_Spain_LNCA |
45 |
3.1 × 10⁻⁴ |
0.037 |
0.100 |
7.6 % |
8.7 % |
7.9 % |
0.012 |
One-way model with RhineMeuse_LNB_BB alone: England_CA_EBA P
= 2.79 × 10⁻¹⁰, 1.49 × 10⁻⁹, 3.82 × 10⁻⁷. England_BB P = 0.607, 0.838,
0.881, 0.188 (setup 4).
Correlations between log₁₀ source n and log₁₀ P
across the seven non-Rhine-Meuse sources: Pearson r = −0.348 (p = 0.445) setup 2, −0.321 (p = 0.482) setup 3; Spearman ρ = −0.321
(p = 0.482) and
−0.179 (p = 0.702).
D.2 The hunter-gatherer discriminator
Group-level distal models (Balkan_N + WHG +
Germany_CordedWare), Supplementary Table 9, setup 1:
|
Group |
Balkan_N |
WHG |
CordedWare |
WHG/(WHG+Balkan_N) |
|
RhineMeuse_LNA_Vlaardingen/CW |
0.424 |
0.410 |
0.165 |
0.492 |
|
RhineMeuse_LNB_BB |
0.105 |
0.069 |
0.826 |
0.397 |
|
England_BB |
0.116 |
0.068 |
0.816 |
0.370 |
|
RhineMeuse_EBA |
0.146 |
0.066 |
0.788 |
0.311 |
|
France_BB_Steppe |
0.288 |
0.094 |
0.618 |
0.246 |
|
England_BB_highEEF |
0.267 |
0.084 |
0.649 |
0.239 |
|
Czechia_BB |
0.237 |
0.072 |
0.691 |
0.233 |
|
France_BB_NoSteppe |
0.755 |
0.220 |
0.026 |
0.226 |
|
SEGermany_BB |
0.252 |
0.070 |
0.677 |
0.217 |
England_CA_EBA does not appear in this table.
England_N's own composition, from the 14 England_N
individuals present in Supplementary Table 7: mean Balkan_N 0.785, mean WHG
0.204, ratio 0.207, individual median 0.207, range 0.180–0.230. That
places English Neolithic inside the 0.217–0.246 continental cluster and about
0.19 away from the Rhine–Meuse sources.
Limitation. Differencing England_BB against
England_BB_highEEF by mass balance gives the added component a ratio of
0.089–0.101 across setups, below England_N's 0.207 and below every candidate.
The three-way distal model is evidently not cleanly additive across groups with
different Corded Ware fractions (0.816 versus 0.649), so that figure should not
be treated as an estimate of the added component's true hunter-gatherer
content. The comparison that stands is the direct one between England_N and the
continental groups.
D.3 IBD coverage of the English individuals
Supplementary Table 14 contains 5,846 pairs with segments
above 12 cM. Pairs involving an English individual, by group:
|
Pair type |
Pairs |
|
England_BB × RhineMeuse_LNB_BB |
26 |
|
England_CA_EBA × RhineMeuse_LNB_BB |
20 |
|
England_BB × RhineMeuse_EBA |
7 |
|
England_N × RhineMeuse_MN_Tiel |
7 |
|
England_BB_highEEF × RhineMeuse_LNB_BB |
5 |
|
England_CA_EBA × RhineMeuse_EBA |
4 |
|
England_CA_EBA × RhineMeuse_MN_Wartberg |
2 |
|
England_BB × RhineMeuse_MLN_Belgium |
2 |
|
England_CA_EBA × RhineMeuse_LNA_Vlaardingen/CW |
1 |
|
England_BB × RhineMeuse_LNA_Vlaardingen/CW |
1 |
|
England_N × RhineMeuse_MN_Swifterbant |
1 |
|
England_N × RhineMeuse_MN_Wartberg |
1 |
|
England_BB × RhineMeuse_MN_Wartberg |
1 |
|
England × England |
0 |
Individuals available: England_BB 16 of 28, England_CA_EBA
12 of 39, England_N 6 of 27, England_BB_highEEF 1 of 4. The nine England_N ×
Rhine–Meuse Neolithic pairs have longest segments of 12.7–20.0 cM.
D.4 Structure within England_CA_EBA
Largest absolute residuals against the common-effect
estimate of 8.55 per cent:
|
ID |
Site |
Estimate |
z |
|
I2462 |
East Kent Access |
35.3 ± 3.8 % |
+7.04 |
|
I6777 |
Wilsford G.54 |
0.0 ± 4.0 % |
−2.14 |
|
I2597 |
Amesbury Down |
0.9 ± 4.0 % |
−1.91 |
|
I5373 |
Carsington Pasture Cave |
0.0 ± 4.5 % |
−1.90 |
|
|
Mean |
Q / df |
p |
I² |
τ |
|
With I2462 |
8.55 % |
81.0 / 27 |
2.6 × 10⁻⁷ |
66.7 % |
6.36 pp |
|
Without I2462 |
7.15 % |
28.9 / 26 |
0.32 |
10.0 % |
1.51 pp |
Residual against calibrated date, I2462 removed: r =
+0.501, p = 0.009. By genetic sex, I2462 removed: male n = 16,
mean 6.58 %; female n = 10, mean 6.35 % (Mann–Whitney p = 0.83).
Sites with more than one dated individual, I2462 removed: Amesbury Down n
= 5, mean 2.10 %; Baston and Langtoft n = 2, mean 7.60 %; East Kent
Access n = 3, mean 10.43 %.
Appendix F — Specification and power of the IBD
test
F.1 The comparison is absent by construction
Checked against Supplementary Table 14's own
population-label columns rather than by external mapping: of 5,846 reported
pairs, 1,157 carry population labels for both members, and no pair has
English individuals on both sides. Every one of the 78 pairs involving an
English individual has a Lower Rhine–Meuse partner, which is the table's stated
scope.
Individuals available under the table's own labels: 19
England_N (including England_N_Megalithic and kinship-annotated variants), 11
England_C_EBA. Full group sizes from Supplementary Table 3 are 27 and 39.
F.2 The method has resolution at the required
depth
Date gaps among the 1,157 pairs with dates for both members:
|
Gap (years) |
Pairs |
Median longest segment |
|
0–200 |
483 |
16.1 cM |
|
200–400 |
224 |
14.3 cM |
|
400–800 |
274 |
14.1 cM |
|
800–1200 |
124 |
13.8 cM |
|
1200–1600 |
33 |
13.6 cM |
|
1600–1961 |
19 |
13.1 cM |
Spearman ρ between
gap and longest segment = −0.341 (p = 5.9 × 10⁻³³). The required
comparison spans roughly 1,600 years — England_N has a median date near 5500
BP, England_CA_EBA near 3800 BP. Detection at that depth occurs in the existing
data: nineteen pairs exceed a 1,600-year gap, and the longest recorded is 1,961
years. Segment lengths compress toward the 12 cM reporting threshold but do not
vanish. The clearest single instance is KD070.SG (England_Northumberland_EBA,
4306 BP) sharing 16.0 cM with an MN_Wartberg individual at 5164 BP across 858
years.
F.3 Benchmark detection rates
|
Comparison |
Individuals |
Possible pairs |
Detected |
Rate |
Median longest |
|
England_BellBeaker × LNB_Bell_Beaker |
16 × 11 |
176 |
30 |
17.0 % |
15.6 cM |
|
England_C_EBA × LNB_Bell_Beaker |
11 × 11 |
121 |
18 |
14.9 % |
14.2 cM |
|
England_N × MN_Tiel |
19 × 2 |
38 |
17 |
44.7 % |
14.5 cM |
The first row is the appropriate benchmark:
near-contemporary individuals with essentially all ancestry from the partner
population. The third row's high rate reflects only two MN_Tiel individuals, at
least one of whom shares with many England_N individuals, and should not be
used for scaling.
F.4 Expected yield
Modelling detections as Poisson with rate proportional to
the ancestry fraction f = 0.076 times the benchmark rate of 0.170 (30
detections in 176 possible pairs), times a decay factor for the 1,600-year
separation:
|
Decay assumption |
Coverage |
Pairs |
Expected detections |
P(≥
1) |
P(≥
3) |
|
None |
current (19 × 11) |
209 |
2.7 |
0.93 |
0.51 |
|
None |
full groups (27 × 39) |
1,053 |
13.6 |
1.00 |
1.00 |
|
50 % |
current |
209 |
1.4 |
0.74 |
0.16 |
|
50 % |
full groups |
1,053 |
6.8 |
1.00 |
0.97 |
|
75 % |
current |
209 |
0.7 |
0.49 |
0.03 |
|
75 % |
full groups |
1,053 |
3.4 |
0.97 |
0.66 |
Under the continental-source null the England_N-specific
excess is zero, so any detection above the shared-deep-ancestry background is
evidence for a British source. At full group coverage the test discriminates
decisively across the whole range of decay assumptions; at the coverage used in
the published table it does not.
F.5 Design requirements
Prune for relatedness. At least three of the 19
England_N individuals carry explicit kinship annotations, and those annotations
reference a further five identifiers that also appear in the table — I21393,
I21389, I30334, I30304, I21395, I30332, I30302. A pedigree therefore spans
something like eight of the nineteen. The Poisson calculation above treats
pairs as independent and consequently overstates power; the effective
independent sample is materially smaller and should be reduced to one
representative per pedigree before testing.
Run a matched-baseline arm. Time decay is the largest
nuisance parameter and cannot be estimated well from 19 pairs. Running the
England_CA_EBA arm simultaneously against date-matched continental candidates
makes decay common to both arms and cancels it in the ratio. Candidates already
present in the table, with median dates:
|
Population |
n |
Median date (BP) |
|
MN_Wartberg |
26 |
5140 |
|
France_MontAime |
5 |
5162 |
|
Ireland_MN |
8 |
5410 |
|
MN_Hazendonk |
2 |
5470 |
|
MN_Tiel |
2 |
5650 |
|
Czechia_C_Baalberge |
7 |
5850 |
|
France_N |
18 |
6550 |
Ireland_MN and MN_Wartberg are the best-matched to
England_N's median of about 5500 BP. France_N at 6550 BP is a poor baseline and
Denmark_SouthScandinavia_LN at 4124 BP is too late.
Relax the coverage threshold. Only 12 of the 39
England_CA_EBA individuals and 6 of the 27 England_N individuals appear in
Supplementary Table 14 under the external group mapping; the table's own labels
raise this to 11 and 19. Either way, roughly two-thirds of the relevant
individuals are excluded by whatever coverage criterion was applied. The power
calculation above shows this is the difference between a decisive test and an
inconclusive one.
Caveats on the calculation. The linear scaling of
detection rate in the ancestry fraction is an approximation that ignores the
distribution of segment lengths contributed by a minority ancestry component,
which will be shifted downward relative to a full-ancestry comparison and so
will lose disproportionately at a fixed 12 cM threshold. The benchmark rate
itself rests on 30 detections. And the Poisson independence assumption is
violated by relatedness in both arms. Taken together these push the true power
below the tabulated figures, which is why the full-group coverage matters
rather than being a refinement.
All figures above are produced by the accompanying script reproduce.py,
available at https://claude.ai/public/artifacts/97845468-1666-45b9-877c-c9ac70cf7fc0
which takes the three published supplementary workbooks as input and prints
every number in the order it appears here. The random seed for the censoring
simulation is fixed at 20260802; simulation figures are stable to the second
decimal place across seeds.
Required files:
-
Booth et al. 2021 Supplementary Table 1 (Cambridge Core,
supplementary material to doi:10.1017/S0959774321000019)
-
Olalde et al. 2018 Supplementary Tables (Nature 555,
doi:10.1038/nature25738)
-
Olalde et al. 2026 Supplementary Tables (Nature,
doi:10.1038/s41586-026-10111-8)
Dependencies: Python 3, numpy, pandas, scipy, openpyxl.
Sources
-
Booth, T.J., Brück, J., Brace, S. & Barnes,
I. 2021. Tales from the supplementary information: ancestry change in
Chalcolithic–Early Bronze Age Britain was gradual with varied kinship
organization. Cambridge Archaeological Journal 31(3): 379–400.
doi:10.1017/S0959774321000019.
-
Olalde, I. et al. 2018. The Beaker
phenomenon and the genomic transformation of northwest Europe. Nature
555: 190–96. doi:10.1038/nature25738.
-
Olalde, I. et al. 2026. Lower Rhine–Meuse
forager ancestry and the Bell Beaker expansion. Nature.
doi:10.1038/s41586-026-10111-8.
Methodological background on qpAdm behaviour, model
rejection and rotating source analysis is not cited above because none of the
calculations here depend on it, but readers evaluating the degrees-of-freedom
convention in Appendix A and the interpretation of non-rejection in §7 should
consult Harney et al. 2021 (Genetics 217) and Maier et al.
2023 (eLife 12).