Further on beyond cognacy

Model adequacy and meaningful branch lengths in alignment-based phylogenetics

Author

Gerhard Jäger

Published

June 11, 2026

A walkthrough of the QC-LESTE talk (Paris, 11 June 2026). The slides and code are available separately. Script names below point into that repository.

Background: align, don’t classify

Phylogenetic linguistics infers language trees from word lists. The standard pipeline first needs a cognate-classification step — manual (labour-intensive, geographically skewed) or automatic (usable, but imperfect). Cognate classification is the bottleneck.

The alternative is to align all reflexes of a concept, regardless of cognacy (a profile-HMM plus T-Coffee), and treat each alignment column as a phylogenetic character — no cognate judgments at all.

One concept aligned across languages; column 4 binarized (Akavarapu & Bhattacharya 2024). From Jäger 2025, SIGTYP.

On the Lexibench benchmark, MSA characters win on all three metrics against both cognate classes and PMI (GQD to Glottolog 0.054, AIC for Grambank 104,787, Pythia difficulty 0.21). Scaled to the whole Lexibank sample, this yields the Lexibank World Tree of 3,397 languages — GQD 0.038 to Glottolog.

The Lexibank World Tree: 3,397 languages, maximum-likelihood (RAxML-NG), macroarea-coloured.

This companion takes that tree and the underlying alignments as given (Jäger 2026, doi.org/10.57754/FDAT.ccfpp-z0113) and answers the two questions the paper left open.

A yardstick: Häuser, Jäger & Stamatakis (2026)

“Current limitations of cognate-based phylogenetic inference” (Open Research Europe 5:258) builds parameter-rich multistate cognate models and finds they can’t be applied yet — the cognate datasets are too small. It names four pathologies of cognate matrices, a checklist for any new representation:

① Data scarcity ② Overparameterized ③ Poor generalization ④ ML doesn’t transfer
entropy ≤ ~2,500 vs DNA > 10,000 GTR collapses by κ = 5 (AIC) CV error ≥ 7× on κ-subsets Pythia MAE 0.17 vs 0.08

Do alignment-based matrices escape them? We re-run Häuser’s diagnostics on the MSA + Dolgopolsky sound-class encoding, per Lexibank dataset (one rich MULTI12 matrix per dataset, no κ-subsets; 56 datasets, 10–115 taxa).

① Information content recovers

The Dolgo-MSA matrices carry 0.52–1.17 bits/column, versus binary cognate’s ~0.1–0.4, overlapping the low end of DNA (0.8–1.8). The binary-vs-κ-subset dilemma is gone: one matrix replaces both.

Per-dataset column entropy (56 datasets). Dolgo-MSA overlaps the molecular band; the binary-cognate band is empty. Scripts: code/adequacy/analyse_charmatrices.jl, figure figures/make_entropy_fig.py.

②③ No overparameterization

AIC, per dataset: GTR is preferred over MK in all 56 datasets (ΔAIC −685 to −69,014) — easily enough data for GTR’s 76 parameters. Held-out cross-validation (Häuser’s §3.7 design: GTR/MK fit on 60% of columns, frozen, scored on the held-out 40%): GTR predicts unseen columns better in 56/56 datasets, generalization gap e ≈ 0.016, versus Häuser’s 0.21–0.40 on overparameterized κ-subsets.

GTR behaves like molecular data, not cognate: in the well-behaved zone, the κ-subset overparameterization band empty. Scripts: code/adequacy/crossval_dolgo.jl, code/adequacy/cv_intermediate.py, code/adequacy/finalize_perdataset.py; figure figures/make_cv_fig.py.

The pathologies were representation artifacts, not bad luck. (Pathology ④, Pythia predictor transfer, is not fixed by representation — but the difficulty score itself drops 0.59 → 0.21.)

Branch lengths: the big caveat

The world-tree paper’s own caveat is that branch lengths are not meaningful: ML/CTMC lengths from the binary MSA are terminal-inflated, too much length piled on the tips. The diagnostic is the γ-shape of the branch-length distribution (γα): real and simulated trees sit at γα ≈ 0.75–1.2; the CTMC tree is at γα = 2.31, far outside.

Under ML/CTMC, terminal branches dwarf internal ones: median terminal ≈ 2.3× internal, 67% of tree length on the tips. Scripts: code/eval/bl_geom.py, code/eval/ks_terminal_internal.py; figure figures/make_branch_violin.py.

A real dated tree (Bouckaert et al. 2022/2026) and birth-death simulations give γα ≈ 0.75–1.2 (lengths roughly decay); the CTMC distribution is a hump (γα = 2.31). Scripts: code/eval/bouckaert_gamma.py, code/eval/bd_sampling_calib.py; figure figures/make_branch_dist.py.

This is a familiar problem in molecular phylogenetics: branch lengths estimate expected substitutions, not time; saturation compresses deep branches; the node-density effect and incomplete sampling inflate tips (Stadler 2009).

Three sources of branch-length distortion: saturation, node-density, incomplete sampling. Figure figures/make_brlen_schematics.py.

Biology tames this with relaxed clocks, calibrated dating, tree priors, and continuous-trait models — Brownian motion and Ornstein–Uhlenbeck (OU). We repurpose OU to re-estimate linguistic branch lengths.

OU pulls a trait toward an optimum at rate α; the phylogenetic half-life is t½ = ln 2 / α. Figure figures/make_ou_schematic.py.

From sound classes to continuous traits

Logistic-PCA of the global Dolgo MSA (3,393 languages, presence-controlled) gives K = 200 continuous traits per language. Presence-controlled means the factors are lexical, not attestation artifacts (r(PC, attestation) ≈ 0). We then re-estimate world-tree branch lengths under a multivariate OU process on the fixed topology — a MAP point estimate with a per-trait reversion rate (hierarchical α), a terminal nugget (per-tip noise), and no birth-process prior.

Why K = 200? The phylogenetic signal is ~200-dimensional. Pagel’s λ of each sound-class PC over the world tree (Grafen branch lengths) stays ≈ 0.98–0.99, uniform through PC ≈ 190, reaching a noise floor only by ≈ 225.

Per-PC Pagel’s λ: uniform signal through ~190 components, collapsing by ~225 — a ~200-dimensional phylogenetic signal. Scripts: code/eval/logistic_pca.py, code/eval/pagel_lambda_fast.py, code/eval/pagel_lambda_pcs.R; figure figures/make_pagel_lambda_fig.py.

The leading components are macro-areal contrasts — each varimax factor a bipolar areal gradient, deepest first.

Leading varimax factors (1/2): Sinosphere vs Altaic ≈ 24 ka, then Australia, MSEA, W-Eurasia contrasts at ≈ 19–22 ka. Half-lives from Bayesian OU over the Bouckaert posterior. Scripts: code/eval/dolgo_wt_varimax_leading.py, code/eval/dolgo_wt_varimax_maps.py.

Leading varimax factors (2/2): shallower factors track more recent structure; Austronesian ≈ 12 ka.

OU restores a sensible geometry — and it’s data-driven

OU branch lengths de-inflate the terminal pathology to γα = 1.16, inside the sensible range (0.75–1.2); CTMC’s 2.31 is far outside. Two controls show it’s the signal, not the method: ablating the terminal nugget barely moves it (1.16 → 1.24); shuffling traits across tips pushes γα to ≈ 1.82, back outside the band (terminal share 72%, as terminal-heavy as CTMC).

γα for CTMC (2.31), tip-shuffle null (≈1.82), and OU (1.16), against the sensible 0.75–1.2 band. Scripts: code/eval/fit_ou_dating.py, code/eval/bl_geom.py; figure figures/make_ou_gamma_bar.py.

The γα fix is about the shape of the distribution. The terminal share — how much length sits on the tips — is a second diagnostic. OU pulls it from CTMC’s 67% down to 59% (treeness 0.33 → 0.41), part-way toward a real calibrated dated tree (Bouckaert: 47%). Closing the rest is a clock/calibration job — the same division of labour as in biology.

Terminal vs internal branch lengths, mean-normalised per tree: CTMC (0.67) → OU (0.59) → Bouckaert calibrated (0.47). Figure figures/make_termshare_3panel.py.

The result is an ultrametric, alignment-based world tree with a sensible, data-driven geometry on a relative timescale — ready to calibrate to absolute dates.

Ultrametric Dolgo–OU world tree, macroarea-coloured (relative time, root = 1). Scripts: code/eval/build_tip_meta.py, code/eval/plot_worldtree_ggtree.R.

Download world tree with OU branch lengths (Newick, 96 kB)

Side by side with the independent Bouckaert et al. (2022/2026) MCC: the major macro-areal blocks recur in both (Austronesian, Sino-Tibetan, Indo-European, Atlantic-Congo, Pama-Nyungan).

Left: Dolgo–OU world tree (3,393 tips, ultrametric, relative time). Right: Bouckaert MCC (6,636 tips, dated).

Further beyond cognacy

Alignment is a viable alternative to manual cognate coding: competitive on topology, and a well-behaved phylogenetic model on two axes — model adequacy (it avoids the cognate κ-subset pathologies) and branch lengths (continuous OU embeddings turn meaningless ML lengths into a sensible, interpretable geometry with areal structure). A natural next step is combined-evidence inference (alignments + cognate classification + Grambank), and a horizontal/contact layer from the OU off-tree residual.

Appendix: robustness

PC half-life spectrum, Bayesian MCMC over the Bouckaert posterior (independent, no circularity): depth declines from ≈20 ka (leading) to ≈2 ka (deepest), Spearman r = −0.91. Script: code/eval/plot_halflife_spectrum.py.

Prior-sensitivity check: every OU-parameter prior widened 5×, node ages barely move (cophenetic r = 0.998, median |Δ| = 0.001 of root depth). The geometry is set by the data, not the priors. Script: code/eval/ou_prior_sensitivity.py.

Code and slides written with assistance by Claude Code.