Further on beyond cognacy
Model adequacy and meaningful branch lengths in alignment-based phylogenetics
A walkthrough of the QC-LESTE talk (Paris, 11 June 2026). The slides and code are available separately. Script names below point into that repository.
Background: align, don’t classify
Phylogenetic linguistics infers language trees from word lists. The standard pipeline first needs a cognate-classification step — manual (labour-intensive, geographically skewed) or automatic (usable, but imperfect). Cognate classification is the bottleneck.
The alternative is to align all reflexes of a concept, regardless of cognacy (a profile-HMM plus T-Coffee), and treat each alignment column as a phylogenetic character — no cognate judgments at all.

On the Lexibench benchmark, MSA characters win on all three metrics against both cognate classes and PMI (GQD to Glottolog 0.054, AIC for Grambank 104,787, Pythia difficulty 0.21). Scaled to the whole Lexibank sample, this yields the Lexibank World Tree of 3,397 languages — GQD 0.038 to Glottolog.

This companion takes that tree and the underlying alignments as given (Jäger 2026, doi.org/10.57754/FDAT.ccfpp-z0113) and answers the two questions the paper left open.
A yardstick: Häuser, Jäger & Stamatakis (2026)
“Current limitations of cognate-based phylogenetic inference” (Open Research Europe 5:258) builds parameter-rich multistate cognate models and finds they can’t be applied yet — the cognate datasets are too small. It names four pathologies of cognate matrices, a checklist for any new representation:
| ① Data scarcity | ② Overparameterized | ③ Poor generalization | ④ ML doesn’t transfer |
|---|---|---|---|
| entropy ≤ ~2,500 vs DNA > 10,000 | GTR collapses by κ = 5 (AIC) | CV error ≥ 7× on κ-subsets | Pythia MAE 0.17 vs 0.08 |
Do alignment-based matrices escape them? We re-run Häuser’s diagnostics on the MSA + Dolgopolsky sound-class encoding, per Lexibank dataset (one rich MULTI12 matrix per dataset, no κ-subsets; 56 datasets, 10–115 taxa).
① Information content recovers
The Dolgo-MSA matrices carry 0.52–1.17 bits/column, versus binary cognate’s ~0.1–0.4, overlapping the low end of DNA (0.8–1.8). The binary-vs-κ-subset dilemma is gone: one matrix replaces both.

code/adequacy/analyse_charmatrices.jl, figure figures/make_entropy_fig.py.②③ No overparameterization
AIC, per dataset: GTR is preferred over MK in all 56 datasets (ΔAIC −685 to −69,014) — easily enough data for GTR’s 76 parameters. Held-out cross-validation (Häuser’s §3.7 design: GTR/MK fit on 60% of columns, frozen, scored on the held-out 40%): GTR predicts unseen columns better in 56/56 datasets, generalization gap e ≈ 0.016, versus Häuser’s 0.21–0.40 on overparameterized κ-subsets.

code/adequacy/crossval_dolgo.jl, code/adequacy/cv_intermediate.py, code/adequacy/finalize_perdataset.py; figure figures/make_cv_fig.py.The pathologies were representation artifacts, not bad luck. (Pathology ④, Pythia predictor transfer, is not fixed by representation — but the difficulty score itself drops 0.59 → 0.21.)
Branch lengths: the big caveat
The world-tree paper’s own caveat is that branch lengths are not meaningful: ML/CTMC lengths from the binary MSA are terminal-inflated, too much length piled on the tips. The diagnostic is the γ-shape of the branch-length distribution (γα): real and simulated trees sit at γα ≈ 0.75–1.2; the CTMC tree is at γα = 2.31, far outside.

code/eval/bl_geom.py, code/eval/ks_terminal_internal.py; figure figures/make_branch_violin.py.
code/eval/bouckaert_gamma.py, code/eval/bd_sampling_calib.py; figure figures/make_branch_dist.py.This is a familiar problem in molecular phylogenetics: branch lengths estimate expected substitutions, not time; saturation compresses deep branches; the node-density effect and incomplete sampling inflate tips (Stadler 2009).

figures/make_brlen_schematics.py.Biology tames this with relaxed clocks, calibrated dating, tree priors, and continuous-trait models — Brownian motion and Ornstein–Uhlenbeck (OU). We repurpose OU to re-estimate linguistic branch lengths.

figures/make_ou_schematic.py.From sound classes to continuous traits
Logistic-PCA of the global Dolgo MSA (3,393 languages, presence-controlled) gives K = 200 continuous traits per language. Presence-controlled means the factors are lexical, not attestation artifacts (r(PC, attestation) ≈ 0). We then re-estimate world-tree branch lengths under a multivariate OU process on the fixed topology — a MAP point estimate with a per-trait reversion rate (hierarchical α), a terminal nugget (per-tip noise), and no birth-process prior.
Why K = 200? The phylogenetic signal is ~200-dimensional. Pagel’s λ of each sound-class PC over the world tree (Grafen branch lengths) stays ≈ 0.98–0.99, uniform through PC ≈ 190, reaching a noise floor only by ≈ 225.

code/eval/logistic_pca.py, code/eval/pagel_lambda_fast.py, code/eval/pagel_lambda_pcs.R; figure figures/make_pagel_lambda_fig.py.The leading components are macro-areal contrasts — each varimax factor a bipolar areal gradient, deepest first.

code/eval/dolgo_wt_varimax_leading.py, code/eval/dolgo_wt_varimax_maps.py.
OU restores a sensible geometry — and it’s data-driven
OU branch lengths de-inflate the terminal pathology to γα = 1.16, inside the sensible range (0.75–1.2); CTMC’s 2.31 is far outside. Two controls show it’s the signal, not the method: ablating the terminal nugget barely moves it (1.16 → 1.24); shuffling traits across tips pushes γα to ≈ 1.82, back outside the band (terminal share 72%, as terminal-heavy as CTMC).

code/eval/fit_ou_dating.py, code/eval/bl_geom.py; figure figures/make_ou_gamma_bar.py.The γα fix is about the shape of the distribution. The terminal share — how much length sits on the tips — is a second diagnostic. OU pulls it from CTMC’s 67% down to 59% (treeness 0.33 → 0.41), part-way toward a real calibrated dated tree (Bouckaert: 47%). Closing the rest is a clock/calibration job — the same division of labour as in biology.

figures/make_termshare_3panel.py.The result is an ultrametric, alignment-based world tree with a sensible, data-driven geometry on a relative timescale — ready to calibrate to absolute dates.

code/eval/build_tip_meta.py, code/eval/plot_worldtree_ggtree.R.Download world tree with OU branch lengths (Newick, 96 kB)
Side by side with the independent Bouckaert et al. (2022/2026) MCC: the major macro-areal blocks recur in both (Austronesian, Sino-Tibetan, Indo-European, Atlantic-Congo, Pama-Nyungan).

Further beyond cognacy
Alignment is a viable alternative to manual cognate coding: competitive on topology, and a well-behaved phylogenetic model on two axes — model adequacy (it avoids the cognate κ-subset pathologies) and branch lengths (continuous OU embeddings turn meaningless ML lengths into a sensible, interpretable geometry with areal structure). A natural next step is combined-evidence inference (alignments + cognate classification + Grambank), and a horizontal/contact layer from the OU off-tree residual.
Appendix: robustness

code/eval/plot_halflife_spectrum.py.
code/eval/ou_prior_sensitivity.py.Code and slides written with assistance by Claude Code.