Naxos Summer School 2026
  • Home
  • Schedule
  • Reading List

Reading List

Phylogenetic Methods in Historical Linguistics · Naxos Summer School 2026

This course introduces students to phylogenetic computational methods and their application in historical linguistics, combining theoretical foundations, case studies, and hands-on applications across five lectures.

About this list Compiled and updated for the 11th Naxos Summer School on Diachronic Linguistics. Items marked LANDMARK are essential reading; others provide depth or technical background.
§ 1

Overview & Foundations

Start here
2016
Jäger, Gerhard & Johann-Mattis List
Statistical and computational elaborations of the classical comparative method
Unpublished manuscript, Tübingen & Paris. Submitted for The Oxford Handbook of Diachronic and Historical Linguistics, but not published there
A comprehensive overview of the field: sequence comparison, cognate detection, and phylogenetic reconstruction. A natural starting point for this course.
Landmark Overview Jäger
↗ PDF
2005
Atkinson, Quentin D. & Russell D. Gray
Curious Parallels and Curious Connections — Phylogenetic Thinking in Biology and Historical Linguistics
Systematic Biology 54(4), 513–526
An accessible and well-argued case for why phylogenetic thinking translates from evolutionary biology to historical linguistics. Essential for understanding the conceptual bridge.
Landmark Overview
2014
List, Johann-Mattis
Sequence Comparison in Historical Linguistics
Dissertations in Language and Cognition, Vol. 1. Düsseldorf University Press
Thorough treatment of sequence comparison methods and their linguistic applications. Especially strong on alignment algorithms and cognate detection.
Overview Methods
2011
Nunn, Charles L.
The Comparative Approach in Evolutionary Anthropology and Biology
The University of Chicago Press
Broadens perspective to the comparative method in anthropology and biology — useful for understanding the intellectual context of the phylogenetic approach.
Overview
2018
Forkel, Robert, Johann-Mattis List, Simon J. Greenhill, Christoph Rzymski, et al.
Cross-Linguistic Data Formats, advancing data sharing and re-use in comparative linguistics
Scientific Data 5, 180205
Defines CLDF, now the de-facto standard for sharing cross-linguistic datasets. It underpins Lexibank, Grambank, and the data used in both labs of this course.
Overview Methods Data
↗ DOI
2022
List, Johann-Mattis, Robert Forkel, Simon J. Greenhill, Christoph Rzymski, Johannes Englisch & Russell D. Gray
Lexibank, a public repository of standardized wordlists with computed phonological and lexical features
Scientific Data 9, 316
A unified, CLDF-based collection of hundreds of wordlists with automatically computed features. The practical backbone of large-scale lexical phylogenetics today.
Overview Methods Data
↗ DOI
2008
Greenhill, Simon J., Robert Blust & Russell D. Gray
The Austronesian Basic Vocabulary Database: From Bioinformatics to Lexomics
Evolutionary Bioinformatics 4, 271–283
The expert-coded cognate database behind most Austronesian phylogenetics: 210-meaning wordlists for around 2,000 varieties, with some 19,000 cognate sets. Together with IE-CoR it is one of the two great hand-coded resources, and it shows what that level of annotation costs.
Overview Methods Data
↗ DOI
2020
Dellert, Johannes, Thora Daneyko, Alla Münch, Alina Ladygina, Armin Buch, Natalie Clarius, Ilja Grigorjew, Mohamed Balabel, Hizniye Isabella Boga, Zalina Baysarova, Roland Mühlenbernd, Johannes Wahle & Gerhard Jäger
NorthEuraLex: a wide-coverage lexical database of Northern Eurasia
Language Resources and Evaluation 54(1), 273–301
1,016 concepts across 107 languages from 21 families, all in a uniform IPA transcription generated from the orthographies — which is what makes comparison across families possible. Deliberately supplies no cognacy annotation: it was built as a benchmark for inferring it. Discussed in Lecture 2.
Overview Methods Data
↗ DOI
2001
Kessler, Brett
The Significance of Word Lists: Statistical Tests for Investigating Historical Connections Between Languages
CSLI Publications, Stanford
The source of the permutation logic that LexStat rests on: to decide whether a sound correspondence is real, compare how often it is attested against how often shuffled wordlists produce it by chance. Also the origin of squaring the frequencies, on the argument that recurrence counts as evidence faster than linearly.
Overview Methods
2023
Skirgård, Hedvig, Hannah J. Haynie, Damián E. Blasi, Harald Hammarström, et al.
Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss
Science Advances 9(16), eadg6175
The first large-scale grammatical database, covering ~2,400 languages and 195 structural features. It opens up phylogenetic and areal study of grammar, and quantifies the looming loss of structural diversity.
Landmark Overview Data
↗ DOI
§ 2

Background in Bioinformatics

Technical depth
2014
Chen, Ming-Hui, Lynn Kuo & Paul O. Lewis
Bayesian Phylogenetics: Methods, Algorithms and Applications
CRC Press, Abingdon
The standard reference for Bayesian phylogenetic methods. Mathematically rigorous; recommended for those wanting a deep understanding of the inference machinery behind BEAST and MrBayes.
Bioinformatics Methods
1998
Durbin, Richard, Sean R. Eddy, Anders Krogh & Graeme Mitchison
Biological Sequence Analysis
Cambridge University Press
The classic textbook on probabilistic models of sequences. Covers HMMs, alignment algorithms, and phylogenetics. A foundational reference for the computational side of the course.
Landmark Bioinformatics
2005
Ewens, Warren & Gregory Grant
Statistical Methods in Bioinformatics: An Introduction
Springer
A gentler statistical introduction. Useful for students who need to build up the probabilistic background before tackling the methods literature.
Bioinformatics
§ 3

Research Highlights & Case Studies

Exemplary applications
2023
Heggarty, Paul, Cormac Anderson, Matthew Scarborough, et al.
Language trees with sampled ancestors support a hybrid model for the origin of Indo-European languages
Science 381(6656), eabg0818
The current reference point for Indo-European phylogenetics. A 161-language analysis with a sampled-ancestor (fossilised birth-death) model; its "hybrid" result places a deep Indo-Anatolian split before a later steppe expansion.
Landmark Case Study
↗ DOI
2025
Kassian, Alexei S. & George Starostin
Do 'language trees with sampled ancestors' really support a 'hybrid model' for the origin of Indo-European? Thoughts on the most recent attempt at yet another IE phylogeny
Humanities and Social Sciences Communications 12, 682
A critique of Heggarty et al. (2023) discussed in Lecture 5. Argues that the cognacy coding does not consistently separate true cognacy from areal diffusion, that the deep topology is barely resolved ("a meaningless rake"), and that 170 concepts are too few for 161 taxa; defends the linguistic-palaeontology argument from reconstructed *kʷekʷlo- 'wheel' for a much later break-up.
Case Study
↗ DOI
2019
Blasi, Damián E., Steven Moran, Scott R. Moisik, Paul Widmer, Dan Dediu & Balthasar Bickel
Human sound systems are shaped by post-Neolithic changes in bite configuration
Science 363(6432), eaav3218
A second target paper for the course. Argues that labiodentals (f, v) spread after the Neolithic as softer diets preserved the overbite and overjet, using a phylogenetic comparative method to control for descent and contact; the rebuttal by Tarasov & Uyeda (2020) contests exactly that phylogenetic model.
Landmark Case Study
↗ DOI
2025
Lazaridis, Iosif, Nick Patterson, David Reich, et al.
The genetic origin of the Indo-Europeans
Nature (2025)
Ancient-DNA evidence identifying a Caucasus–Lower Volga population (~6,500 BP) as the common source of both Anatolian and steppe (Yamnaya) lineages. Brings genetics directly to bear on the homeland debate, broadly consistent with the Heggarty et al. hybrid model.
Case Study
↗ DOI
2019
Sagart, Laurent, Guillaume Jacques, Yunfan Lai, Robin J. Ryder, et al.
Dated language phylogenies shed light on the ancestry of Sino-Tibetan
Proceedings of the National Academy of Sciences 116(21), 10317–10322
A Bayesian phylogeny of 50 Sino-Tibetan languages, dating the family to ~7,200 BP in northern China. The Sino-Tibetan counterpart to the Indo-European homeland debate.
Case Study
↗ DOI
2019
Zhang, Menghan, Shi Yan, Wuyun Pan & Li Jin
Phylogenetic evidence for Sino-Tibetan origin in northern China in the Late Neolithic
Nature 569, 112–115
An independent Bayesian analysis reaching a younger root (~4,000–6,000 BP). Read alongside Sagart et al. (2019): the same northern homeland, notably different dates, a useful object lesson in dating sensitivity.
Case Study
↗ DOI
2018
Bouckaert, Remco R., Claire Bowern & Quentin D. Atkinson
The origin and expansion of Pama–Nyungan languages across Australia
Nature Ecology & Evolution 2, 741–749
A continental-scale phylogeography of Australia's largest family, with a mid-Holocene origin (~6,000 BP) in the Gulf Plains. A model application of Bayesian phylogeography outside Eurasia.
Case Study
↗ DOI
2018
Kolipakam, Vishnupriya, Fiona M. Jordan, Michael Dunn, Simon J. Greenhill, et al.
A Bayesian phylogenetic study of the Dravidian language family
Royal Society Open Science 5(3), 171504
Cognate-coded Bayesian analysis dating Dravidian to ~4,500 BP and recovering its major branches. Broadens the set of carefully studied families beyond the usual European and Pacific cases.
Case Study
↗ DOI
2021
Meloni, Carlo, Shauli Ravfogel & Yoav Goldberg
Ab Antiquo: Neural Proto-language Reconstruction
Proc. NAACL-HLT 2021, 4460–4473
Casts proto-form reconstruction as neural sequence-to-sequence learning, approaching expert quality on Romance. Representative of the deep-learning turn in computational historical linguistics.
Case Study Methods
↗ ACL
2003
Gray, Russell D. & Quentin D. Atkinson
Language-tree divergence times support the Anatolian theory of Indo-European origin
Nature 426(6965), 435–439
The paper that put computational historical linguistics on the map. Applies Bayesian phylogenetics to Indo-European and generates a dateable tree. Hugely influential — and contested.
Landmark Case Study
2012
Bouckaert, Remco, Philippe Lemey, Michael Dunn, Simon J. Greenhill, et al.
Mapping the Origins and Expansion of the Indo-European Language Family
Science 337(6097), 957–960
A landmark study combining Bayesian phylogenetics with geographic diffusion modeling to trace Indo-European origins. Essential reading on the Anatolian vs. steppe debate.
Landmark Case Study
2015
Chang, Will, Chundra Cathcart, David Hall & Andrew Garrett
Ancestry-constrained phylogenetic analysis supports the Indo-European steppe hypothesis
Language 91(1), 194–244
A methodological improvement on Bouckaert et al. (2012), imposing linguistically motivated clade constraints. Arrives at steppe rather than Anatolian origins. An important counterpoint.
Landmark Case Study Methods
2013
Pagel, Mark, Quentin D. Atkinson, Andreea S. Calude & Andrew Meade
Ultraconserved words point to deep language ancestry across Eurasia
Proceedings of the National Academy of Sciences 110(21), 8471–8476
Uses rates of lexical evolution inferred from phylogenies to identify ultra-stable words potentially linking major Eurasian families (the "Eurasiatic" hypothesis). Provocative and methodologically interesting.
Case Study
2007
Pagel, Mark, Quentin D. Atkinson & Andrew Meade
Frequency of word-use predicts rates of lexical evolution throughout Indo-European history
Nature 449(7163), 717–720
Shows that word frequency is a strong predictor of evolutionary stability — a robust and elegant result with practical implications for data selection in phylogenetic studies.
Case Study
2010
Wichmann, Søren, André Müller & Viveka Velupillai
Homelands of the world's language families: A quantitative approach
Diachronica 27(2), 247–276
Applies quantitative distance methods at a global scale to infer where major language families originated. A good example of applying computational methods beyond Indo-European.
Case Study
§ 4

Selected Work by G. Jäger

Instructor's contributions
2015
Jäger, Gerhard
Support for linguistic macrofamilies from weighted sequence alignment
Proceedings of the National Academy of Sciences 112(41), 12752–12757
Uses empirically weighted sequence alignment to provide quantitative support for proposed macro-families. Introduces the PMI-based distance measure used in later work.
Jäger Methods Case Study
2013
Jäger, Gerhard
Phylogenetic inference from word lists using weighted alignment with empirically determined weights
Language Dynamics and Change 3(2), 245–291
Introduces the PMI-weighted alignment approach for computing phonetic distances and inferring phylogenies from raw word lists without pre-coded cognacy judgments.
Jäger Methods
2018
Jäger, Gerhard
Global-scale phylogenetic linguistic inference from lexical resources
Scientific Data 5, 180189
Builds a distance-based tree of roughly 6,000 languages from ASJP using PMI-weighted alignment, recovering most established families automatically. The full-scale realisation of the world-tree programme begun in Jäger & Wichmann (2016).
Jäger Case Study Methods
↗ DOI
2026
Jäger, Gerhard
A Phylogenetic Tree of 3,397 World Languages via Multiple Sequence Alignment
FDAT, University of Tübingen (data paper). Earlier preprint: Beyond cognacy, arXiv:2507.03005
Compares expert-cognate inference with two fully automated alternatives, including multiple-sequence alignment from a pair-hidden Markov model, evaluated against Glottolog and Grambank. MSA-based trees track expert classifications most closely (96.4% quartet agreement), pointing toward global phylogenies without a cognate-annotation bottleneck.
Jäger Methods
↗ DOI
2019
Jäger, Gerhard
Computational historical linguistics
Theoretical Linguistics 45(3-4), 151-182
Target-article survey of the field: what computational methods have contributed to historical linguistics, and where the open problems are. Recommended after Lecture 1.
Jäger Overview
↗ DOI
2025
Jäger, Gerhard
Computational Typology
Manuscript, University of Tübingen. arXiv:2504.15642
Typology classifies languages by structural features rather than by descent. A short illustration (19 pages) of what computational statistical modelling contributes: analysing typological data at scale, and testing hypotheses about language structure and evolution. The overview companion to the two word-order papers below.
Jäger Overview
↗ arXiv
2021
Jäger, Gerhard & Johannes Wahle
Phylogenetic typology
Frontiers in Psychology 12, Article 682132
Estimates the frequency distribution of a typological variable while controlling for statistical non-independence due to shared ancestry, then tests candidate word-order correlations across the 1,626 WALS languages that carry at least one word-order feature (175 Glottolog lineages, 81 of them isolates).
Jäger Methods
↗ DOI
2026
Jäger, Gerhard
Tracking the change that leads to typological variation: Word-order universals and the phylogenetic comparative method
Zeitschrift für Sprachwissenschaft, in press
Applies the phylogenetic comparative method to word-order universals, focusing on the diachronic changes that produce typological variation. Journal version of the plenary given at the DGfS annual meeting, Mainz 2025.
Jäger Methods
§ 5

Software & Practical Resources

For hands-on sessions
—
LingPy
Python library for quantitative historical linguistics
The primary software tool for sequence comparison, cognate detection, and alignment in computational historical linguistics. Developed by Johann-Mattis List and collaborators.
MethodsSoftware
↗ lingpy.org
—
BEAST 2
Bayesian Evolutionary Analysis by Sampling Trees
The standard platform for Bayesian phylogenetic inference. Used in most major studies including Bouckaert et al. (2012). Tutorials available at beast2.org.
MethodsSoftware
↗ beast2.org
—
CLDF & Glottolog / ASJP
Cross-Linguistic Data Formats; key databases
CLDF is the emerging standard for cross-linguistic data. Glottolog provides language classification and geographic data; ASJP provides a large database of Swadesh lists in a standardized transcription.
MethodsData
↗ cldf.clld.org
2016
RevBayes
Höhna et al., Systematic Biology 65(4), 726–736
A flexible, graphical-model framework for Bayesian phylogenetics that has become the research-standard complement to MrBayes and BEAST when a model needs to be built from scratch. Worth knowing about beyond the MrBayes lab.
MethodsSoftware
↗ revbayes.github.io
2017
List, Johann-Mattis, Simon J. Greenhill & Russell D. Gray
The Potential of Automatic Word Comparison for Historical Linguistics
PLOS ONE 12(1), e0170046
Introduces the LexStat-Infomap variant used as the standard baseline for automatic cognate detection, and assesses honestly how far the automatic methods can be trusted. Background for Lab 1.
MethodsSoftware
↗ DOI
2019
Rama, Taraka & Johann-Mattis List
An Automated Framework for Fast Cognate Detection and Bayesian Phylogenetic Inference in Computational Historical Linguistics
Proc. ACL 2019, 6225–6235
Defines the multi-family benchmark on which cognate-detection systems are now compared, and reports the B-cubed F scores quoted in Lecture 2.
MethodsSoftware
↗ ACL Anthology
2024
Akavarapu, V. S. D. S. Mahesh & Arnab Bhattacharya
Automated Cognate Detection as a Supervised Link Prediction Task with Cognate Transformer
Proc. EACL 2024, 965–975
The neural state of the art, and the reason the numbers in Lecture 2 are close: it edges past the LexStat-Infomap baseline, but only with supervision, and expert annotation remains the reference standard.
MethodsSoftware
↗ ACL Anthology
Gerhard Jäger · Universität Tübingen Updated for Naxos 2026