Phylogenetic inference from cognate word forms
David M. Goldstein
1,*
,
Shawn H. McCreight
2
,
Éva Buchi
3
, and
John P. Huelsenbeck
4
1
Department of Linguistics,
University of California, Los Angeles, UCLA, Los Angeles, CA 90095-1543, USA
2
Nytril LLC
, 2904 Horsehead Bay Dr. NW, Gig Harbor, WA 98335, USA
3
Centre national de la recherche scientifique
, Nancy, FR
and Université de Lorraine
, Nancy, FR
4
Department of Integrative Biology,
University of California, Berkeley, UC Berkeley, Berkeley, CA 94720, USA
Abstract
Linguistic phylogenies are commonly inferred from abstract cognate classifications that encode relationships among lexemes. Although widespread,
this practice has well-recognized limitations: it discards the phylogenetic signal contained in segmental word forms; restricts the range of
evolutionary questions that can be addressed; and treats cognacy judgments, which are hypotheses, as observed data. We introduce a comparative
framework that addresses these limitations by modeling the evolution of aligned cognate word forms directly. Our approach adapts the TKF91
model of molecular evolution, originally developed to account for insertion and deletion events in DNA sequences, to the domain of linguistic data.
By operating on segmental strings rather than abstract character codings, the framework enables phylogenetic inference from observable word
forms and supports quantitative investigation of sound change. We demonstrate its utility through analyses that illuminate patterns of segmental
stability and the evolution of phonological inventories.
White Paper
Published version of the paper
Supplementary Material
Extra derivations and charts to supplement the main paper
Diagnostics
Summary of the data used in the experiment with
diagnostic views
Graphs
Graphs of links between facts
Cite
Citation information and links to the original publication
Overview
Overview of results
Replication
Guide to replicating the experiment
Acknowledgments
Acknowledgments and Author Biographies