Dataset Open Access

Vietic 116 item phylogenetic lexicon

Sidwell, Paul; Alves, Mark

The file is a116 item lexicostatistical dataset for classification of the Vietic languages. The set includes 30 Vietic doculects, Proto-Vietic, plus Khmu and Jahai as out-groups. Included is a listing of the sources, and the NEXUS file with our cognate value assignments, which we created to run on SplitsTree to generate phylograms and NeighborNets. The 116-item list was the outcome of beginning with the Swadesh 100 and 200 lists and reconciling these with the available data with the aim of achieving at least 80% coverage for each lect in the analysis. Procedurally, sources were selected and lexicons aggregated in a spreadsheet, with rows identified with Swadesh 100 and 200 items, subject to semantic and phonological adjustments as we judged necessary. For most of the languages, full coverage of the Swadesh 100 categories was not possible, with 20 or more gaps being common. Some 40 additional categories were added from the Swadesh 200 list, based on the 40 best represented items in the aggregated data, seeking to achieve a 120-item list with at least 100 items coverage for all lects, ultimately settling on 116 items.

The dataset was created for a paper provisionally entitled "The Vietic Languages: A Phylogenetic Analysis". The paper is submitted for journal publication and a version submitted for presentation at ICAAL9, November 2021. We encourage sharing for the purpose of testing/reproducing results, and augmented or derived studies under Creative Commons Attribution licence.
Files (199.8 kB)
Name Size
199.8 kB Download
All versions This version
Views 140140
Downloads 4141
Data volume 8.2 MB8.2 MB
Unique views 126126
Unique downloads 3838


Cite as