Dataset Open Access

Vietic 116 item phylogenetic lexicon

Sidwell, Paul; Alves, Mark

DataCite XML Export

<?xml version='1.0' encoding='utf-8'?>
<resource xmlns:xsi="" xmlns="" xsi:schemaLocation="">
  <identifier identifierType="DOI">10.5281/zenodo.5263195</identifier>
      <creatorName>Sidwell, Paul</creatorName>
      <nameIdentifier nameIdentifierScheme="ORCID" schemeURI="">0000-0002-9162-5668</nameIdentifier>
      <affiliation>University of Sydney</affiliation>
      <creatorName>Alves, Mark</creatorName>
      <nameIdentifier nameIdentifierScheme="ORCID" schemeURI="">0000-0001-7055-9182</nameIdentifier>
      <affiliation>Montgomery College</affiliation>
    <title>Vietic 116 item phylogenetic lexicon</title>
    <subject>Vietic, Austroasiatic, Swadesh list, phylogentics, lexicostatistics, Nexus file</subject>
    <date dateType="Issued">2021-08-26</date>
  <resourceType resourceTypeGeneral="Dataset"/>
    <alternateIdentifier alternateIdentifierType="url"></alternateIdentifier>
    <relatedIdentifier relatedIdentifierType="DOI" relationType="IsVersionOf">10.5281/zenodo.5263194</relatedIdentifier>
    <relatedIdentifier relatedIdentifierType="URL" relationType="IsPartOf"></relatedIdentifier>
  <version>First version (26 Aug 2021)eng</version>
    <rights rightsURI="">Creative Commons Attribution 4.0 International</rights>
    <rights rightsURI="info:eu-repo/semantics/openAccess">Open Access</rights>
    <description descriptionType="Abstract">&lt;p&gt;The file is a116 item lexicostatistical dataset for classification of the Vietic languages. The set includes 30 Vietic doculects, Proto-Vietic, plus Khmu and Jahai as out-groups. Included is a listing of the sources, and the NEXUS file with our cognate value assignments, which we created to run on SplitsTree to generate phylograms and NeighborNets. The 116-item list was the outcome of beginning with the Swadesh 100 and 200 lists and reconciling these with the available data with the aim of achieving at least 80% coverage for each lect in the analysis. Procedurally,&amp;nbsp;sources were selected and lexicons aggregated in a spreadsheet, with rows identified with Swadesh 100 and 200 items, subject to semantic and phonological adjustments as we judged necessary. For most of the languages, full coverage of the Swadesh 100 categories was not possible, with 20 or more gaps being common. Some 40 additional categories were added from the Swadesh 200 list, based on the 40 best represented items in the aggregated data, seeking to achieve a 120-item list with at least 100 items coverage for all lects, ultimately settling on 116 items.&lt;strong&gt; &lt;/strong&gt;&lt;/p&gt;</description>
    <description descriptionType="Other">The dataset was created for a paper provisionally entitled "The Vietic Languages: A Phylogenetic Analysis". The paper is submitted for journal publication and a version submitted for presentation at ICAAL9, November 2021. We encourage sharing for the purpose of testing/reproducing results, and augmented or derived studies under Creative Commons Attribution licence.</description>
All versions This version
Views 138138
Downloads 4141
Data volume 8.2 MB8.2 MB
Unique views 124124
Unique downloads 3838


Cite as