Cross-Linguistic Data Formats, advancing data sharing and re-use in comparative linguistics

Abstract
The amount of available digital data for the languages of the world is constantly increasing. Unfortunately, most of the digital data are provided in a large variety of formats and therefore not amenable for comparison and re-use. The Cross-Linguistic Data Formats initiative proposes new standards for two basic types of data in historical and typological language comparison (word lists, structural datasets) and a framework to incorporate more data types (e.g. parallel texts, and dictionaries). The new specification for cross-linguistic data formats comes along with a software package for validation and manipulation, a basic ontology which links to more general frameworks, and usage examples of best practices.
Authors
Citation
Forkel R, List J-M, Greenhill SJ, Bank S, Rzymski C, Cysouw M, Hammarström H, Haspelmath M & Kaiping GA & Gray RD. 2018. Cross-linguistic Data Formats, advancing data sharing and reuse in comparative linguistics. Scientific Data, 5:180205.
Published 20 June 2018
@article{Forkel2018,
title = {Cross-Linguistic Data Formats, advancing data sharing and re-use in comparative linguistics},
author = {Forkel, Robert and List, Johann-Mattis and Greenhill, Simon J. and Bank, Sebastian and Rzymski, Christoph and Cysouw, Michael and Hammarström, Harald and Haspelmath, Martin and Kaiping, Gereon A. and Gray, Russell D.},
journal = {Scientific Data},
volume = {5},
number = {1},
publisher = {Springer Science and Business Media LLC},
year = {2018},
doi = {10.1038/sdata.2018.205},
url = {https://simon.net.nz/articles/cross-linguistic-data-formats/},
}