Informace o publikaci

Enlargement of the Czech Question-Answering Dataset to SQAD v2.0

Autoři

ŠULGANOVÁ Terézia MEDVEĎ Marek HORÁK Aleš

Rok publikování 2017
Druh Článek ve sborníku
Konference Proceedings of the Eleventh Workshop on Recent Advances in Slavonic Natural Language Processing, RASLAN 2017
Fakulta / Pracoviště MU

Fakulta informatiky

Citace
www http://raslan2017.nlp-consulting.net/proceedings
Obor Informatika
Klíčová slova question answering; QA dataset; SQAD
Popis In this paper, we present the second version of Czech question-answering dataset called SQAD v2.0 (Simple Question Answering Database). The new version represents a large extension of our original SQAD database. In the current release, the dataset contains nearly 9,000 question-answer pairs completed with manual annotation of question and answer types. All texts in the dataset (the source documents, the question and the respective answer) are provided with complete morphological annotation in plain textual format. We offer detailed statistics of the SQAD v2.0 dataset based on the new QA annotation.
Související projekty: