Zde se nacházíte:
Informace o publikaci
Can large language models recognize complex language errors such as zeugma?
| Autoři | |
|---|---|
| Rok publikování | 2026 |
| Druh | Recenzovaný odborný článek |
| Časopis / Zdroj | Engineering Applications of Artificial Intelligence |
| Fakulta / Pracoviště MU | |
| Citace | |
| www | https://www.sciencedirect.com/science/article/pii/S095219762601907X |
| Doi | https://doi.org/10.1016/j.engappai.2026.115623 |
| Klíčová slova | Dataset; Error analysis; Explainable artificial intelligence; Knowledge transfer; Large language models; Zeugma error detection |
| Popis | Recently, pre-trained large language models have opened up the possibility of efficiently detecting grammatical errors across many languages. This article focuses on zeugma, a linguistic structure that appears in many languages either as an intentional rhetorical device or as a grammatical error. Here, we examine zeugma primarily as a grammatical error. To identify erroneous zeugmatic constructions, we developed Zeugma Dataset 2.0, comprising over 10,000 sentences with various types of coordinated structures, including zeugma. We then fine-tuned and evaluated several pre-trained language models to develop ZeugBERT (Zeugma Bidirectional Encoder Representations from Transformers). The best-performing model, based on RobeCzech, achieved 90% sentence-level detection accuracy, significantly outperforming larger models such as RoBERTa. For comparison, we also evaluated the open-source GPT-OSS model (Generative Pretrained Transformer for Open-Source Systems) using zero-shot and few-shot prompting to assess whether a larger reasoning model can capture the nature of the task. The article describes the dataset and analyzes how ZeugBERT transfers knowledge from coordinated verbs to coordinated non-verbal constituents. We evaluate the model’s ability to generalize to previously unseen syntactic structures and show that performance improves when non-predicate coordinations are included in the training data. Additional fine-tuning experiments with different training-set compositions support these findings. To improve interpretability, we provide a detailed error analysis using Local Interpretable Model-agnostic Explanations and Transformer Interpret attention visualization. The dataset and trained models are publicly available to facilitate replication and future research. |
| Související projekty: |