Informace o publikaci

Can large language models recognize complex language errors such as zeugma?

Autoři

MEDKOVÁ Helena HORÁK Aleš

Rok publikování 2026
Druh Recenzovaný odborný článek
Časopis / Zdroj Engineering Applications of Artificial Intelligence
Fakulta / Pracoviště MU

Filozofická fakulta

Citace
www https://www.sciencedirect.com/science/article/pii/S095219762601907X
Doi https://doi.org/10.1016/j.engappai.2026.115623
Klíčová slova Dataset; Error analysis; Explainable artificial intelligence; Knowledge transfer; Large language models; Zeugma error detection
Popis Recently, pre-trained large language models have opened up the possibility of efficiently detecting grammatical errors across many languages. This article focuses on zeugma, a linguistic structure that appears in many languages either as an intentional rhetorical device or as a grammatical error. Here, we examine zeugma primarily as a grammatical error. To identify erroneous zeugmatic constructions, we developed Zeugma Dataset 2.0, comprising over 10,000 sentences with various types of coordinated structures, including zeugma. We then fine-tuned and evaluated several pre-trained language models to develop ZeugBERT (Zeugma Bidirectional Encoder Representations from Transformers). The best-performing model, based on RobeCzech, achieved 90% sentence-level detection accuracy, significantly outperforming larger models such as RoBERTa. For comparison, we also evaluated the open-source GPT-OSS model (Generative Pretrained Transformer for Open-Source Systems) using zero-shot and few-shot prompting to assess whether a larger reasoning model can capture the nature of the task. The article describes the dataset and analyzes how ZeugBERT transfers knowledge from coordinated verbs to coordinated non-verbal constituents. We evaluate the model’s ability to generalize to previously unseen syntactic structures and show that performance improves when non-predicate coordinations are included in the training data. Additional fine-tuning experiments with different training-set compositions support these findings. To improve interpretability, we provide a detailed error analysis using Local Interpretable Model-agnostic Explanations and Transformer Interpret attention visualization. The dataset and trained models are publicly available to facilitate replication and future research.
Související projekty:

Používáte starou verzi internetového prohlížeče. Doporučujeme aktualizovat Váš prohlížeč na nejnovější verzi.

Další info