Informace o publikaci

Automating Semantic Annotation in Low-Resource Languages: Evaluating GPT-4 for Urdu NLP

Autoři

RAHMAN Gohar

Rok publikování 2026
Druh Recenzovaný odborný článek
Časopis / Zdroj AI-Linguistica. Linguistic Studies on AI-Generated Texts and Discourses
Fakulta / Pracoviště MU

Pedagogická fakulta

Citace
www https://doi.org/10.62408/ai-ling.v3i1.40
Doi https://doi.org/10.62408/ai-ling.v3i1.40
Klíčová slova semantic annotation; Urdu NLP; low-resource languages; prompt engineering; named entity recognition
Popis Semantic annotation is a fundamental yet labor-intensive process essential for building effective Natural Language Processing systems, particularly for low-resource languages such as Urdu. The limited availability of large, manually annotated datasets has constrained advancements in Urdu NLP. This study explores the potential of automating semantic annotation using GPT-4, a state-of-the-art large language model, through structured prompt engineering without task-specific fine-tuning. A corpus of 50,000 Urdusentences spanning news articles, social media posts, and literary texts was used to evaluate three core tasks: Named Entity Recognition,semantic similarity, and sentiment analysis. GPT-4 demonstrated strong performance, achieving an F1-score of 92% for NER, a Pearson correlation of 0.87 for semantic similarity, and an accuracy of 88% with a macro-F1 of 87% for sentiment classification. These results indicate that LLMs guided by instruction-based prompts can reliably perform complex NLP tasks in low-resource contexts. Nonetheless, challenges with idiomatic expressions, sarcasm, and rare entities highlight the need for carefully designed prompts and potential human-AI collaboration

Používáte starou verzi internetového prohlížeče. Doporučujeme aktualizovat Váš prohlížeč na nejnovější verzi.

Další info