Publication details

Detecting Subtle Sense Shift with Polysemy-Aware Trends

Investor logo
Authors

HERMAN Ondřej RYCHLÝ Pavel

Year of publication 2026
Type Paper in proceedings
Conference Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers)
MU Faculty or unit

Faculty of Informatics

Citation
web
Keywords lexical semantic change
Attached files
Description Language changes faster than dictionaries can be revised, yet automatic tools still struggle to spot the subtle, short-term shifts in meaning that precede a formal update. We present a language-independent pipeline that detects word-sense shifts in large, time-stamped web corpora. The method couples a robust re-implementation of the Adaptive Skip-Gram model, which induces multiple sense vectors per lemma without any external inventory, with a second stage that tracks each sense through time under three alternative frequency normalizations. Linear Regression and the robust Mann-Kendall/Theil–Sen estimator then test whether a sense’s frequency slope deviates significantly from zero, producing a ranked list of headwords whose semantics are drifting. We evaluate the system on the English (12,B tokens) and Czech (1,B tokens) Timestamped corpora for May 2023–May 2025. Expert annotation of the top-100 candidates for each model variant shows that 50.7,% of Czech and 25.7,% of English headwords exhibit genuine sense shifts, despite web-scale noise.
Related projects:

You are running an old browser version. We recommend updating your browser to its latest version.

More info