tidyX

tidyX turns raw Spanish-language social-media text into clean corpora ready for NLP models. I co-authored it after running into the same preprocessing problems again and again while measuring online xenophobia.
It handles lemmatization, emoji and special Unicode characters, stop-word removal and tokenization tuned for Spanish social-media data. Install from PyPI
