Research

SugarTextNet Targets Sugar-Dating Content Detection on Social Media

A preprint introduced SugarTextNet, a transformer-based system for detecting sugar-dating-related posts and euphemistic language on social media.

SugarTextNet is a research preprint proposing a specialized transformer model to identify sugar-dating-related posts on social media, where euphemisms and severe class imbalance make simple keyword filters unreliable. The authors evaluated the approach on 3,067 manually annotated Chinese-language posts collected from Sina Weibo. The architecture combines a pretrained transformer encoder, an attention-based cue extractor and a contextual phrase encoder to capture explicit and indirect language.

A Context-Aware Focal Loss function gives more weight to difficult minority-class examples, addressing the fact that relevant posts are rare compared with ordinary content. The authors report that SugarTextNet outperformed traditional machine-learning models, deep-learning baselines and large language models across multiple evaluation metrics. Ablation tests—removing individual components—were used to argue that each part of the model contributed to performance.

Domain-specific moderation may catch coded solicitation that general models miss, but accuracy on one curated dataset does not establish readiness for enforcement. False positives could suppress education, news, research or consensual discussion, while language changes quickly after moderation rules become known. Human review, appeal rights and subgroup testing are essential. The work is an arXiv preprint and had not necessarily completed peer review. The dataset is small, Chinese-language and platform-specific; the abstract does not establish performance across regions, dialects or changing euphemisms.

arXivPublished Nov 9, 2025
Read original ↗