شناسایی گروه‌های سخت‌قانون در پیام‌رسان تلگرام با استفاده از شبکه عصبی گرافی گراف‌سیج و داده‌های شبیه‌سازی‌شده دارای ویژگی‌های زمانی

نوع مقاله : مقاله پژوهشی

نویسندگان

1 دانشجوی دکتری، دانشکده مهندسی کامپیوتر، گروه مهندسی نرم‌افزار، دانشگاه یزد، یزد، ایران

2 دانشیار، دانشکده مهندسی کامپیوتر، گروه مهندسی نرم‌افزار، دانشگاه یزد، یزد، ایران

10.22034/abmir.2026.24432.1231

چکیده

این پژوهش به شناسایی «گروه‌های سخت‌قانون» در پیام‌رسان تلگرام می‌پردازد، گروه‌هایی که کاربران را برای عضویت یا ادامه فعالیت به انجام اقدامات اجباری مانند دعوت از دیگران یا عضویت در کانال‌های خاص ملزم کرده و روند رشد طبیعی گروه‌ها را دست‌کاری می‌کنند. تشخیص این گروه‌ها برای حفظ یکپارچگی پلتفرم و جلوگیری از سوءاستفاده اهمیت بالایی دارد. کمبود داده‌های برچسب‌دار واقعی، ماهیت پویا و زمانی رفتارها و ترکیب هم‌زمان سیگنال‌های ساختاری و رفتاری از چالش‌های این حوزه هستند. برای غلبه بر این چالش‌ها، مجموعه‌داده‌ای شبیه‌سازی‌شده در مقیاس بزرگ شامل ۱۰۰٬۰۰۰ کاربر، ۱۰٬۰۰۰ گروه و ۱۲ تصویر زمانی تولید شد که ویژگی‌های ایستا و زمانی کاربران و گروه‌ها را در برمی‌گیرد. برچسب‌ها با استفاده از یک امتیاز ریسک مبتنی بر تحلیل مؤلفه‌های اصلی تعیین و یک گراف دوبخشی کاربر-گروه ساخته شد. سپس یک مدل شبکه عصبی گرافی مبتنی بر گراف‌سیج با پروتکل ارزیابی کاملاً استقرایی، تابع زیان آنتروپی متقاطع دودویی وزن‌دار و بهینه‌سازی آستانه تصمیم آموزش داده و در کنار شش مدل پایه ارزیابی گردید. مدل پیشنهادی بر روی مجموعه آزمون مستقل به صحت ۲/۹۸ درصد، دقت کلاس مثبت ۲/۹۷ درصد، فراخوانی ۷/۹۳ درصد، امتیاز F۱ برابر ۴/۹۵ درصد، AUC برابر ۸۷/۹۹ درصد و میانگین دقت برابر ۵۱/۹۹ درصد دست‌یافت و در تمام بذرها پایدار بود. نتایج نشان داد به دلیل وابستگی ساختاری برچسب‌ها به ویژگی‌های ایستای گروه، مدل‌های جدولی با هزینه محاسباتی کمتر عملکرد اندکی بهتر دارند، درحالی‌که گراف‌سیج در گراف‌های بزرگ و پردرجه مقاوم می‌ماند. این یافته‌ها صرفاً بر روی داده‌های شبیه‌سازی‌شده تأیید شده‌اند و تعمیم آن‌ها به داده‌های واقعی تلگرام، به دلیل نبود داده‌های برچسب‌دار واقعی، همچنان چالشی باز است

کلیدواژه‌ها

موضوعات


عنوان مقاله [English]

Detection of Hard-Law Groups in Telegram Messenger Using GraphSAGE Graph Neural Networks and Simulated Data with Temporal Features

نویسندگان [English]

  • Maedeh Arab Bafrani 1
  • Mohammad Ali Zare Chahooki 2
1 Ph.D. Candidate, Department of Software Engineering, Faculty of Computer Engineering, Yazd University, Yazd, Iran
2 Associate Professor, Department of Software Engineering, Faculty of Computer Engineering, Yazd University, Yazd, Iran
چکیده [English]

This study addresses the detection of ‘Hard-law’ groups in the Telegram messaging platform, defined as groups that manipulate their organic growth by requiring users to perform mandatory actions, such as inviting additional users or joining specific channels, as a prerequisite for joining or continuing participation. Detecting such groups is essential for preserving platform integrity and mitigating abusive behaviors. The scarcity of labeled real-world data, the dynamic temporal nature of user behavior, and the need to jointly exploit structural and behavioral signals constitute the primary challenges in this domain. To address these challenges, a large-scale simulated dataset comprising 100,000 users, 10,000 groups, and 12 temporal snapshots was generated, capturing both static and temporal features of users and groups. Group labels were assigned using a principal component analysis (PCA)-based risk score, and a user–group bipartite graph was constructed. A GraphSAGE-based graph neural network was then trained using a fully inductive evaluation protocol, a weighted binary cross-entropy loss function, and decision threshold optimization, and its performance was compared with six baseline models. On an independent test set, the proposed model achieved an accuracy of 98.2٪, a positive-class precision of 97.2٪, a recall of 93.7٪, an F1-score of 95.4٪, an area under the ROC curve (AUC) of 99.87٪, and an average precision (AP) of 99.51٪, while demonstrating stable performance across all random seeds. The results indicate that, because the labels are structurally dependent on static group attributes, tabular machine learning models achieve slightly better performance with lower computational cost, whereas GraphSAGE remains robust on large-scale, high-degree graphs. These findings have been validated only on simulated data, and their generalization to real-world Telegram data remains an open challenge due to the lack of labeled datasets.

کلیدواژه‌ها [English]

  • Hard-law groups
  • Telegram messenger
  • Graph neural networks
  • GraphSAGE
  • Simulated dataset
  • Temporal features
  • Coercive behavior detection
  • Graph-based deep learning
[1]     D. Javed, N. Z. Jhanjhi, N. A. Khan, S. K. Ray, A. Al Mazroa, F. Ashfaq, and S. R. Das, "Towards the future of bot detection: A comprehensive taxonomical review and challenges on Twitter/X," Computer Networks, vol. 254, Art. no. 110808, 2024.
[2]     A. Beutel, W. Xu, V. Guruswami, C. Palow, and C. Faloutsos, "CopyCatch: Stopping group attacks by spotting lockstep behavior in social networks," in Proc. 22nd Int. World Wide Web Conf. (WWW), Rio de Janeiro, Brazil, 2013, pp. 119–130.
[3]     A. Buitrago López et al., "Agent-based simulation of online social networks and disinformation," arXiv preprint arXiv:2512.22082, 2025.
[4]     Y. Dou, Z. Liu, L. Sun, Y. Deng, H. Peng, and P. S. Yu, "Enhancing graph neural network-based fraud detectors against camouflaged fraudsters," in Proc. 29th ACM Int. Conf. on Information and Knowledge Management (CIKM), 2020, pp. 315–324.
[5]     A. Davies and N. Ajmeri, "Realistic synthetic social networks with graph neural networks," arXiv preprint arXiv:2212.07843, 2022.
[6]     N. Jiang, F. Yin, B. Wang, and A. T. Crooks, "A large-scale geographically explicit synthetic population with social networks for the United States," Scientific Data, vol. 11, Art. no. 1204, 2024, doi: 10.1038/s41597-024-03970-1.
[7]     N. Jiang, A. T. Crooks, H. Kavak, A. Burger, and W. G. Kennedy, "A method to create a synthetic population with social networks for geographically-explicit agent-based models," Computational Urban Science, vol. 2, Art. no. 7, 2022, doi: 10.1007/s43762-022-00034-1.
[8]     W. L. Hamilton, R. Ying, and J. Leskovec, "Inductive representation learning on large graphs," in Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 1024–1034.
[9]     A. Bojchevski, O. Shchur, D. Zügner, and S. Günnemann, "NetGAN: Generating graphs via random walks," in Proc. 35th Int. Conf. on Machine Learning (ICML), PMLR, vol. 80, 2018, pp. 610–619.
[10] J. You, R. Ying, X. Ren, W. L. Hamilton, and J. Leskovec, "GraphRNN: Generating realistic graphs with deep auto-regressive models," in Proc. 35th Int. Conf. on Machine Learning (ICML), PMLR, 2018, pp. 5708–5717.
[11] A. Zareie et al., "Identifying coordination in online social networks through anomalous sharing behaviour," Online Social Networks and Media, vol. 50, Art. no. 100341, 2025.
[12] R. Rogers and N. Righetti, "Coordinated inauthentic behaviour on Facebook? A typology of manufactured attention," Platforms & Society, vol. 2, Art. no. 29768624251369784, 2025.
[13] M. Cinelli et al., "Coordinated inauthentic behavior and information spreading on Twitter," Decision Support Systems, vol. 160, Art. no. 113819, 2022.
[14] L. Luceri et al., "Coordinated inauthentic behavior on TikTok: Challenges and opportunities for detection in a video-first ecosystem," in Proc. Int. AAAI Conf. on Web and Social Media, vol. 20, no. 1, 2026.
[15] F. Cinus et al., "Exposing cross-platform coordinated inauthentic activity in the run-up to the 2024 US election," in Proc. ACM Web Conf. 2025, 2025.
[16] P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, "Graph attention networks," arXiv preprint arXiv:1710.10903, 2017.
[17] Y. Liu et al., "Heterogeneous graph convolutional network for rumor detection with multi-level interactive fusion and graph reconstruction," Scientific Reports, vol. 15, no. 1, Art. no. 31639, 2025.
[18] S. Shehnepoor et al., "Spatio-temporal graph representation learning for fraudster group detection," IEEE Trans. Neural Netw. Learn. Syst., vol. 35, no. 5, pp. 6628–6642, 2022.
[19] M. Boyapati and R. Aygun, "BalancerGNN: Balancer graph neural networks for imbalanced datasets: A case study on fraud detection," Neural Networks, vol. 182, Art. no. 106926, 2025.
[20] D. Saldaña-Ulloa, G. De Ita Luna, and J. R. Marcial-Romero, "A temporal graph network algorithm for detecting fraudulent transactions on online payment platforms," Algorithms, vol. 17, no. 12, Art. no. 552, 2024.
[21] Z. Qu et al., "Temporal enhanced multimodal graph neural networks for fake news detection," IEEE Trans. Comput. Soc. Syst., vol. 11, no. 6, pp. 7286–7298, 2024.
[22] T. N. Kipf and M. Welling, "Semi-supervised classification with graph convolutional networks," in Proc. Int. Conf. on Learning Representations (ICLR), 2017.
[23] Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and P. S. Yu, "A comprehensive survey on graph neural networks," IEEE Transactions on Neural Networks and Learning Systems, vol. 32, no. 1, pp. 4–24, 2021.
[24] X. Ma, J. Wu, S. Xue, J. Yang, C. Zhou, Q. Z. Sheng, H. Xiong, and L. Akoglu, "A comprehensive survey on graph anomaly detection with deep learning," IEEE Transactions on Knowledge and Data Engineering, vol. 35, no. 12, pp. 12012–12038, 2023.
[25] H. Qiao, H. Tong, B. An, I. King, C. Aggarwal, and G. Pang, "Deep graph anomaly detection: A survey and new perspectives," arXiv preprint arXiv:2409.09957, 2024.
[26] B. Khemani, S. Patil, K. Kotecha, and S. Tanwar, "A review of graph neural networks: concepts, architectures, techniques, challenges, datasets, applications, and future directions," Journal of Big Data, vol. 11, Art. no. 18, 2024.
[27] M. Fey and J. E. Lenssen, "Fast graph representation learning with PyTorch Geometric," in ICLR Workshop on Representation Learning on Graphs and Manifolds, 2019.
[28] N. Patki, R. Wedge, and K. Veeramachaneni, "The synthetic data vault," in Proc. IEEE Int. Conf. on Data Science and Advanced Analytics (DSAA), 2016, pp. 399–410.
[29] P. W. Holland, K. B. Laskey, and S. Leinhardt, "Stochastic blockmodels: First steps," Social Networks, vol. 5, no. 2, pp. 109–137, 1983.
[30] I. T. Jolliffe, Principal Component Analysis, 2nd ed. New York, NY: Springer, 2002.
[31] W. Chen, K. Yang, Z. Yu, Y. Shi, and C. L. P. Chen, "A survey on imbalanced learning: latest research, applications and future directions," Artificial Intelligence Review, vol. 57, no. 6, Art. no. 137, 2024.
[32] M. McPherson, L. Smith-Lovin, and J. M. Cook, "Birds of a feather: Homophily in social networks," Annual Review of Sociology, vol. 27, pp. 415–444, 2001.
[33] G. E. P. Box, G. M. Jenkins, G. C. Reinsel, and G. M. Ljung, Time Series Analysis: Forecasting and Control, 5th ed. Hoboken, NJ: Wiley, 2015.
[34] S. Sahu, "A Fast Parallel Approach for Neighborhood-based Link Prediction by Disregarding Large Hubs," arXiv preprint arXiv:2401.11415, 2024.
[35] H. H. Rashidi, S. Albahra, B. Hu, and B. P. Rubin, "A novel and fully automated platform for synthetic tabular data generation and validation," Scientific Reports, vol. 14, 2024.
[36] P. A. Apellániz, A. Jiménez, B. Arroyo Galende, J. Parras, and S. Zazo, "Synthetic Tabular Data Validation: A Divergence-Based Approach," IEEE Access, vol. 12, 2024.
[37] N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, "Dropout: A simple way to prevent neural networks from overfitting," Journal of Machine Learning Research, vol. 15, pp. 1929–1958, 2014.
[38] Q. Li, Z. Han, and X.-M. Wu, "Deeper insights into graph convolutional networks for semi-supervised learning," in Proc. 32nd AAAI Conf. on Artificial Intelligence, 2018, pp. 3538–3545.
[39] K. Oono and T. Suzuki, "Graph neural networks exponentially lose expressive power for node classification," in Proc. Int. Conf. on Learning Representations (ICLR), 2020.
[40] L. Breiman, "Random forests," Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
[41] T. Chen and C. Guestrin, "XGBoost: A scalable tree boosting system," in Proc. 22nd ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining (KDD), 2016, pp. 785–794.
[42] D. P. Kingma and J. Ba, "Adam: A method for stochastic optimization," in Proc. Int. Conf. on Learning Representations (ICLR), 2015.
[43] T. Saito and M. Rehmsmeier, "The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets," PLoS ONE, vol. 10, no. 3, Art. no. e0118432, 2015.
[44] [44] J. Davis and M. Goadrich, "The relationship between precision-recall and ROC curves," in Proc. 23rd Int. Conf. on Machine Learning (ICML), 2006, pp. 233–240.
[45] A. Foucart, A. Elskens, and C. Decaestecker, "Ranking the scores of algorithms with confidence," in Proc. ESANN 2025.
[46] J. Demšar, "Statistical comparisons of classifiers over multiple data sets," Journal of Machine Learning Research, vol. 7, pp. 1–30, 2006.
[47] Y. Zheng, L. Yi, and Z. Wei, "A survey of dynamic graph neural networks," Frontiers of Computer Science, vol. 19, no. 6, Art. no. 196323, 2025.
[48] E. Rossi, B. Chamberlain, F. Frasca, D. Eynard, F. Monti, and M. Bronstein, "Temporal graph networks for deep learning on dynamic graphs," in ICML Workshop on Graph Representation Learning, 2020.