Computational Social Science Approach to Online Community Dynamics: Discovering Collective Identity Formation through Language Models

Authors

  • Henrik Powell Department of Electrical Engineering and Computer Science, University of Kansas, Lawrence, KS, USA.
  • Terry L. Becker School of Information Technology, University of Cincinnati, Cincinnati, OH, USA.
  • Surieaj Krishnean Department of Computer Science, University of Central Florida, Orlando, FL, USA.

Keywords:

computational social science, online communities, collective identity, language models, natural language processing, socio-technical systems, community governance, fairness, infrastructure

Abstract

Online communities function as dynamic socio-technical systems in which collective identity continuously emerges through discourse, shared practices, and evolving norms. Understanding how collective identity forms and transforms within these digitally mediated spaces remains a central challenge for both social theory and platform design. This paper presents a computational social science framework that leverages large-scale language models to discover and trace collective identity formation across diverse online communities. We articulate a systems-level architecture that integrates natural language processing pipelines with theoretical constructs from social identity theory, enabling the extraction of community-level narratives, value structures, and boundary-making language. The discussion emphasizes structural trade-offs in designing such discovery systems, including tensions between batch and stream processing, model interpretability, computational scalability, and privacy preservation. We examine governance implications arising from algorithmic surfacing of collective identity, addressing fairness, representation, and the risk of essentializing community voices. Further, we analyze infrastructure requirements for sustainable deployment and the robustness of identity models under linguistic drift and platform evolution. Through cross-domain case illustrations spanning open-source developer communities, health support forums, and fan subcultures, we highlight how language model outputs can inform platform moderation policies, community health diagnostics, and inclusive design. The paper contributes a holistic, systems-oriented perspective on computationally studying collective identity, foregrounding the interplay between technical architecture, social theory, and responsible deployment in large-scale online environments.

References

1. Lazer, D., Pentland, A., Adamic, L., Aral, S., Barabasi, A. L., Brewer, D., Christakis, N., Contractor, N., Fowler, J., Gutmann, M., Jebara, T., King, G., Macy, M., Roy, D., & Van Alstyne, M. (2009). Computational social science. Science, 323(5915), 721–723.

2. Tajfel, H., & Turner, J. C. (1986). The social identity theory of intergroup behavior. In S. Worchel & W. G. Austin (Eds.), Psychology of intergroup relations (pp. 7–24). Nelson-Hall.

3. boyd, d., & Ellison, N. B. (2007). Social network sites: Definition, history, and scholarship. Journal of Computer-Mediated Communication, 13(1), 210–230.

4. Kozinets, R. V. (2010). Netnography: Doing ethnographic research online. Sage Publications.

5. Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent Dirichlet allocation. Journal of Machine Learning Research, 3, 993–1022.

6. Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.

7. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186).

8. Hamilton, W. L., Leskovec, J., & Jurafsky, D. (2016). Diachronic word embeddings reveal statistical laws of semantic change. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (pp. 1489–1501).

9. Nguyen, D., Dogruoz, A. S., Rose, C. P., & de Jong, F. (2020). Computational sociolinguistics: A survey. Computational Linguistics, 46(3), 567–639.

10. Solanki, D., Hsu, H. M., Zhao, O., Zhang, R., Bi, W., & Kannan, R. (2020, July). The way we think about ourselves. In International Conference on Human-Computer Interaction (pp. 276-285). Cham: Springer International Publishing.

11. Bamman, D., Eisenstein, J., & Schnoebelen, T. (2014). Gender identity and lexical variation in social media. Journal of Sociolinguistics, 18(2), 135–160.

12. Kreps, J., Narkhede, N., & Rao, J. (2011). Kafka: A distributed messaging system for log processing. In Proceedings of the NetDB 2011 Workshop.

13. Armbrust, M., Fox, A., Griffith, R., Joseph, A. D., Katz, R. H., Konwinski, A., Lee, G., Patterson, D. A., Rabkin, A., Stoica, I., & Zaharia, M. (2010). A view of cloud computing. Communications of the ACM, 53(4), 50–58.

14. Barocas, S., Hardt, M., & Narayanan, A. (2019). Fairness and machine learning. fairmlbook.org.

15. Gillespie, T. (2018). Custodians of the Internet: Platforms, content moderation, and the hidden decisions that shape social media. Yale University Press.

16. Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J. F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. In Advances in Neural Information Processing Systems 28 (pp. 2503–2511).

17. Zliobaite, I. (2010). Change with delayed labeling: When is it detectable? In 2010 IEEE International Conference on Data Mining Workshops (pp. 843–850).

18. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623).

19. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645–3650).

20. Metcalf, J., & Crawford, K. (2016). Where are human subjects in big data research? The emerging ethics divide. Big Data & Society, 3(1), 1–14. 21.Zhou, D. (2026, May). A Code Visualization Graph-Based Method for Vulnerability Severity Assessment Using a Multi-Scale Feature Fusion Network. In 2026 3rd International Conference on Image Processing and Artificial Intelligence (ICIPAI) (pp. 306-309). IEEE.

Downloads

Published

2026-05-01