Multimodal Large Language Models for Understanding Human Self-Concept Evolution in Digital Social Environments

Authors

  • Milos J. Baker School of Information Technology, University of Cincinnati, Cincinnati, OH, USA.
  • Casper J. Kennedy Department of Computer Science, University of Houston, Houston, TX, USA.

Keywords:

multimodal large language models, self-concept, digital identity, social media, system architecture, algorithmic fairness, data governance

Abstract

The digitally mediated social world is now a primary arena in which human self-concept is formed, revised, and expressed. The proliferation of multimodal platforms, from image-centric social networks to short-video ecosystems, has transformed identity construction into a continuous, cross-channel, and algorithmically shaped process. Understanding how individuals’ self-representations evolve across these environments poses profound challenges for methods rooted in single-modality analysis or episodic self-report. This paper examines the potential of multimodal large language models (MLLMs) as systems instruments capable of capturing the textured, longitudinal dynamics of self-concept within digital social environments. We adopt a systems-level perspective that moves beyond model accuracy to interrogate the broader sociotechnical infrastructure required to deploy MLLMs responsibly. The discussion is organized around structural trade-offs, data governance, architectural choices, fairness considerations, robustness, sustainability, and policy implications. We argue that MLLMs, by virtue of their capacity to process and integrate text, image, video, and interaction metadata at scale, can serve as computational mirrors that reveal how platform affordances and algorithmic curation co-produce evolving self-views. However, harnessing this capacity demands rigorous attention to the inherent tensions between representational richness and privacy, between personalization and ecological validity, and between interpretability and model scale. The paper provides a conceptual analysis of these tensions, sketches design principles for a federated MLLM-based research infrastructure, and maps the regulatory landscape that must be navigated if such systems are to be deployed in ways that uphold human autonomy, fairness, and long-term trust.

References

1. Turkle, S. (1995). Life on the screen: Identity in the age of the Internet. Simon & Schuster.

2. boyd, d. (2014). It’s complicated: The social lives of networked teens. Yale University Press.

3. Goffman, E. (1959). The presentation of self in everyday life. Doubleday.

4. Markus, H. R., & Kitayama, S. (1991). Culture and the self: Implications for cognition, emotion, and motivation. Psychological Review, 98(2), 224–253.

5. Higgins, E. T. (1987). Self-discrepancy: A theory relating self and affect. Psychological Review, 94(3), 319–340.

6. Oyserman, D., Elmore, K., & Smith, G. (2012). Self, self-concept, and identity. In M. R. Leary & J. P. Tangney (Eds.), Handbook of self and identity (2nd ed., pp. 69–104). Guilford Press.

7. Pariser, E. (2011). The filter bubble: What the Internet is hiding from you. Penguin Press.

8. Sunstein, C. R. (2017). #Republic: Divided democracy in the age of social media. Princeton University Press.

9. McAdams, D. P. (2001). The psychology of life stories. Review of General Psychology, 5(2), 100–122.

10. Solanki, D., Hsu, H. M., Zhao, O., Zhang, R., Bi, W., & Kannan, R. (2020, July). The way we think about ourselves. In International Conference on Human-Computer Interaction (pp. 276-285). Cham: Springer International Publishing.

11. Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., ... Simonyan, K. (2022). Flamingo: a visual language model for few-shot learning. Advances in Neural Information Processing Systems, 35, 23716-23736.

12. Li, J., Li, D., Savarese, S., & Hoi, S. (2023). BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning (pp. 19730-19742). PMLR.

13. Liu, H., Li, C., Wu, Q., & Lee, Y. J. (2024). Visual instruction tuning. Advances in Neural Information Processing Systems, 36.

14. OpenAI. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.

15. Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., Silver, D., Petrov, S., Johnson, M., Antonoglou, I., Schrittwieser, J., Glaese, A., Chen, J., Pitler, E., ... Vinyals, O. (2023). Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.

16. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (pp. 8748-8763). PMLR.

17. Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1-35.

18. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610-623).

19. Zuboff, S. (2019). The age of surveillance capitalism. PublicAffairs.

20. Mittelstadt, B. D., Allo, P., Taddeo, M., Wachter, S., & Floridi, L. (2016). The ethics of algorithms: Mapping the debate. Big Data & Society, 3(2), 2053951716679679.

21. Floridi, L., & Cowls, J. (2019). A unified framework of five principles for AI in society. Harvard Data Science Review, 1(1).

22. Selbst, A. D., Boyd, D., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 59-68).

23. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645-3650).

24. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., ... Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.

25. Zhou, D. (2026, May). A Code Visualization Graph-Based Method for Vulnerability Severity Assessment Using a Multi-Scale Feature Fusion Network. In 2026 3rd International Conference on Image Processing and Artificial Intelligence (ICIPAI) (pp. 306-309). IEEE.

Downloads

Published

2026-05-01