Energy-Efficient LLM Inference for Real-Time Edge Intelligence Applications
Keywords:
edge intelligence; large language models; energy-efficient inference; real-time systems; model compression; sustainability; algorithmic governanceAbstract
Large language models have become central to natural language processing, but their inference demands create substantial energy, memory, and latency challenges for real-time edge intelligence systems. Edge devices operate under strict power envelopes, thermal limits, and connectivity constraints, while many language model inference workloads are memory-bound and produce highly variable token-level computational demands. This paper examines energy-efficient large language model inference from a system-level perspective rather than treating efficiency as an isolated algorithmic concern. It analyzes structural trade-offs across model architecture, compression, quantization, runtime scheduling, memory management, heterogeneous hardware, and edge-cloud partitioning. The discussion emphasizes interdependence among these architectural and operational choices, showing that gains in one layer can be undermined by inefficiencies in another. The paper further considers the governance, fairness, sustainability, and policy dimensions of deploying compressed language models at the edge. It argues that efficiency cannot be reduced to floating-point operations per inference, because the organizational and environmental costs of edge artificial intelligence are shaped by data movement, lifecycle carbon emissions, auditability, and the distribution of performance across diverse linguistic and demographic groups. A coordinated research agenda is proposed in which efficiency, robustness, equity, and regulatory compliance are treated as co-design requirements rather than post hoc constraints. The conclusion identifies future directions for energy-aware runtime systems, transparent reporting standards, and hardware-software co-design for sustainable edge intelligence.
References
1. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
2. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shinn, T., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., & Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
3. Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645–3650.
4. Schwartz, R., Dodge, J., Smith, N. A., & Etzioni, O. (2020). Green AI. Communications of the ACM, 63(12), 54–63.
5. Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., & Dean, J. (2021). Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350.
6. Han, S., Mao, H., & Dally, W. J. (2016). Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. International Conference on Learning Representations.
7. Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531.
8. Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2704–2713.
9. Frankle, J., & Carbin, M. (2019). The lottery ticket hypothesis: Finding sparse, trainable neural networks. International Conference on Learning Representations.
10. Tan, M., & Le, Q. V. (2019). EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning, 6105–6114.
11. Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Cowan, M., Shen, H., Wang, L., Hu, Y., Ceze, L., Guestrin, C., & Krishnamurthy, A. (2018). TVM: An automated end-to-end optimizing compiler for deep learning. Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation, 578–594.
12. Kang, Y., Hauswald, J., Gao, C., Rovinski, A., Mudge, T., Mars, J., & Tang, L. (2017). Neurosurgeon: Collaborative intelligence between the cloud and mobile edge. Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems, 615–629.
13. Li, Q. (2026, May). TrajLite PPO: Scalable Reasoning Path Filtering and Alignment for Small Parameter Large Language Models. In 2026 2nd International Conference on Artificial Intelligence and Digital Ethics (ICAIDE) (pp. 29-32). IEEE.
14. Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2022). LLM.int8(): 8-bit matrix multiplication for transformers at scale. Advances in Neural Information Processing Systems, 35, 30318–30332.
15. Mao, Y., You, C., Zhang, J., Huang, K., & Letaief, K. B. (2017). A survey on mobile edge computing: The communication perspective. IEEE Communications Surveys & Tutorials, 19(4), 2322–2358.
16. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–623.
17. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N. S., Chen, A. S., Creel, K. A., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etzioni, O., Fiete, I., Finn, C., Garg, A., Goh, G., Goodman, N., Grusky, E., Gururangan, S., Hashimoto, T., Hashimoto, T. B., He, H., Horvitz, E., Koyejo, O., Kraut, R. E., Krishnan, A., Ladhak, F., Lang, H., Lee, J., Liang, P., Ma, S., Ma, Y., Manning, C. D., Meek, C., Mitchell, M., Miller, J., Milstein, B., Mohamed, S., Mozer, M., Narayanan, D., Neubig, G., Newton, L., Orr, L., Passos, A., Pierson, E., Qin, L., Radford, A., Ranzato, M., Riedel, S., Rogers, A., Romera-Paredes, B., Ruder, S., Russell, S. J., Sahoo, D., Schoenick, C., Sermanet, P., Sharma, P., Shridhar, M., Sidor, J., Simchowitz, M., Snell, C., Song, D., Spelda, P., Strubell, E., Subramanian, S., Tamkin, A., Tenenbaum, J. B., Tishby, N., Toulouse, T., Varoquaux, G., Vries, H. D., Wang, A., Wang, C., Wang, X., Weidinger, L., Weller, A., Weld, D. S., Wilson, A. G., Winstein, K., Wu, J., Xie, S. M., Yatskar, M., Yurochkin, M., Zettlemoyer, L., & Liang, P. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
18. Dhar, P. (2020). The carbon impact of artificial intelligence. Nature Machine Intelligence, 2(8), 423–425.
19. European Commission. (2024). Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union.
20. Luccioni, A. S., Viguier, S., & Ligozat, A.-L. (2023). Estimating the carbon footprint of BLOOM, a 176B parameter language model. Journal of Machine Learning Research, 24(253), 1–15.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Engineering Systems and Digital Innovation

This work is licensed under a Creative Commons Attribution 4.0 International License.