Building energy forecasting suffers from unavailable sensor data and hard-to-model long-range nonlinear time correlations; classic ML models rely on heavy feature engineering.
A GPT-style Transformer teacher model (GPT-BEF) learns full multivariate time series to capture complex building energy temporal patterns.
Knowledge distillation transfers teacher knowledge to a lightweight student model using merely two temperature input features.
GPT-BEF+KD cuts RMSE by ~25% versus two-feature Random Forest, retaining strong short- and long-term temporal capture capacity.
This work validates GPT-based distillation for sparse-sensor building energy prediction and supports follow-up energy management research.
| [1] | Jia R., Jin B., Jin M., Zhou Y., Konstantakopoulos I.C., Zou H., et al. (2018). Design automation for smart building systems. Proc. IEEE 106(9):1680−1699. DOI:10.1109/JPROC.2018.2860343 |
| [2] | Buckman A.H., Mayfield M., Beck S.B.M. (2014). What is a smart building. Smart Sustain. Built Environ. 3(2):92−109. DOI:10.1080/20416622.2014.911037 |
| [3] | Zhang X.P., Cheng X.M. (2009). Energy consumption, carbon emissions, and economic growth in china. Ecol. Econ. 68(10):2706−2712. DOI:10.1016/j.ecolecon.2009.05.011 |
| [4] | Bansal M.A., Sharma D.R., Kathuria D.M. (2022). A systematic review on data scarcity problem in deep learning. ACM Comput. Surv. 54(10s):1−29. DOI:10.1145/3563425 |
| [5] | Jiang R., Zeng S., Song Q., Wu Z. (2022). Deep-chain echo state network with explainable temporal dependence for complex building energy prediction. IEEE Trans. Ind. Inform. 19(1):426−435. DOI:10.1109/TII.2022.3170672 |
| [6] | Garza A., Challu C., Mergenthaler-Canseco M. (2023). Timegpt-1. arXiv preprint arXiv:2310.03589. DOI:10.48550/arXiv.2310.03589 |
| [7] | Liu Y., Zhang H., Li C., Huang X., Wang J., Long M. (2024). Timer: Generative pre-trained transformers are large time series models. arXiv preprint arXiv:2402.02368. DOI:10.48550/arXiv.2402.02368 |
| [8] | Mirchandani S., Xia F., Florence P., Ichter B., Driess D., Arenas M.G., et al. (2023). Large language models as general pattern machines. arXiv preprint arXiv:2307.04721. DOI:10.48550/arXiv.2307.04721 |
| [9] | Kraus M. (1987). Energy forecasting: The epistemological context. Futures 19(3):254−275. DOI:10.1016/0016-3287(87)90018-7 |
| [10] | Dai X., Cheng S., Chong A. (2023). Deciphering optimal mixed-mode ventilation in the tropics using reinforcement learning with explainable artificial intelligence. Energy Build. 278:112629. DOI:10.1016/j.enbuild.2023.112629 |
| [11] | Pérez-Lombard L., Ortiz J., Pout C. (2008). A review on buildings energy consumption information. Energy Build. 40(3):394−398. DOI:10.1016/j.enbuild.2007.03.007 |
| [12] | Wei N., Li C., Peng X., Zeng F., Lu X. (2019). Conventional models and artificial intelligence-based models for energy consumption forecasting. J. Pet. Sci. Eng. 181:106187. DOI:10.1016/j.petrol.2019.106187 |
| [13] | Ahmed N.K., Atiya A.F., Gayar N.E., El-Shishiny H. (2010). An empirical comparison of machine learning models for time series forecasting. Econometr. Rev. 29(5-6):594−621. DOI:10.1080/07474938.2010.481027 |
| [14] | Hammad M.A., Jereb B., Rosi B., Dragan D. (2020). Methods and models for electric load forecasting. Logist. Supply Chain Sustain. Glob. Chall. 11(1):51−76. DOI:10.3390/logistics11010051 |
| [15] | Vagropoulos S.I., Chouliaras G., Kardakos E.G., Simoglou C.K., Bakirtzis A.G. (2016). Comparison of sarimax, sarima, modified sarima and ann-based models for short-term pv generation forecasting. In: 2016 IEEE International Energy Conference (ENERGYCON), pp:1–6. DOI:10.1109/ENERGYCON.2016.7514066 |
| [16] | El-Kenawy E.S.M., Ibrahim A., Alhussan A.A., Khafaga D.S., Ahmed A., Eid M., et al. (2026). Smart city electricity load forecasting using greylag goose optimization-enhanced time series analysis. Arab. J. Sci. Eng. 51(6):8359−8377. DOI:10.1007/s13369-025-09421 |
| [17] | Gardner E.S. Jr. (1985). Exponential smoothing: The state of the art. J. Forecast. 4(1):1−28. DOI:10.1002/for.39800401 |
| [18] | Chatfield C., Yar M. (1988). Holt-winters forecasting: some practical issues. Statistician 37(2):129−140. DOI:10.2307/2348673 |
| [19] | Dai X., Liu J., Zhang X. (2020). A review of studies applying machine learning models to predict occupancy and window-opening behaviours in smart buildings. Energy Build. 223:110159. DOI:10.1016/j.enbuild.2020.110159 |
| [20] | Dhiman H.S., Deb D., Guerrero J.M. (2019). Hybrid machine intelligent svr variants for wind forecasting and ramp events. Renew. Sustain. Energy Rev. 108:369−379. DOI:10.1016/j.rser.2019.03.032 |
| [21] | Wang Z., Wang Y., Zeng R., Srinivasan R.S., Ahrentzen S. (2018). Random forest based hourly building energy prediction. Energy Build. 171:11−25. DOI:10.1016/j.enbuild.2018.04.032 |
| [22] | Chen T., Guestrin C. (2016). Xgboost: A scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp:785–794. DOI:10.1145/2939672.2939785 |
| [23] | Chen C., Zhang Q., Ma Q., Yu B. (2019). Lightgbm-ppi: Predicting protein-protein interactions through multi-information fusion. Chemom. Intell. Lab. Syst. 191:54−64. DOI:10.1016/j.chemolab.2019.05.007 |
| [24] | Breiman L. (2001). Random forests. Mach. Learn. 45:5−32. DOI:10.1023/A:1010933404324 |
| [25] | Kremic E., Subasi A. (2016). Performance of random forest and svm in face recognition. Int. Arab J. Inf. Technol. 13(2):287−293. DOI:10.34069/iajit.2016.13.2.7 |
| [26] | Alhussan A.A., El-Kenawy E.S.M., Eid M.M., Khodadadi N. (2026). Hybrid al-biruni and puma optimization (berpo) for boosting the classification of quality-of-service (qos) in 5g networks. J. Netw. Comput. Appl. :104462. DOI:10.1016/j.jnca.2026.104462 |
| [27] | Lim B., Arık S.Ö., Loeff N., Pfister T. (2021). Temporal fusion transformers for interpretable multi-horizon time series forecasting. Int. J. Forecast. 37(4):1748−1764. DOI:10.1016/j.ijforecast.2021.03.012 |
| [28] | Zhou H., Zhang S., Peng J., Zhang S., Li J., Xiong H., et al. (2021). Informer: Beyond efficient transformer for long sequence time-series forecasting. In: AAAI Conference Proceedings, vol.35, pp:11106–11115. DOI:10.1609/aaai.v35i12.17325 |
| [29] | Wu H., Xu J., Wang J., Long M. (2021). Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Adv. Neural Inf. Process. Syst. 34:22419−22430. DOI:10.48550/arXiv.2106.13008 |
| [30] | Zhou T., Ma Z., Wen Q., Wang X., Sun L., Jin R. (2022). Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. In: International Conference on Machine Learning, PMLR, pp:27268–27286. DOI:10.48550/arXiv.2201.12740 |
| [31] | Radwan M., Ibrahim A., Abdelsalam M.M., Alhussan A.A., Mattar E.A., El-Kenawy E.S.M., et al. (2026). Optimizing solar and wind forecasting with ihow optimization algorithm and multi-scale attention networks. Sci. Rep. DOI:10.1038/s41598-026-84671. |
| [32] | Dai T.Y., Niyogi D., Nagy Z. (2025). Citytft: A temporal fusion transformer-based surrogate model for urban building energy modeling. Appl. Energy 389:125712. DOI:10.1016/j.apenergy.2025.125712 |
| [33] | Gu Y., Jazizadeh F., Wang X. (2025). Toward large energy models: A comparative study of transformers’ efficacy for energy forecasting. Appl. Energy 384:125358. DOI:10.1016/j.apenergy.2025.125358 |
| [34] | Xing Z., Pan Y., Yang Y., Yuan X., Liang Y., Huang Z., et al. (2024). Transfer learning integrating similarity analysis for short-term and long-term building energy consumption prediction. Appl. Energy 365:123276. DOI:10.1016/j.apenergy.2024.123276 |
| [35] | Liang X., Chen S., Zhu X., Jin X., Du Z. (2023). Domain knowledge decomposition of building energy consumption and a hybrid data-driven model for 24-h ahead predictions. Appl. Energy 344:121244. DOI:10.1016/j.apenergy.2023.121244 |
| [36] | Hu Z., Gao Y., Ji S., Mae M., Imaizumi T. (2024). Improved multistep ahead photovoltaic power prediction model based on lstm and self-attention with weather forecast data. Appl. Energy 359:122709. DOI:10.1016/j.apenergy.2024.122709 |
| [37] | Karijadi I., Chou S.Y. (2022). A hybrid rf-lstm based on ceemdan for improving the accuracy of building energy consumption prediction. Energy Build. 259:111908. DOI:10.1016/j.enbuild.2022.111908 |
| [38] | El-Kenawy E.S.M., Khodadadi N., Mirjalili S., Zaki A.M., Ibrahim A., Alhussan A.A., et al. (2026). Glider snake optimizer (gso): a nature-inspired metaheuristic algorithm for global and engineering optimization problems. Artif. Intell. Rev. 59:91. DOI:10.1007/s10462-025-10876 |
| [39] | Jang J., Han J., Leigh S.B. (2022). Prediction of heating energy consumption with operation pattern variables for non-residential buildings using lstm networks. Energy Build. 255:111647. DOI:10.1016/j.enbuild.2022.111647 |
| [40] | Liu J., Zhang Y., Wen K., Ding Y. (2025). Analysis of different neural network models based on variational mode decomposition and dung beetle optimizer for the prediction of air-conditioning energy consumption in multifunctional complex large public buildings. Energy Build. 334:115518. DOI:10.1016/j.enbuild.2025.115518 |
| [41] | Brants T., Popat A., Xu P., Och F.J., Dean J. (2007). Large language models in machine translation. In: EMNLP-CoNLL Conference Proceedings, pp:858–867. DOI:10.3115/1079540. |
| [42] | Dathathri S., Madotto A., Lan J., Hung J., Frank E., Molino P., et al. (2019). Plug and play language models: A simple approach to controlled text generation. arXiv preprint arXiv:1912.02164. DOI:10.48550/arXiv.1912.02164. |
| [43] | Izadi M., Katzy J., Van Dam T., Otten M., Popescu R.M., Van Deursen A. (2024). Language models for code completion. In: IEEE/ACM 46th International Conference on Software Engineering, pp:1–13. DOI:10.1145/3597643.3597682. |
| [44] | Yu X., Chen Z., Ling Y., Dong S., Liu Z., Lu Y. (2023). Temporal data meets llm–explainable financial time series forecasting. arXiv preprint arXiv:2306.11025. DOI:10.48550/arXiv.2306.11025 |
| [45] | de Zarzà I., de Curtò J., Roig G., Calafate C.T. (2023). Llm multimodal traffic accident forecasting. Sensors 23(22):9225. DOI:10.3390/s23229225 |
| [46] | Thirunavukarasu A.J., Ting D.S.J., Elangovan K., Gutierrez L., Tan T.F., Ting D.S.W. (2023). Large language models in medicine. Nat. Med. 29(8):1930−1940. DOI:10.1038/s41591-023-02488 |
| [47] | Jin M., Wang S., Ma L., Chu Z., Zhang J.Y., Shi X., et al. (2023). Time-llm: Time series forecasting by reprogramming large language models. arXiv preprint arXiv:2310.01728. DOI:10.48550/arXiv.2310.01728 |
| [48] | Liang Y., Wen H., Nie Y., Jiang Y., Jin M., Song D., et al. (2024). Foundation models for time series analysis: A tutorial and survey. In: 30th ACM SIGKDD Conference Proceedings, pp:6555–6565. DOI:10.1145/3637462.3637560 |
| [49] | Kottapalli S.R.K., Hubli K., Chandrashekhara S., Jain G., Hubli S., Botla, et al. (2025). Foundation models for time series. arXiv preprint arXiv:2504.04011. DOI:10.48550/arXiv.2504.04011 |
| [50] | Yang S.D., Ali Z.A., Wong B.M. (2023). Fluid-gpt (fast learning to understand and investigate dynamics with a generative pre-trained transformer): Efficient predictions of particle trajectories and erosion. Ind. Eng. Chem. Res. 62(37):15278−15289. DOI:10.1021/acs.iecr.3c01639 |
| [51] | Jeong E., Oh S., Kim H., Park J., Bennis M., Kim S.L. (2018). Communication-efficient on-device machine learning: Federated distillation and augmentation under non-iid private data. arXiv preprint arXiv:1811.11479. DOI:10.48550/arXiv.1811.11479 |
| [52] | Barshandeh S., Khodadadi N., Abdollahzadeh B., Mohammadzadeh A., El-Kenawy E.S.M., Eid M.M., et al. (2026). Gray langurs optimizer: a multi-group bio-inspired optimization algorithm. Artif. Intell. Rev. DOI:10.1007/s10462-026-10987 |
| [53] | Ma X., Zhang P., Zhang S., Duan N., Hou Y., Zhou M., et al. (2019). A tensorized transformer for language modeling. Adv. Neural Inf. Process. Syst. 32. DOI:10.48550/arXiv.1906.02744 |
| [54] | Han K., Xiao A., Wu E., Guo J., Xu C., Wang Y. (2021). Transformer in transformer. Adv. Neural Inf. Process. Syst. 34:15908−15919. DOI:10.48550/arXiv.2103.04697 |
| [55] | Liao W., Wang S., Yang D., Yang Z., Fang J., Rehtanz C., et al. (2025). Timegpt in load forecasting: A large time series model perspective. Appl. Energy 379:124973. DOI:10.1016/j.apenergy.2025.124973 |
| [56] | Park Y.J., Germain F., Liu J., Wang Y., Koike-Akino T., Wichern G., et al. (2025). Probabilistic forecasting for building energy systems using time-series foundation models. arXiv preprint arXiv:2506.00630. DOI:10.48550/arXiv.2506.00630 |
| [57] | Zhang C., Lu J., Zhao Y. (2024). Generative pre-trained transformers (gpt)-based automated data mining for building energy management: Advantages, limitations and the future. Energy Built Environ. 5(1):143−169. DOI:10.1016/j.enbenv.2024.01.004 |
| [58] | Hinton G., Vinyals O., Dean J. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531. DOI:10.48550/arXiv.1503.02531 |
| [59] | Dai X., Chen R., Guan S., Li W.T., Yuen C. (2025). Buildinggym: An open-source toolbox for ai-based building energy management. In: Building Simulation, Springer, pp:1–19. DOI:10.1007/9464-024-00876 |
| [60] | Huang L., Qin J., Zhou Y., Zhu F., Liu L., Shao L. (2023). Normalization techniques in training dnns. IEEE Trans. Pattern Anal. Mach. Intell. 45(8):10173−10196. DOI:10.1109/TPAMI.2023.3250241 |
| [61] | Hota H., Handa R., Shrivas A.K. (2017). Time series data prediction using sliding window based rbf neural network. Int. J. Comput. Intell. Res. 13(5):1145−1156. DOI:10.37423/IJOCIR.2017.13.5.11 |
| [62] | Monfet D., Corsi M., Choinière D., Arkhipova E. (2014). Development of an energy prediction tool for commercial buildings using case-based reasoning. Energy Build. 81:152−160. DOI:10.1016/j.enbuild.2014.06.014 |
| [63] | Chicco D., Warrens M.J., Jurman G. (2021). The coefficient of determination r-squared is more informative than smape, mae, mape, mse and rmse. PeerJ Comput. Sci. 7:e623. DOI:10.7717/peerj-cs.623 |
| [64] | Willmott C.J., Matsuura K. (2005). Advantages of the mean absolute error (mae) over the root mean square error (rmse) in assessing average model performance. Clim. Res. 30(1):79−82. DOI:10.3354/cr0030079 |
| Wang Y., Dai X., Lin Z., et al. (2026). Distilling Temporal Knowledge: A GPT-BEF Based Framework for Building Energy Consumption Forecasting. Energy Use 2:100056. https://doi.org/10.59717/ipj.energy-use.2026.100056 |
To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.
(A) Overview of the Transformer encoder—decoder architecture
(A) General principle of knowledge distillation in deep neural networks
Overall pipeline of GPT-BEF for sparse-feature building load forecasting and cross-climate-zone trans- fer.
(A) U.S. climate zones and representative EPW weather files used in this study: The map shows the eight U.S. DOE Building America climate zones. For each zone we select a representative city and use its typical meteorological year (TMY) EPW file as the weather boundary condition in our EnergyPlus simulations. (B) Demonstration of a sliding-window time segmentation approach for building energy forecasting.
(A)Comparison of typical device energy consumption predictions by multiple traditional models and the proposed GPT- BEF against the ground truth using five input features. (B) Percentage error of typical device energy consumption predictions across samples using the GPT-BEF. (C) & (D) Comparison of prediction performance using different sequence windows in GPT-BEF model.
(A)Comparison of typical device energy consumption predictions by multiple traditional models and the proposed GPT- BEF against the ground truth over a two-day window using five input features. (B) Comparison of typical device energy consumption predictions by multiple traditional models and the proposed GPT- BEF against the ground truth over a two-day window using two input features.
(A) RMSE comparison among the proposed GPT-BEF model with five features, GPT-BEF+KD model with two features, and RF/XGBoost baselines using either five or two features. (B) Regime-based RMSE comparison for the 8A extreme-climate case. The analysis stratifies errors by load quantiles and rapid-ramp periods, highlighting where GPT-BEF, GPT-BEF+KD, RF, and XGBoost succeed or fail under rare operating regimes.
Typical energy prediction results of different models in the 3B → 3C cross-climate-zone transfer setting.
Typical energy prediction results of different models in the 3B → 3C cross-climate-zone transfer setting.