Article Contents
REVIEW   Open Access     Cite

Distilling Temporal Knowledge: A GPT-BEF Based Framework for Building Energy Consumption Forecasting

    Fund Project: This research is supported by AI Singapore under its project of Development of NetZero BEMS through AI-based HVAC System Control (AISG2-TC-2023-008-SGKR), and Chongqing Overseas Talent Recruitment Program (CSTB2025YCJH-KYXM0119).
More Information
  • Corresponding author: xileidai@outlook.com 
  • DownLoad: Full size image
    1. Building energy forecasting suffers from unavailable sensor data and hard-to-model long-range nonlinear time correlations; classic ML models rely on heavy feature engineering.

      A GPT-style Transformer teacher model (GPT-BEF) learns full multivariate time series to capture complex building energy temporal patterns.

      Knowledge distillation transfers teacher knowledge to a lightweight student model using merely two temperature input features.

      GPT-BEF+KD cuts RMSE by ~25% versus two-feature Random Forest, retaining strong short- and long-term temporal capture capacity.

      This work validates GPT-based distillation for sparse-sensor building energy prediction and supports follow-up energy management research.

  • With the rapid growth of smart-building applications, accurate forecasting of building energy consump- tion faces two critical challenges: some operational features (e.g., occupancy or equipment load) cannot be predicted accurately or may be entirely unavailable due to sensor malfunctions at prediction time, and com- plex nonlinear, long-range temporal dependencies are difficult to capture with traditional machine-learning models. Conventional approaches, such as Seasonal AutoRegressive Integrated Moving Average with eXoge- nous regressors (SARIMAX), Random Forest (RF), eXtreme Gradient Boosting (XGBoost), and Support Vector Machine (SVM), require extensive feature engineering, often assume linearity or stationarity, and show reduced accuracy when only sparse inputs are provided. To overcome these limitations, we introduce a Knowledge Distillation (KD) framework in which a GPT-style Transformer-based Building Energy Forecast- ing (GPT-BEF) teacher model is first trained on rich historical multivariate time series to capture complex temporal dependencies in building energy use. Its knowledge is then transferred into a reduced-feature GPT- BEF+KD student model that uses only two temperature features. The resulting GPT-BEF+KD model achieves approximately 25% lower Root Mean Squared Error (RMSE) than a two-feature RF baseline, sug- gesting its potential for lightweight forecasting under limited sensor availability. The analysis demonstrates that the student model effectively inherits part of the teacher’s ability to balance long-term and short-term temporal signals. By combining archive-driven full-feature learning with feature-constrained inference, this work provides an initial demonstration of GPT-style knowledge distillation for sparse-feature building en- ergy prediction and lays the foundation for future advances in multimodal fusion, closed-loop control, and uncertainty-aware energy management.
  • 加载中
  • [1] Jia R., Jin B., Jin M., Zhou Y., Konstantakopoulos I.C., Zou H., et al. (2018). Design automation for smart building systems. Proc. IEEE 106(9):1680−1699. DOI:10.1109/JPROC.2018.2860343

    View in Article CrossRef Google Scholar

    [2] Buckman A.H., Mayfield M., Beck S.B.M. (2014). What is a smart building. Smart Sustain. Built Environ. 3(2):92−109. DOI:10.1080/20416622.2014.911037

    View in Article CrossRef Google Scholar

    [3] Zhang X.P., Cheng X.M. (2009). Energy consumption, carbon emissions, and economic growth in china. Ecol. Econ. 68(10):2706−2712. DOI:10.1016/j.ecolecon.2009.05.011

    View in Article CrossRef Google Scholar

    [4] Bansal M.A., Sharma D.R., Kathuria D.M. (2022). A systematic review on data scarcity problem in deep learning. ACM Comput. Surv. 54(10s):1−29. DOI:10.1145/3563425

    View in Article CrossRef Google Scholar

    [5] Jiang R., Zeng S., Song Q., Wu Z. (2022). Deep-chain echo state network with explainable temporal dependence for complex building energy prediction. IEEE Trans. Ind. Inform. 19(1):426−435. DOI:10.1109/TII.2022.3170672

    View in Article CrossRef Google Scholar

    [6] Garza A., Challu C., Mergenthaler-Canseco M. (2023). Timegpt-1. arXiv preprint arXiv:2310.03589. DOI:10.48550/arXiv.2310.03589

    View in Article Google Scholar

    [7] Liu Y., Zhang H., Li C., Huang X., Wang J., Long M. (2024). Timer: Generative pre-trained transformers are large time series models. arXiv preprint arXiv:2402.02368. DOI:10.48550/arXiv.2402.02368

    View in Article Google Scholar

    [8] Mirchandani S., Xia F., Florence P., Ichter B., Driess D., Arenas M.G., et al. (2023). Large language models as general pattern machines. arXiv preprint arXiv:2307.04721. DOI:10.48550/arXiv.2307.04721

    View in Article Google Scholar

    [9] Kraus M. (1987). Energy forecasting: The epistemological context. Futures 19(3):254−275. DOI:10.1016/0016-3287(87)90018-7

    View in Article CrossRef Google Scholar

    [10] Dai X., Cheng S., Chong A. (2023). Deciphering optimal mixed-mode ventilation in the tropics using reinforcement learning with explainable artificial intelligence. Energy Build. 278:112629. DOI:10.1016/j.enbuild.2023.112629

    View in Article CrossRef Google Scholar

    [11] Pérez-Lombard L., Ortiz J., Pout C. (2008). A review on buildings energy consumption information. Energy Build. 40(3):394−398. DOI:10.1016/j.enbuild.2007.03.007

    View in Article CrossRef Google Scholar

    [12] Wei N., Li C., Peng X., Zeng F., Lu X. (2019). Conventional models and artificial intelligence-based models for energy consumption forecasting. J. Pet. Sci. Eng. 181:106187. DOI:10.1016/j.petrol.2019.106187

    View in Article CrossRef Google Scholar

    [13] Ahmed N.K., Atiya A.F., Gayar N.E., El-Shishiny H. (2010). An empirical comparison of machine learning models for time series forecasting. Econometr. Rev. 29(5-6):594−621. DOI:10.1080/07474938.2010.481027

    View in Article CrossRef Google Scholar

    [14] Hammad M.A., Jereb B., Rosi B., Dragan D. (2020). Methods and models for electric load forecasting. Logist. Supply Chain Sustain. Glob. Chall. 11(1):51−76. DOI:10.3390/logistics11010051

    View in Article CrossRef Google Scholar

    [15] Vagropoulos S.I., Chouliaras G., Kardakos E.G., Simoglou C.K., Bakirtzis A.G. (2016). Comparison of sarimax, sarima, modified sarima and ann-based models for short-term pv generation forecasting. In: 2016 IEEE International Energy Conference (ENERGYCON), pp:1–6. DOI:10.1109/ENERGYCON.2016.7514066

    View in Article Google Scholar

    [16] El-Kenawy E.S.M., Ibrahim A., Alhussan A.A., Khafaga D.S., Ahmed A., Eid M., et al. (2026). Smart city electricity load forecasting using greylag goose optimization-enhanced time series analysis. Arab. J. Sci. Eng. 51(6):8359−8377. DOI:10.1007/s13369-025-09421

    View in Article CrossRef Google Scholar

    [17] Gardner E.S. Jr. (1985). Exponential smoothing: The state of the art. J. Forecast. 4(1):1−28. DOI:10.1002/for.39800401

    View in Article CrossRef Google Scholar

    [18] Chatfield C., Yar M. (1988). Holt-winters forecasting: some practical issues. Statistician 37(2):129−140. DOI:10.2307/2348673

    View in Article CrossRef Google Scholar

    [19] Dai X., Liu J., Zhang X. (2020). A review of studies applying machine learning models to predict occupancy and window-opening behaviours in smart buildings. Energy Build. 223:110159. DOI:10.1016/j.enbuild.2020.110159

    View in Article CrossRef Google Scholar

    [20] Dhiman H.S., Deb D., Guerrero J.M. (2019). Hybrid machine intelligent svr variants for wind forecasting and ramp events. Renew. Sustain. Energy Rev. 108:369−379. DOI:10.1016/j.rser.2019.03.032

    View in Article CrossRef Google Scholar

    [21] Wang Z., Wang Y., Zeng R., Srinivasan R.S., Ahrentzen S. (2018). Random forest based hourly building energy prediction. Energy Build. 171:11−25. DOI:10.1016/j.enbuild.2018.04.032

    View in Article CrossRef Google Scholar

    [22] Chen T., Guestrin C. (2016). Xgboost: A scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp:785–794. DOI:10.1145/2939672.2939785

    View in Article Google Scholar

    [23] Chen C., Zhang Q., Ma Q., Yu B. (2019). Lightgbm-ppi: Predicting protein-protein interactions through multi-information fusion. Chemom. Intell. Lab. Syst. 191:54−64. DOI:10.1016/j.chemolab.2019.05.007

    View in Article CrossRef Google Scholar

    [24] Breiman L. (2001). Random forests. Mach. Learn. 45:5−32. DOI:10.1023/A:1010933404324

    View in Article CrossRef Google Scholar

    [25] Kremic E., Subasi A. (2016). Performance of random forest and svm in face recognition. Int. Arab J. Inf. Technol. 13(2):287−293. DOI:10.34069/iajit.2016.13.2.7

    View in Article CrossRef Google Scholar

    [26] Alhussan A.A., El-Kenawy E.S.M., Eid M.M., Khodadadi N. (2026). Hybrid al-biruni and puma optimization (berpo) for boosting the classification of quality-of-service (qos) in 5g networks. J. Netw. Comput. Appl. :104462. DOI:10.1016/j.jnca.2026.104462

    View in Article Google Scholar

    [27] Lim B., Arık S.Ö., Loeff N., Pfister T. (2021). Temporal fusion transformers for interpretable multi-horizon time series forecasting. Int. J. Forecast. 37(4):1748−1764. DOI:10.1016/j.ijforecast.2021.03.012

    View in Article CrossRef Google Scholar

    [28] Zhou H., Zhang S., Peng J., Zhang S., Li J., Xiong H., et al. (2021). Informer: Beyond efficient transformer for long sequence time-series forecasting. In: AAAI Conference Proceedings, vol.35, pp:11106–11115. DOI:10.1609/aaai.v35i12.17325

    View in Article Google Scholar

    [29] Wu H., Xu J., Wang J., Long M. (2021). Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Adv. Neural Inf. Process. Syst. 34:22419−22430. DOI:10.48550/arXiv.2106.13008

    View in Article CrossRef Google Scholar

    [30] Zhou T., Ma Z., Wen Q., Wang X., Sun L., Jin R. (2022). Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting. In: International Conference on Machine Learning, PMLR, pp:27268–27286. DOI:10.48550/arXiv.2201.12740

    View in Article Google Scholar

    [31] Radwan M., Ibrahim A., Abdelsalam M.M., Alhussan A.A., Mattar E.A., El-Kenawy E.S.M., et al. (2026). Optimizing solar and wind forecasting with ihow optimization algorithm and multi-scale attention networks. Sci. Rep. DOI:10.1038/s41598-026-84671.

    View in Article Google Scholar

    [32] Dai T.Y., Niyogi D., Nagy Z. (2025). Citytft: A temporal fusion transformer-based surrogate model for urban building energy modeling. Appl. Energy 389:125712. DOI:10.1016/j.apenergy.2025.125712

    View in Article CrossRef Google Scholar

    [33] Gu Y., Jazizadeh F., Wang X. (2025). Toward large energy models: A comparative study of transformers’ efficacy for energy forecasting. Appl. Energy 384:125358. DOI:10.1016/j.apenergy.2025.125358

    View in Article CrossRef Google Scholar

    [34] Xing Z., Pan Y., Yang Y., Yuan X., Liang Y., Huang Z., et al. (2024). Transfer learning integrating similarity analysis for short-term and long-term building energy consumption prediction. Appl. Energy 365:123276. DOI:10.1016/j.apenergy.2024.123276

    View in Article CrossRef Google Scholar

    [35] Liang X., Chen S., Zhu X., Jin X., Du Z. (2023). Domain knowledge decomposition of building energy consumption and a hybrid data-driven model for 24-h ahead predictions. Appl. Energy 344:121244. DOI:10.1016/j.apenergy.2023.121244

    View in Article CrossRef Google Scholar

    [36] Hu Z., Gao Y., Ji S., Mae M., Imaizumi T. (2024). Improved multistep ahead photovoltaic power prediction model based on lstm and self-attention with weather forecast data. Appl. Energy 359:122709. DOI:10.1016/j.apenergy.2024.122709

    View in Article CrossRef Google Scholar

    [37] Karijadi I., Chou S.Y. (2022). A hybrid rf-lstm based on ceemdan for improving the accuracy of building energy consumption prediction. Energy Build. 259:111908. DOI:10.1016/j.enbuild.2022.111908

    View in Article CrossRef Google Scholar

    [38] El-Kenawy E.S.M., Khodadadi N., Mirjalili S., Zaki A.M., Ibrahim A., Alhussan A.A., et al. (2026). Glider snake optimizer (gso): a nature-inspired metaheuristic algorithm for global and engineering optimization problems. Artif. Intell. Rev. 59:91. DOI:10.1007/s10462-025-10876

    View in Article CrossRef Google Scholar

    [39] Jang J., Han J., Leigh S.B. (2022). Prediction of heating energy consumption with operation pattern variables for non-residential buildings using lstm networks. Energy Build. 255:111647. DOI:10.1016/j.enbuild.2022.111647

    View in Article CrossRef Google Scholar

    [40] Liu J., Zhang Y., Wen K., Ding Y. (2025). Analysis of different neural network models based on variational mode decomposition and dung beetle optimizer for the prediction of air-conditioning energy consumption in multifunctional complex large public buildings. Energy Build. 334:115518. DOI:10.1016/j.enbuild.2025.115518

    View in Article CrossRef Google Scholar

    [41] Brants T., Popat A., Xu P., Och F.J., Dean J. (2007). Large language models in machine translation. In: EMNLP-CoNLL Conference Proceedings, pp:858–867. DOI:10.3115/1079540.

    View in Article Google Scholar

    [42] Dathathri S., Madotto A., Lan J., Hung J., Frank E., Molino P., et al. (2019). Plug and play language models: A simple approach to controlled text generation. arXiv preprint arXiv:1912.02164. DOI:10.48550/arXiv.1912.02164.

    View in Article Google Scholar

    [43] Izadi M., Katzy J., Van Dam T., Otten M., Popescu R.M., Van Deursen A. (2024). Language models for code completion. In: IEEE/ACM 46th International Conference on Software Engineering, pp:1–13. DOI:10.1145/3597643.3597682.

    View in Article Google Scholar

    [44] Yu X., Chen Z., Ling Y., Dong S., Liu Z., Lu Y. (2023). Temporal data meets llm–explainable financial time series forecasting. arXiv preprint arXiv:2306.11025. DOI:10.48550/arXiv.2306.11025

    View in Article Google Scholar

    [45] de Zarzà I., de Curtò J., Roig G., Calafate C.T. (2023). Llm multimodal traffic accident forecasting. Sensors 23(22):9225. DOI:10.3390/s23229225

    View in Article CrossRef Google Scholar

    [46] Thirunavukarasu A.J., Ting D.S.J., Elangovan K., Gutierrez L., Tan T.F., Ting D.S.W. (2023). Large language models in medicine. Nat. Med. 29(8):1930−1940. DOI:10.1038/s41591-023-02488

    View in Article CrossRef Google Scholar

    [47] Jin M., Wang S., Ma L., Chu Z., Zhang J.Y., Shi X., et al. (2023). Time-llm: Time series forecasting by reprogramming large language models. arXiv preprint arXiv:2310.01728. DOI:10.48550/arXiv.2310.01728

    View in Article Google Scholar

    [48] Liang Y., Wen H., Nie Y., Jiang Y., Jin M., Song D., et al. (2024). Foundation models for time series analysis: A tutorial and survey. In: 30th ACM SIGKDD Conference Proceedings, pp:6555–6565. DOI:10.1145/3637462.3637560

    View in Article Google Scholar

    [49] Kottapalli S.R.K., Hubli K., Chandrashekhara S., Jain G., Hubli S., Botla, et al. (2025). Foundation models for time series. arXiv preprint arXiv:2504.04011. DOI:10.48550/arXiv.2504.04011

    View in Article Google Scholar

    [50] Yang S.D., Ali Z.A., Wong B.M. (2023). Fluid-gpt (fast learning to understand and investigate dynamics with a generative pre-trained transformer): Efficient predictions of particle trajectories and erosion. Ind. Eng. Chem. Res. 62(37):15278−15289. DOI:10.1021/acs.iecr.3c01639

    View in Article CrossRef Google Scholar

    [51] Jeong E., Oh S., Kim H., Park J., Bennis M., Kim S.L. (2018). Communication-efficient on-device machine learning: Federated distillation and augmentation under non-iid private data. arXiv preprint arXiv:1811.11479. DOI:10.48550/arXiv.1811.11479

    View in Article Google Scholar

    [52] Barshandeh S., Khodadadi N., Abdollahzadeh B., Mohammadzadeh A., El-Kenawy E.S.M., Eid M.M., et al. (2026). Gray langurs optimizer: a multi-group bio-inspired optimization algorithm. Artif. Intell. Rev. DOI:10.1007/s10462-026-10987

    View in Article Google Scholar

    [53] Ma X., Zhang P., Zhang S., Duan N., Hou Y., Zhou M., et al. (2019). A tensorized transformer for language modeling. Adv. Neural Inf. Process. Syst. 32. DOI:10.48550/arXiv.1906.02744

    View in Article Google Scholar

    [54] Han K., Xiao A., Wu E., Guo J., Xu C., Wang Y. (2021). Transformer in transformer. Adv. Neural Inf. Process. Syst. 34:15908−15919. DOI:10.48550/arXiv.2103.04697

    View in Article CrossRef Google Scholar

    [55] Liao W., Wang S., Yang D., Yang Z., Fang J., Rehtanz C., et al. (2025). Timegpt in load forecasting: A large time series model perspective. Appl. Energy 379:124973. DOI:10.1016/j.apenergy.2025.124973

    View in Article CrossRef Google Scholar

    [56] Park Y.J., Germain F., Liu J., Wang Y., Koike-Akino T., Wichern G., et al. (2025). Probabilistic forecasting for building energy systems using time-series foundation models. arXiv preprint arXiv:2506.00630. DOI:10.48550/arXiv.2506.00630

    View in Article Google Scholar

    [57] Zhang C., Lu J., Zhao Y. (2024). Generative pre-trained transformers (gpt)-based automated data mining for building energy management: Advantages, limitations and the future. Energy Built Environ. 5(1):143−169. DOI:10.1016/j.enbenv.2024.01.004

    View in Article CrossRef Google Scholar

    [58] Hinton G., Vinyals O., Dean J. (2015). Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531. DOI:10.48550/arXiv.1503.02531

    View in Article Google Scholar

    [59] Dai X., Chen R., Guan S., Li W.T., Yuen C. (2025). Buildinggym: An open-source toolbox for ai-based building energy management. In: Building Simulation, Springer, pp:1–19. DOI:10.1007/9464-024-00876

    View in Article Google Scholar

    [60] Huang L., Qin J., Zhou Y., Zhu F., Liu L., Shao L. (2023). Normalization techniques in training dnns. IEEE Trans. Pattern Anal. Mach. Intell. 45(8):10173−10196. DOI:10.1109/TPAMI.2023.3250241

    View in Article CrossRef Google Scholar

    [61] Hota H., Handa R., Shrivas A.K. (2017). Time series data prediction using sliding window based rbf neural network. Int. J. Comput. Intell. Res. 13(5):1145−1156. DOI:10.37423/IJOCIR.2017.13.5.11

    View in Article CrossRef Google Scholar

    [62] Monfet D., Corsi M., Choinière D., Arkhipova E. (2014). Development of an energy prediction tool for commercial buildings using case-based reasoning. Energy Build. 81:152−160. DOI:10.1016/j.enbuild.2014.06.014

    View in Article CrossRef Google Scholar

    [63] Chicco D., Warrens M.J., Jurman G. (2021). The coefficient of determination r-squared is more informative than smape, mae, mape, mse and rmse. PeerJ Comput. Sci. 7:e623. DOI:10.7717/peerj-cs.623

    View in Article CrossRef Google Scholar

    [64] Willmott C.J., Matsuura K. (2005). Advantages of the mean absolute error (mae) over the root mean square error (rmse) in assessing average model performance. Clim. Res. 30(1):79−82. DOI:10.3354/cr0030079

    View in Article CrossRef Google Scholar

  • Cite this article:

    Wang Y., Dai X., Lin Z., et al. (2026). Distilling Temporal Knowledge: A GPT-BEF Based Framework for Building Energy Consumption Forecasting. Energy Use 2:100056. https://doi.org/10.59717/ipj.energy-use.2026.100056
    Wang Y., Dai X., Lin Z., et al. (2026). Distilling Temporal Knowledge: A GPT-BEF Based Framework for Building Energy Consumption Forecasting. Energy Use 2:100056. https://doi.org/10.59717/ipj.energy-use.2026.100056

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(9)    

Supplementary Information

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(391) PDF downloads(129)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint