Article Contents
ARTICLE   Open Access     Cite

Benchmarking Knowledge and Capability of Large Language Models in Building Science Domain

More Information
  • DownLoad: Full size image
  • Large language models (LLMs) are increasingly adopted across scientific and engineering fields. However, applying general-purpose LLMs to specialized engineering domains imposes stringent requirements for structured knowledge, rigorous reasoning, and technical precision. Thus, the suitability of current general-purpose LLMs for practical applications in engineering domains remains questionable. To understand the mastery level of LLMs in the building science domain as one broad but specific engineering domain, in this paper, we perform a comprehensive benchmark analysis (with benchmark dataset of 1,487 questions) to evaluate abilities of 15 state-of-the-art (SOTA) LLMs across 12 core subject topics in the building science domain. To enable scalable and robust evaluation, we propose and validate an AI-Judger for assessment across five dimensions of abilities, i.e., knowledge and concept, logic and consistency, clarity of expression, and reflection and exploratory. Overall, SOTA general-purposes LLMs achieve only ~50% accuracy on average in answering different types of questions. The capabilities of LLMs decrease progressively from linguistic expression and factual knowledge to logical reasoning, then reflection and exploratory thinking. For different tasks, LLMs exhibit notably low accuracy on calculation (~13%), short-answer (~23%), and cloze tasks (~30%), contrast to stronger performance on single-choice (74%) and multiple-choice questions (63%). Finally, pronounced variance of LLM performance exists across topics, with relatively low accuracy on physics fundamental and HVAC&R-related questions (median of 20%-40%) compared to ~80% for building standards and codes. These identified gaps highlight the limitations of general-purpose LLMs in engineering contexts, clearly pointing to the necessity of developing domain-specific LLMs tailored for engineering applications.
  • 加载中
  • [1] Bubeck S., Chandrasekaran V., Eldan R., et al. (2023). Sparks of artificial general intelligence: Early experiments with GPT-4. Preprint at arXiv. DOI:10.48550/arXiv.2303.12712.

    View in Article Google Scholar

    [2] Ge Y., Hua W., Mei K., et al. (2023). OpenAGI: When LLM meets domain experts. Preprint at arXiv. DOI:10.48550/arXiv.2304.04370.

    View in Article Google Scholar

    [3] Hendrycks D., Burns C., Kadavath S., et al. (2021). Measuring mathematical problem solving with the MATH dataset. Preprint at arXiv. DOI:10.48550/arXiv.2103.03874.

    View in Article Google Scholar

    [4] Cobbe K., Kosaraju V., Bavarian M., et al. (2021). Training verifiers to solve math word problems. Preprint at arXiv. DOI:10.48550/arXiv.2110.14168.

    View in Article Google Scholar

    [5] Hendrycks D., Burns C., Basart S., et al. (2020). Measuring massive multitask language understanding. Preprint at arXiv. DOI:10.48550/arXiv.2009.03300.

    View in Article Google Scholar

    [6] Zhong W., Cui R., Guo Y., et al. (2023). AGIEval: A human-centric benchmark for evaluating foundation models. Preprint at arXiv. DOI:10.48550/arXiv.2304.06364.

    View in Article Google Scholar

    [7] Rein D., Hou B. L., Stickland A. C., et al. (2023). GPQA: A graduate-level Google-proof Q&A benchmark. Preprint at arXiv. DOI:10.48550/arXiv.2311.12022.

    View in Article Google Scholar

    [8] Du X., Yao Y., Ma K., et al. (2025). SuperGPQA: Scaling LLM evaluation across 285 graduate disciplines. Preprint at arXiv. DOI:10.48550/arXiv.2502.14739.

    View in Article Google Scholar

    [9] Lu P., Mishra S., Xia T., et al. (2022). Learn to explain: Multimodal reasoning via thought chains for science question answering. Preprint at arXiv. DOI:10.48550/arXiv.2209.09513.

    View in Article Google Scholar

    [10] Wang Y., Ma X., Zhang G., et al. (2024). MMLU-Pro: A more robust and challenging multi-task language understanding benchmark. Preprint at arXiv. DOI:10.48550/arXiv.2406.01574.

    View in Article Google Scholar

    [11] Yue X., Ni Y., Zhang K., et al. (2023). MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. Preprint at arXiv. DOI:10.48550/arXiv.2311.16502.

    View in Article Google Scholar

    [12] Phan L., Gatti A., Han Z., et al. (2025). Humanity’s last exam. Preprint at arXiv. DOI:10.48550/arXiv.2501.14249.

    View in Article Google Scholar

    [13] Zhao Q., Huang Y., Lv T., et al. (2024). MMLU-CF: A contamination-free multi-task language understanding benchmark. Preprint at arXiv. DOI:10.48550/arXiv.2412.15194.

    View in Article Google Scholar

    [14] Sullivan G. M. and Artino A. R. (2013). Analyzing and interpreting data from Likert-type scales. J. Grad. Med. Educ. 5:541−542. DOI:10.4300/JGME-5-4-18

    View in Article CrossRef Google Scholar

    [15] Hoyos-Osorio J. K. and Sanchez-Giraldo L. G. (2023). The representation Jensen–Shannon divergence. Preprint at arXiv. DOI:10.48550/arXiv.2305.16446.

    View in Article Google Scholar

    [16] Shlens J. (2014). Notes on Kullback–Leibler divergence and likelihood. Preprint at arXiv. DOI:10.48550/arXiv.1404.2000.

    View in Article Google Scholar

    [17] Jiang G., Chen Y., Wang Z., et al. (2024). A deep learning-based Bayesian framework for high-resolution calibration of building energy models. Energy Build. 323:114755. DOI:10.1016/j.enbuild.2024.114755

    View in Article CrossRef Google Scholar

    [18] Langrené N. and Warin X. (2021). Fast multivariate empirical cumulative distribution function with connection to kernel density estimation. Comput. Stat. Data Anal. 157:107267. DOI:10.1016/j.csda.2021.107267

    View in Article CrossRef Google Scholar

    [19] Yang A., Li A., Yang B., et al. (2025). Qwen3 technical report. Preprint at arXiv. DOI:10.48550/arXiv.2505.09388.

    View in Article Google Scholar

    [20] Achiam J., Adler S., Agarwal S., et al. (2023). GPT-4 technical report. Preprint at arXiv. DOI:10.48550/arXiv.2303.08774.

    View in Article Google Scholar

    [21] Comanici G., Bieber E., Schaekermann M., et al. (2025). Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next-generation agentic capabilities. Preprint at arXiv. DOI:10.48550/arXiv.2507.06261.

    View in Article Google Scholar

    [22] Liu A., Feng B., Xue B., et al. (2024). DeepSeek-V3 technical report. Preprint at arXiv. DOI:10.48550/arXiv.2412.19437.

    View in Article Google Scholar

    [23] Jiang G., Ma Z., Zhang L. and Chen J. (2024). EPlus-LLM: A large language model-based computing platform for automated building energy modeling. Appl. Energy 367:123431. DOI:10.1016/j.apenergy.2024.123431

    View in Article CrossRef Google Scholar

    [24] Kaplan J., McCandlish S., Henighan T., et al. (2020). Scaling laws for neural language models. Preprint at arXiv. DOI:10.48550/arXiv.2001.08361.

    View in Article Google Scholar

    [25] Zhang Q., Hu C., Upasani S., et al. (2025). Agentic context engineering: Evolving contexts for self-improving language models. Preprint at arXiv. DOI:10.48550/arXiv.2510.04618.

    View in Article Google Scholar

    [26] Jiang G., Ma Z., Zhang L. and Chen J. (2025). Prompt engineering to inform large language model in automated building energy modeling. Energy 316:134548. DOI:10.1016/j.energy.2025.134548

    View in Article CrossRef Google Scholar

    [27] Jiang G. and Chen J. (2025). Efficient fine-tuning of large language models for automated building energy modeling in complex cases. Autom. Constr. 175:106223. DOI:10.1016/j.autcon.2025.106223

    View in Article CrossRef Google Scholar

  • Cite this article:

    Jiang G., Chen K., Liu J., et al. (2025). Benchmarking Knowledge and Capability of Large Language Models in Building Science Domain. Energy Use 1:100026. https://doi.org/10.59717/ipj.energy-use.2025.100026
    Jiang G., Chen K., Liu J., et al. (2025). Benchmarking Knowledge and Capability of Large Language Models in Building Science Domain. Energy Use 1:100026. https://doi.org/10.59717/ipj.energy-use.2025.100026

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(8)    

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(4199) PDF downloads(951)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint