| [1] | Bubeck S., Chandrasekaran V., Eldan R., et al. (2023). Sparks of artificial general intelligence: Early experiments with GPT-4. Preprint at arXiv. DOI:10.48550/arXiv.2303.12712. |
| [2] | Ge Y., Hua W., Mei K., et al. (2023). OpenAGI: When LLM meets domain experts. Preprint at arXiv. DOI:10.48550/arXiv.2304.04370. |
| [3] | Hendrycks D., Burns C., Kadavath S., et al. (2021). Measuring mathematical problem solving with the MATH dataset. Preprint at arXiv. DOI:10.48550/arXiv.2103.03874. |
| [4] | Cobbe K., Kosaraju V., Bavarian M., et al. (2021). Training verifiers to solve math word problems. Preprint at arXiv. DOI:10.48550/arXiv.2110.14168. |
| [5] | Hendrycks D., Burns C., Basart S., et al. (2020). Measuring massive multitask language understanding. Preprint at arXiv. DOI:10.48550/arXiv.2009.03300. |
| [6] | Zhong W., Cui R., Guo Y., et al. (2023). AGIEval: A human-centric benchmark for evaluating foundation models. Preprint at arXiv. DOI:10.48550/arXiv.2304.06364. |
| [7] | Rein D., Hou B. L., Stickland A. C., et al. (2023). GPQA: A graduate-level Google-proof Q&A benchmark. Preprint at arXiv. DOI:10.48550/arXiv.2311.12022. |
| [8] | Du X., Yao Y., Ma K., et al. (2025). SuperGPQA: Scaling LLM evaluation across 285 graduate disciplines. Preprint at arXiv. DOI:10.48550/arXiv.2502.14739. |
| [9] | Lu P., Mishra S., Xia T., et al. (2022). Learn to explain: Multimodal reasoning via thought chains for science question answering. Preprint at arXiv. DOI:10.48550/arXiv.2209.09513. |
| [10] | Wang Y., Ma X., Zhang G., et al. (2024). MMLU-Pro: A more robust and challenging multi-task language understanding benchmark. Preprint at arXiv. DOI:10.48550/arXiv.2406.01574. |
| [11] | Yue X., Ni Y., Zhang K., et al. (2023). MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI. Preprint at arXiv. DOI:10.48550/arXiv.2311.16502. |
| [12] | Phan L., Gatti A., Han Z., et al. (2025). Humanity’s last exam. Preprint at arXiv. DOI:10.48550/arXiv.2501.14249. |
| [13] | Zhao Q., Huang Y., Lv T., et al. (2024). MMLU-CF: A contamination-free multi-task language understanding benchmark. Preprint at arXiv. DOI:10.48550/arXiv.2412.15194. |
| [14] | Sullivan G. M. and Artino A. R. (2013). Analyzing and interpreting data from Likert-type scales. J. Grad. Med. Educ. 5:541−542. DOI:10.4300/JGME-5-4-18 |
| [15] | Hoyos-Osorio J. K. and Sanchez-Giraldo L. G. (2023). The representation Jensen–Shannon divergence. Preprint at arXiv. DOI:10.48550/arXiv.2305.16446. |
| [16] | Shlens J. (2014). Notes on Kullback–Leibler divergence and likelihood. Preprint at arXiv. DOI:10.48550/arXiv.1404.2000. |
| [17] | Jiang G., Chen Y., Wang Z., et al. (2024). A deep learning-based Bayesian framework for high-resolution calibration of building energy models. Energy Build. 323:114755. DOI:10.1016/j.enbuild.2024.114755 |
| [18] | Langrené N. and Warin X. (2021). Fast multivariate empirical cumulative distribution function with connection to kernel density estimation. Comput. Stat. Data Anal. 157:107267. DOI:10.1016/j.csda.2021.107267 |
| [19] | Yang A., Li A., Yang B., et al. (2025). Qwen3 technical report. Preprint at arXiv. DOI:10.48550/arXiv.2505.09388. |
| [20] | Achiam J., Adler S., Agarwal S., et al. (2023). GPT-4 technical report. Preprint at arXiv. DOI:10.48550/arXiv.2303.08774. |
| [21] | Comanici G., Bieber E., Schaekermann M., et al. (2025). Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next-generation agentic capabilities. Preprint at arXiv. DOI:10.48550/arXiv.2507.06261. |
| [22] | Liu A., Feng B., Xue B., et al. (2024). DeepSeek-V3 technical report. Preprint at arXiv. DOI:10.48550/arXiv.2412.19437. |
| [23] | Jiang G., Ma Z., Zhang L. and Chen J. (2024). EPlus-LLM: A large language model-based computing platform for automated building energy modeling. Appl. Energy 367:123431. DOI:10.1016/j.apenergy.2024.123431 |
| [24] | Kaplan J., McCandlish S., Henighan T., et al. (2020). Scaling laws for neural language models. Preprint at arXiv. DOI:10.48550/arXiv.2001.08361. |
| [25] | Zhang Q., Hu C., Upasani S., et al. (2025). Agentic context engineering: Evolving contexts for self-improving language models. Preprint at arXiv. DOI:10.48550/arXiv.2510.04618. |
| [26] | Jiang G., Ma Z., Zhang L. and Chen J. (2025). Prompt engineering to inform large language model in automated building energy modeling. Energy 316:134548. DOI:10.1016/j.energy.2025.134548 |
| [27] | Jiang G. and Chen J. (2025). Efficient fine-tuning of large language models for automated building energy modeling in complex cases. Autom. Constr. 175:106223. DOI:10.1016/j.autcon.2025.106223 |
| Jiang G., Chen K., Liu J., et al. (2025). Benchmarking Knowledge and Capability of Large Language Models in Building Science Domain. Energy Use 1:100026. https://doi.org/10.59717/ipj.energy-use.2025.100026 |
To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.
Distribution of the benchmark
Sampling strategy and coverage across topics and question types
Statistical analysis of AI-human comparison
Bayesian posterior distributions for capability deviations between AI-Judgers and human expert baseline
Result accuracy of LLMs on the benchmark
Capability scores of LLMs across four aspects (1-5 scale)
Summary of LLM performance
Capability scores of LLMs under prompt strategies