<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">JWA</journal-id><journal-title-group><journal-title>Journal of World Architecture</journal-title></journal-title-group><issn>2208-3480</issn><eissn>2208-3499</eissn><publisher><publisher-name>Bio-Byword Scientific Publishing Pty. Ltd.</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.26689/JWA.v10i4.15268</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>Evaluating Large Language Models in Green Building Design Decisions: A Case-Based Analysis of Energy, Water, and Material Trade-Offs</title><url>https://artdesignp.com/journal/JWA/10/4/10.26689/JWA.v10i4.15268</url><author>AL-qawasReem,ShaoDan</author><pub-date pub-type="publication-year"><year>2026</year></pub-date><volume>10</volume><issue>4</issue><history><date date-type="pub"><published-time>2026-08-31</published-time></date></history><abstract>Background: AI has shown potential across technical professions, yet its application to green building design&amp;mdash;navigating trade-offs between energy efficiency, water conservation, and materials lifecycle assessment (LCA)&amp;mdash;remains unevaluated. No prior study has benchmarked frontier large language models (LLMs) using a professor-validated instrument combining multiple-choice question (MCQ) and True/False (T/F) items probing context-dependent sustainability reasoning. Objective: To comparatively benchmark four frontier LLMs&amp;mdash;Claude, ChatGPT-5.2, DeepSeek-R1, and Gemini 3.1&amp;mdash;on a validated 80-item instrument (40 MCQs and 40 True/False statements) derived from 20 green building design scenarios spanning materials, energy, water, and cross-domain lifecycle trade-off reasoning. Methods: 80 items from 20 original scenarios were validated by two professors. MCQs required selection of the optimal design alternative under competing sustainability constraints; True/False items used context-dependent partial truths (Statements 1&amp;ndash;20: False via contextual override; 21&amp;ndash;40: True) to evaluate resistance to overgeneralization. Items were grounded in LEED v4.1, ASHRAE 90.1, and ISO 14044. Each model was evaluated under identical zero-temperature conditions using binary scoring. Pearson Chi-square tests were applied for inter-model and inter-format comparisons. Results: Claude achieved the highest overall accuracy (95.00%; 95% CI: 90.22&amp;ndash;99.78%), followed by DeepSeek-R1 (93.75%), ChatGPT-5.2 (88.75%), and Gemini 3.1 (83.75%). No statistically significant difference across chatbots was found (x2 = 7.251, df = 3, P = 0.064). No significant MCQ-vs-T/F differential was observed for any model (all P &amp;gt; 0.05); pooled T/F accuracy (91.3%) marginally exceeded pooled MCQ accuracy (89.4%). Conclusions: This is the first study to benchmark multiple frontier LLMs in green building design using a purpose-built, professor-validated instrument evaluating both design-alternative selection and resistance to absolute sustainability heuristics. Consistent performance declines in cross-domain lifecycle scenarios reveal residual multi-framework integration limitations across all models. Findings support LLM integration into green building education and decision-support, while underscoring the need for expert oversight in complex design contexts.</abstract><keywords>Large language models, Green building design, Sustainable architecture, Benchmark evaluation, Energy efficiency, Water conservation, LCA, LEED, Claude, ChatGPT, DeepSeek, Gemini</keywords></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>[1] Biswas SS, 2023, Role of ChatGPT in public health. Annals of Biomedical Engineering, 2023(51): 868&amp;ndash;869.
[2] Ali SR, Dobbs TD, Hutchings HA, et al., 2023, Using ChatGPT to Write Patient Clinic Letters. The Lancet Digital Health, 2023(5): e179&amp;ndash;181.
[3] Wang Y, Liang L, Li R, et al., 2024, Comparison of ChatGPT, Claude, and Bard in Support of Technical Decision-making. Journal of Multidisciplinary Healthcare, 2024(17): 3917&amp;ndash;3929.
[4] DeepSeek-AI, 2025, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv: 2501.12948.
[5] Wu J, Jiang M, Fan J, et al., 2025, Arch-Eval Benchmark for Assessing Chinese Architectural Domain Knowledge in LLMs. Scientific Reports, 2025(15): 1&amp;ndash;12.
[6] He Q, 2024, How Well Do LLMs Perform in Green Building Assessment? iCCPMCE-2024.
[7] Nguyen HC, Dang HP, Nguyen TL, et al., 2025, Accuracy of Latest LLMs in Answering MCQs in Technical Domains. PLoS ONE, 2025(20): e0317423.
[8] U.S. Green Building Council, 2021, LEED v4.1 Building Design and Construction Reference Guide. USGBC; 2021.
[9] Lu C, Li S, Lu Z, 2021, Building Energy Prediction Using Artificial Neural Networks: A Literature Survey. Energy and Buildings, 2021(233): 110919.
[10] Luo X, Zhang Y, Lu J, et al., 2024, Multi-objective Optimization of Office Park Building Envelope for Nearly Zero Energy. Journal of Building Engineering, 2024(86): 108713.
[11] Yu L, Sun Y, Xu Z, et al., 2021, Multi-agent Deep Reinforcement Learning for HVAC Control. IEEE Transactions on Smart Grid, 12(1): 407&amp;ndash;419.
[12] Zhong S, Aseniero BA, Groom AI, et al., 2025, Towards Interactive AI-assisted Material selection for Sustainable Building Design. Companion Publication of the 2025 ACM Designing Interactive Systems Conference, 567&amp;ndash;573.
[13] Raeissi MM, Knapen R, 2025, Applications of Generative LLMs in Environmental Science: A Systematic Review. Advances in Environmental and Engineering Research, 6(2): 1&amp;ndash;20.
[14] Gao Y, Yiu TW, Shen X, et al., 2026, LLMs in Smart Construction: A Systematic Review. Engineering, Construction and Architectural Management, 33(15): 159&amp;ndash;181. https://doi.org/10.1108/ECAM-12-2024-1402
[15] Center for AI Safety, Scale AI, HLE Contributors Consortium, 2026, A Benchmark of Expert-level Academic Questions to Assess AI Capabilities. Nature, 2026(649): 1139&amp;ndash;1146.
[16] Shen Y, Heacock L, Elias J, et al., 2023, ChatGPT and other LLMs are Double-Edged Swords. Radiology, 307(2): e230163. https://doi.org/10.1148/RADIOL.230163
[17] Kambhampati S, 2024, Can LLMs Plan? The Science and Nonsense of LLM Reasoning Claims. Communications of the ACM, 67(4): 34&amp;ndash;41.
[18] Turpin M, Michael J, Perez E, et al., 2023, Language Models Don&amp;rsquo;t Always Say What They Think. Advances in Neural Information Processing Systems, 2023(36): 74952&amp;ndash;74965.</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
