TLDR: A study evaluated seven leading Large Language Models (LLMs) on their ability to recommend energy retrofit decisions for homes. It found that LLMs perform well in maximizing COâ‚‚ reduction (up to 92.8% accuracy for top-5 options) but struggle with minimizing payback periods. While LLMs show structured reasoning and sensitivity to key building features like location, they exhibit low consistency among themselves and often use simplified logic. The research concludes that LLMs hold promise for energy retrofit decisions but require significant improvements in accuracy, consistency, and contextual understanding for practical, reliable application.
The complex world of building energy retrofits, aimed at improving energy efficiency and reducing carbon emissions, has traditionally relied on methods that often struggle with being broadly applicable or easily understood. These conventional approaches, whether physics-based simulations or data-driven models, face hurdles like extensive data requirements, limited scalability, and a lack of transparency in their recommendations.
However, recent advancements in artificial intelligence, particularly Large Language Models (LLMs) like ChatGPT, DeepSeek, Gemini, Grok, Llama, and Claude, are showing promise in overcoming these challenges. A new study explores whether these generative AI tools can effectively make energy retrofit decisions, especially in diverse residential settings.
Evaluating AI for Retrofit Decisions
Researchers evaluated seven prominent LLMs on their ability to recommend optimal retrofit packages for 400 diverse homes across 49 U.S. states. The evaluation focused on two key objectives: maximizing COâ‚‚ reduction (a technical goal) and minimizing the payback period (a socio-economic goal). The models’ performance was assessed across four dimensions: accuracy, consistency, sensitivity to input features, and the quality of their reasoning processes.
The study utilized a subset of the ResStock 2024.2 dataset, which contains detailed information on residential buildings, including their characteristics, equipment, occupant behavior, and energy usage. This data was supplemented with cost estimates for 16 different retrofit packages, ranging from heat pump upgrades and insulation improvements to appliance electrification.
Key Findings: A Mixed Bag of Potential
The results indicate that LLMs can indeed generate effective retrofit recommendations, though their performance varies significantly depending on the objective. For maximizing COâ‚‚ reduction, the models showed stronger capabilities, with accuracy reaching up to 54.5% for pinpointing the single best option and an impressive 92.8% for identifying a solution within the top five most effective measures. This suggests that while finding the absolute optimal solution remains challenging, LLMs can consistently suggest near-optimal choices for environmental benefits.
However, when it came to minimizing the payback period, the LLMs struggled more. Their accuracy in this socio-economic context was considerably lower, with the best-performing model only reaching 14.3% for the top-1 match and 52.5% within the top five. This limitation highlights the difficulty LLMs face in balancing complex economic trade-offs and contextual factors compared to clearer engineering objectives.
Consistency among the LLMs was also found to be low, meaning different models often provided divergent recommendations. Interestingly, the models that performed better in terms of accuracy tended to disagree more with the others, suggesting they might be employing unique, more effective strategies. In terms of sensitivity, most LLMs, like traditional physics-based models, prioritized location and architectural characteristics when making decisions, inferring climate conditions from geographical data.
An examination of the reasoning processes, where available (from ChatGPT o3 and DeepSeek R1), revealed a structured, five-step logic that aligns with engineering principles. This included baseline energy assumptions, adjustments for building envelope improvements, system energy calculations, appliance energy assumptions, and a final outcome comparison. However, this reasoning was often simplified and lacked a deeper understanding of nuanced contextual dependencies.
Also Read:
- Enhancing Language Model Reasoning with Dynamic Confidence Assessment
- Understanding Why Code Changes: A Large-Scale Study with AI
Implications and Future Directions
The study underscores the critical role of prompt engineering – how users phrase their questions – in influencing LLM performance. Even minor changes in wording could significantly alter the models’ reasoning and recommendations. This suggests that clear, explicit guidance is essential to ensure LLMs fully grasp the relevance of all input variables.
Furthermore, the research points to limitations in how LLMs represent context internally, sometimes overlooking less salient but crucial features. Their inference processes can also suffer from oversimplified logic, inconsistency in responses, and unintended biases carried over from previous interactions. To address these issues, strategies such as fine-tuning LLMs with domain-specific data, using retrieval-augmented generation to ground responses in validated sources, and employing hybrid modeling approaches (combining LLMs with physics-based simulations) are suggested.
In conclusion, while LLMs demonstrate promising capabilities in making energy retrofit decisions, particularly for environmental objectives, improvements in accuracy, consistency, and contextual understanding are necessary for their reliable application in real-world practice. This research provides a valuable foundation for understanding the strengths and limitations of current AI in this critical area. You can read the full paper for more details at Can AI Make Energy Retrofit Decisions? An Evaluation of Large Language Models.


