Automated Benchmarking Identifies Topic-Dependent Accuracy of Large Language Models Applied for Indoor Air Quality
Résumé fourni par la source
Abstract Indoor air quality (IAQ) is a major public health concern because modern urban residents spend around 90% of their daily lives indoors. Key concerns include controlling pollutant emissions, managing contaminants, and maintaining the ventilation or filtration performance. With the increasing prevalence of large language models (LLMs), IAQ researchers and building engineers may rely on them to guide their decisions. However, the reliability of these tools remains unclear. We benchmarked nine recently released LLMs─spanning closed- and open-weight architectures─against 174 expert-verified conceptual IAQ questions manually derived from authoritative textbooks and peer-reviewed literature. To capture variability in model responses, we tested each question across five independent trials. Outputs were then automatically evaluated against reference answers using a hybrid framework that integrated overlap-based metrics (BLEU, ROUGE-L), semantic similarity (BERT), and AI-as-judge scoring and has been trained and tested against expert assessments to ensure estimation accuracy. Incorporating the AI-based scoring enhanced the reliability of the automatic evaluation system, especially in identifying incorrect LLM responses. Prompting models with domain-specific instructions slightly improved their accuracy and consistency. Importantly, the accuracy of all LLMs dropped below 80% for particle impaction and sampling efficiency topics, possibly because of limited domain-specific training data and specialized terminologies inherent to these topics. This study provides systematic benchmarks of commonly used LLMs in the context of IAQ, offering guidance on their current capabilities, highlighting areas where reliability is insufficient, and pointing out the need for fine-tuned models to ensure trustworthy decision support for certain IAQ topics.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Automated Benchmarking Identifies Topic-Dependent Accuracy of Large Language Models Applied for Indoor Air Quality
- Date Crossref
- 02/09/2026
- Éditeur
- American Chemical Society (ACS)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.