Aller au contenu principal
2026 article

Automated Benchmarking Identifies Topic-Dependent Accuracy of Large Language Models Applied for Indoor Air Quality

0Citations signalées — pas une note de qualité
3Institutions déclarées
3Pays d’affiliation déclarés

Résumé fourni par la source

Abstract Indoor air quality (IAQ) is a major public health concern because modern urban residents spend around 90% of their daily lives indoors. Key concerns include controlling pollutant emissions, managing contaminants, and maintaining the ventilation or filtration performance. With the increasing prevalence of large language models (LLMs), IAQ researchers and building engineers may rely on them to guide their decisions. However, the reliability of these tools remains unclear. We benchmarked nine recently released LLMs─spanning closed- and open-weight architectures─against 174 expert-verified conceptual IAQ questions manually derived from authoritative textbooks and peer-reviewed literature. To capture variability in model responses, we tested each question across five independent trials. Outputs were then automatically evaluated against reference answers using a hybrid framework that integrated overlap-based metrics (BLEU, ROUGE-L), semantic similarity (BERT), and AI-as-judge scoring and has been trained and tested against expert assessments to ensure estimation accuracy. Incorporating the AI-based scoring enhanced the reliability of the automatic evaluation system, especially in identifying incorrect LLM responses. Prompting models with domain-specific instructions slightly improved their accuracy and consistency. Importantly, the accuracy of all LLMs dropped below 80% for particle impaction and sampling efficiency topics, possibly because of limited domain-specific training data and specialized terminologies inherent to these topics. This study provides systematic benchmarks of commonly used LLMs in the context of IAQ, offering guidance on their current capabilities, highlighting areas where reliability is insufficient, and pointing out the need for fine-tuned models to ensure trustworthy decision support for certain IAQ topics.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Automated Benchmarking Identifies Topic-Dependent Accuracy of Large Language Models Applied for Indoor Air Quality
Date Crossref
02/09/2026
Éditeur
American Chemical Society (ACS)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Air Quality Monitoring and ForecastingIndoor Air Quality and Microbial ExposureAir Quality and Health Impacts

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.