Aller au contenu principal
Accès ouvert déclaré2026dataset

Dataset for: Evaluating multimodal large language model (LLM) systems on a Transport Canada Private Pilot Licence (PPL) Exam

0Citations signalées
1Institutions associées
1Pays d’affiliation

Résumé fourni par la source

This dataset accompanies a study evaluating the performance of contemporary consumer-facing multimodal large language model (LLM) systems on an official Transport Canada Private Pilot Licence (PPL) sample written examination (TP 13014). Fifteen AI system configurations from OpenAI/ChatGPT, Google Gemini, Anthropic Claude, and Microsoft Copilot were evaluated. Each configuration completed the same examination on three independent occasions using a standardised zero-shot prompt and a fresh chat session for each run. The models were supplied with the examination in PDF format and were required to return answers to all 100 multiple-choice questions without iterative prompting or access to the official answer key. The workbook contains the complete model responses for all three runs, the official answer key used for scoring, per-question correct/incorrect classifications, and adjusted sectional performance data. Transport Canada divides the examination into Air Law, Aeronautics, Meteorology, and Navigation sections and requires both an overall passing score and minimum performance within each section. Eleven Navigation questions depended on a Toronto VFR Navigation Chart that was referenced by the examination but was not available with the publicly accessible source material. These questions were retained in the raw response data but excluded from the primary adjusted analysis. Adjusted overall performance is therefore calculated from 89 valid questions, with the Navigation section calculated from the nine questions that could be answered using the material supplied to the models.

Institutions

BNTIC News n’est pas le producteur de ces données. Métadonnées interrogées à la demande auprès de OpenAlex (CC0). Sources et limites.