Aller au contenu principal
Accès ouvert déclaré 2026 report

What Vague Instructions Cost an AI Agent: A Controlled Test of Deterministic Pre-Execution Clarity Screening

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : ca. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Sixty agentic instructions with deposited intent were executed by a tool-enabled LLM agent and hand-scored against per-instruction checklists. Unscreened, the agent completed the depositor's intent in 25 % of runs — 100 % on determinate instructions, 0 % on underspecified ones: the agent executes a plausible reading, not the intended one. A deterministic pre-execution clarity screen (PREEXEC® engine build 5.9.1, no model call in the scoring path) held 8 of 60 instructions at its shipped default; every held instruction would otherwise have failed (gate precision 8/8) and none of the determinate instructions was held (0 false alarms), lifting completion to 38.3 %; a tighter candidate setting reaches 71.7 %. At constant total compute, a successfully completed task costs 35 % fewer tokens with the shipped screen (65 % at the tighter setting) — the same budget buys more finished work — measured at a median 49,000 tokens per agent run over 650 logged runs. The same 60 instructions were put to three alternative detectors, pre-registered before measurement. Two LLM classifiers reached 93–98 % recall but held 7 and 8 of the 15 determinate instructions as well, and changed 3 of 10 verdicts on repetition; a readability formula separated the strata in the wrong direction (AUC 0.341). None is usable as a gate. All raw judgements are published with this record. The corpus, per-instruction checklists, all agent responses, hand judgements and per-instruction screening verdicts are published alongside this record; every figure is recomputable from them. Score values, thresholds and the derivation of the clarity score are withheld as trade secrets and are not required for verification. The architecture of the software under test is specified in the companion Scientific Whitepaper v3.0 (concept DOI 10.5281/zenodo.18375998). Conflict of interest: the author is the developer and vendor of the screening software under test; the corpus, checklists and analysis paths were fixed in writing before execution, and all raw data are published. PREEXEC® is a registered EU trade mark (No. 019301368) of Clarity IP Holdings Limited.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Où se fait cette recherche

  • Institute on Governance pays non établi dans la notice
    Organisation à but non lucratif
  • Noetik Governance Ltd pays non établi dans la notice
    Entreprise

Institute on Governance et Noetik Governance Ltd.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.