When Do Agent Scaffolds Actually Help? Toward Compute-Normalized Evaluation of Workflow-Guided LLM Agents
Résumé fourni par la source
Workflow scaffolds—plans, checklists, verification passes, role decomposition, and artifact tracking—are widely claimed to make LLM agents more reliable. Yet scaffolded workflows also consume more model calls, so apparent gains may reflect extra compute rather than better orchestration. This measurement-agenda preprint develops a compute-normalized framework for evaluating when workflow guidance actually helps. It distinguishes unstructured, repeated-unstructured, plan-first, and plan-plus-check regimes; proposes call-accounted pilot comparisons and stronger token- or dollar-matched follow-up experiments; and identifies task difficulty, verification quality, and coordination overhead as central moderators. The included pilots are deliberately limited and should not be interpreted as a fully powered benchmark. The paper's contribution is a falsifiable evaluation design, explicit accounting rules, and a practical research agenda for separating useful scaffolding from additional inference budget.