Aller au contenu principal
Accès ouvert déclaré 2026 peer-review

Evaluating Self: Consistency and Tree of Thought Reasoning Strategies in Lightweight Language Models on Math and Logic Benchmark

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : ae. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

AbstractThe evaluation of Tree-of-Thought (ToT) reasoning and self-consistency is important whendeploying lightweight language models (LLMs) that can produce capable results whileoperating on constrained hardware. Self-consistency decoders can enhance the robustness ofchain-of-thought style inference for complex reasoning tasks through the use of selfconsistency, by allowing multiple pathways of diverse reasoning to be aggregated throughsample-and-vote based methods. By enabling look-ahead, backtracking, and intentionalprocessing of alternate solution paths, the concept of ToT generalizes these ideas to create asearch tree that represents all the "thoughts" produced at the intermediate level of processing.This version provides a thorough description of both methods of deriving results from usinglightweight language models through self-consistency (SC) and trees of thoughts (ToT)techniques according to an established benchmark suite consisting of three different types ofarithmetic problems: word, logic, and combinatorial, and implementing various instances ofSC and ToT methodologies (with varying numbers of parameters) on small models that can bechecked for performance (accuracy, sample efficiency and latency) against alternativecomputational environments and the respective experiments were constructed using acombination of algorithmic and statistically-based methods to achieve “faithfulness” tomachines' reasoning by employing an agreement between the methods' individual sampletrajectories and ground-truth solutions. The result whereby self-tested consistency alone (trialand error) was greater in performance as compared to ToT were due only to circumstances withhigh “overhead” costs associated to searching tree-based data with additional overheads forToT data structures permitting “non-trivial” savings achieved through incorrect reasoning, i.e.,by escaping “local consistency” through incorrect use of its rules. Finally, we propose a seriesof guidelines to assist in the selection and tuning of SC and ToT for use in lightweight LLMsfor applications requiring mathematical or logical reasoning, and a framework will be createdto serve as an efficiency-based adaptive-toa-of-self-consistency model for future researchendeavors.Keywords: Deep Learning, Large Language Model, Tree of Thought Reasoning, NeuralNetworks, Artificial Intelligence.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Où se fait cette recherche

  • Amity University Department of Computer Science & Engineering pays non établi dans la notice
    Université ou école supérieure

Department of Computer Science & Engineering — Amity University.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.