Researchers are addressing a key limitation in large-scale language model (LLM) inference: the disconnect between how these models are trained and how humans solve problems. Shaojie Wang and Liang Zhang from the Hong Kong University of Science and Technology in Guangzhou, along with colleagues, demonstrated that current post-training methods that rely on supervised fine-tuning and reinforcement learning fail to separate the acquisition of a generalizable strategy from its specific application. Their new framework, inspired by human cognition, explicitly trains the LLM in two stages. First with abstract reasoning patterns using chains of thought, then with a reliability-aware reinforcement learning approach for task adaptation. This innovative technique achieves significant performance improvements across multiple benchmarks, both in-distribution (2.19%) and out-of-distribution (4.63%), while reducing training time by up to 70% and token consumption in half. This suggests that aligning artificial intelligence with human cognitive principles can yield significant gains in both capability and efficiency.
A two-step learning process to improve LLM inference includes a first step.
Scientists have demonstrated a new post-training framework for large-scale language models (LLMs) that more closely mimics human cognitive processes, achieving significant improvements in both generalization and training efficiency. This work addresses a fundamental limitation of current methods that use supervised fine-tuning (SFT) followed by reinforcement learning (RL) to optimize complete inference trajectories, which fail to reflect the way humans naturally solve problems. This study proves that human problem solving involves a two-step process. It is about acquiring abstract strategies or meta-knowledge that can be applied to a variety of problems and then adapting these strategies to specific cases. The team achieved this by deploying a cognitively inspired framework to decouple generalizable strategy acquisition from problem-specific execution.
Specifically, we developed Chain-of-Meta-Thought (CoMT), a supervised learning method that focuses on abstract reasoning patterns and filters out specific execution details to facilitate internalization of meta-knowledge. This is in contrast to traditional methods, which involve abstract strategies and problem-specific steps and inhibit the development of transferable skills. The researchers then implemented confidence-coordinated reinforcement learning (CCRL) to optimize task adaptation and utilize confidence-aware rewards at intermediate steps to prevent error aggravation and increase execution reliability. Experiments conducted across four models and eight benchmarks reveal significant performance improvements with this new approach.
This study revealed a 2.19% improvement in in-distribution performance and a 4.63% improvement in out-of-distribution performance compared to the standard method. Furthermore, this study demonstrates a significant reduction in training time, achieving a 65-70% reduction and a 50% reduction in token consumption, highlighting the improved training efficiency of the proposed framework. These results confirm that aligning post-training with human cognitive principles not only provides better generalization ability but also streamlines the learning process. This breakthrough reveals the path to more robust and efficient LLMs that can tackle new problems with greater accuracy and reliability.
Researchers overcame the limitations inherent in existing paradigms by explicitly separating meta-knowledge acquisition from task adaptation. The confidence adjustment mechanism employed in CCRL is particularly noteworthy as it addresses the overconfidence error problem that often plagues multi-step inference processes. This study paves the way for future research focused on further refining the interaction between abstract strategy formation and concrete execution in LLM, which may lead to more human-like reasoning abilities.
Meta-thinking chain for abstract strategy acquisition enables robust zero shots
Scientists have developed a new post-training framework for large-scale language models, inspired by the two-step cognitive process observed in human problem solving. The research team addressed a fundamental gap in current methods of optimizing complete inference trajectories using supervised fine-tuning (SFT) followed by reinforcement learning (RL). This research pioneers a method that separates the acquisition of abstract strategies, called metaknowledge, from their adaptation to specific problem cases. Initially, this study employed chain of meta-thought (CoMT) to focus supervised learning on abstract reasoning patterns and intentionally exclude execution of specific problems.
This approach allowed the model to obtain generalizable strategies regardless of the context of the individual problem. The researchers then implemented confidence-controlled reinforcement learning (CCRL) to utilize confidence-aware rewards in intermediate inference steps to optimize task adaptation. This innovative approach prevents overconfidence errors from propagating through the inference process, increasing the reliability of the final execution. Experiments were conducted across four different models and evaluated their performance on eight benchmark datasets. The team carefully collected data from these benchmarks and evaluated both in-distribution and out-of-distribution generalization capabilities.
Performance was quantified by measuring the improvement compared to the standard method, revealing an improvement of 2.19% within distribution and 4.63% outside distribution. Furthermore, this study demonstrated significant efficiency gains, achieving a 65-70% reduction in training time and a 50% reduction in token consumption. The experimental setup included a rigorous evaluation protocol to compare the proposed framework with existing CoT-SFT+RL pipelines. The team leveraged the power of intermediate step rewards to adjust step rewards based on the model’s confidence level to guide the learning process. This method achieves good generalization and improved training efficiency, aligns the post-training content to human cognitive principles, and demonstrates the potential for more robust and adaptive LLM inference.
CoMT and CCRL significantly improve your LLM reasoning skills
Scientists have developed a new post-training framework for large-scale language models (LLMs) that more closely aligns with human cognitive processes, resulting in significant improvements in both performance and efficiency. This study addresses important limitations of current methods that treat complete reasoning trajectories as the fundamental unit of learning, rather than separating abstract strategy acquisition from adaptation to specific problems. Experimental results reveal that this new approach, which combines Chain of Thought (CoMT) and Confidence Coordinated Reinforcement Learning (CCRL), achieves 2.19% and 4.63% improvement in within-variance and out-of-variance performance, respectively. The team measured the performance of four models: LLaMA-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Qwen3-4B-Instruct, Qwen3-8B, and eight benchmarks, including GSM8K and GSM-Hard.
Data demonstrates that the CoMT+CCRL framework significantly improves training efficiency by reducing training time by 65-70% and token consumption by 50%. The researchers focused on identifying intermediate results computed during reinforcement learning, classifying numerical tokens into those extracted from the problem statement and those generated computationally. This allowed us to pinpoint where the error occurred and cascade subsequent inference steps. Specifically, this study used an entropy-based analysis of the model’s predictive distribution to measure the confidence in these intermediate steps.
The entropy of each calculated numeric token is calculated, with lower entropy indicating higher reliability. The scientists then used the maximum entropy across all computed numbers to define a reward function scaled by confidence, encouraging confidence when the model was correct and uncertainty when it was wrong. Measurements confirm that this approach effectively prevents overconfidence error cascades and improves the reliability of model execution. The test proves the validity of the confidence-adjusted reward function. This function includes an exponential term to highlight the difference between high and low confidence predictions. In the optimization process, we adopted Proximal Policy Optimization (PPO) using the KL divergent regularization reference model. The results show that aligning post-training with human cognitive principles not only provides good generalization ability but also significantly improves training efficiency and paves the way for more robust and adaptive LLMs.
👉 More information
đź—ž From metathinking to action: Cognitively tailored post-training for generalizable and reliable LLM reasoning.
đź§ ArXiv: https://arxiv.org/abs/2601.21909
