xxang/AStar-Thought-s1.1-32B
xxang/AStar-Thought-s1.1-32B is a 32.8 billion parameter language model developed by xxang, based on the A*-Thought framework. This model is specifically designed for efficient reasoning in low-resource settings by identifying and compressing essential thoughts from reasoning chains. It significantly improves accuracy and efficiency, achieving up to 2.39x accuracy and 2.49x ACU improvements in low-budget scenarios, and up to 33.59% length reduction without substantial accuracy drops in higher-budget settings.
Loading preview...
A*-Thought: Efficient Reasoning for Low-Resource Settings
xxang/AStar-Thought-s1.1-32B is a 32.8 billion parameter model that implements the A*-Thought framework, a novel approach to enhance reasoning efficiency and performance, particularly in environments with limited computational resources. This framework operates by intelligently compressing reasoning chains, identifying and isolating the most critical thoughts.
Key Capabilities & Innovations
- Bidirectional Compression: A*-Thought employs a unique bidirectional importance estimation mechanism at the step level to quantify the significance of each thinking step based on its relevance to both the question and the prospective solution.
- A Search for Path Optimization:* It utilizes A* search at the path level to navigate the exponential search space of reasoning paths, using cost functions that assess path quality and conditional self-information of the solution.
- Significant Efficiency Gains: The model demonstrates substantial improvements in low-budget scenarios, achieving up to 2.39x accuracy and 2.49x ACU (Accuracy-Cost-Utility) improvements. For instance, it boosts the average accuracy of QwQ-32B from 12.3% to 29.4% with an inference budget of 512 tokens.
- Reduced Response Lengths: A*-Thought can reduce average response lengths by up to 33.59% (e.g., from 2826.00 to 1876.66 tokens for QwQ-32B) without significant accuracy degradation in 4096-token settings.
- Generalizability: The framework is compatible with several backbone models, including QwQ-32B and R1-Distill-32B, consistently achieving high ACU scores across various budget conditions.
When to Use This Model
This model is particularly well-suited for applications requiring efficient and accurate reasoning where computational resources or inference budgets are constrained. It excels in scenarios where compact yet effective reasoning paths are crucial, making it ideal for deployment in resource-limited environments or for tasks demanding optimized inference costs without sacrificing performance. Developers should consider this model for tasks that benefit from streamlined thought processes and reduced token generation, such as complex problem-solving, question answering, or logical deduction, especially when operating under tight budget constraints.