amao0o0/InfoDensity-DeepSeek-R1-Distill-Qwen-7B
The amao0o0/InfoDensity-DeepSeek-R1-Distill-Qwen-7B is a 7.6 billion parameter language model, based on the DeepSeek-R1-Distill-Qwen-7B architecture, fine-tuned with InfoDensity, a reinforcement learning reward that prioritizes information-dense reasoning traces. This model is optimized to achieve correct answers with significantly reduced deliberation, demonstrating improved accuracy and efficiency in complex reasoning tasks. It excels in mathematical and general reasoning benchmarks, achieving higher accuracy with shorter generated response lengths compared to its base model.
Loading preview...
Model Overview
The amao0o0/InfoDensity-DeepSeek-R1-Distill-Qwen-7B is a 7.6 billion parameter language model derived from deepseek-ai/DeepSeek-R1-Distill-Qwen-7B. Its key differentiator is the application of InfoDensity, a novel reinforcement learning reward designed to favor information-dense reasoning traces. This method encourages the model to arrive at correct answers with notably less deliberation, making its reasoning processes more efficient.
Key Capabilities & Performance
- Efficient Reasoning: The InfoDensity reward combines an entropy-trajectory quality term with a group-relative length scaling term, applied exclusively to traces leading to correct answers. This trains the model to be more concise and direct in its problem-solving.
- Enhanced Accuracy: On a combined evaluation across AMC23, AIME24, MATH500, and GPQA-Diamond, the InfoDensity-trained model achieved an accuracy of 72.2%, a significant improvement over the base model's 58.1%.
- Reduced Deliberation Length: Concurrently, the model's mean generated length for these tasks was 4.5k tokens, substantially shorter than the base model's 8.5k tokens, indicating greater efficiency.
- Superior Efficiency Score: It demonstrates a positive Accuracy–Efficiency Score (AES) of +1.20, highlighting its balanced improvement in both accuracy and conciseness.
Ideal Use Cases
This model is particularly well-suited for applications requiring:
- Complex Reasoning: Excels in tasks demanding logical deduction and problem-solving, such as mathematical challenges and general question-answering.
- Concise Explanations: Generates shorter, yet highly accurate, reasoning steps, which is beneficial for scenarios where brevity and clarity are paramount.
- Resource-Efficient Inference: Its ability to achieve high accuracy with shorter outputs can lead to more efficient use of computational resources during inference.