MaliDDD/deepseek_legal_L_8B_8n
MaliDDD/deepseek_legal_L_8B_8n is an 8 billion parameter language model fine-tuned from Qwen/Qwen3-8B. This model was trained using the GRPO method, which is designed to enhance mathematical reasoning capabilities. It is optimized for tasks requiring advanced reasoning, particularly in mathematical contexts, leveraging its foundation in the Qwen3-8B architecture.
Loading preview...
Overview
MaliDDD/deepseek_legal_L_8B_8n is an 8 billion parameter language model built upon the Qwen/Qwen3-8B architecture. It has been specifically fine-tuned using the GRPO (Gradient-based Reward Policy Optimization) method, as introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). The training process utilized the TRL library (Transformers Reinforcement Learning).
Key Capabilities
- Enhanced Mathematical Reasoning: The primary focus of this model's fine-tuning is to improve its ability to handle complex mathematical problems and reasoning tasks, leveraging the GRPO method.
- Foundation in Qwen3-8B: Benefits from the robust base capabilities of the Qwen3-8B model, providing a strong general language understanding foundation.
When to Use This Model
- Mathematical Problem Solving: Ideal for applications requiring advanced mathematical reasoning, calculations, and problem-solving.
- Research in Reasoning: Suitable for researchers exploring methods to improve language models' logical and mathematical capabilities.
- Specialized Q&A: Can be applied to question-answering systems where the domain involves numerical or logical deduction.