RLMIA/GRPO-3B-SEARCH-R1
RLMIA/GRPO-3B-SEARCH-R1 is an instruction-tuned causal language model from the Qwen2.5 series, developed by Qwen. This 3.09 billion parameter model features a 32,768 token context length and is optimized for enhanced knowledge, coding, and mathematics capabilities. It excels in instruction following, long text generation, structured data understanding, and multilingual support across 29 languages.
Loading preview...
RLMIA/GRPO-3B-SEARCH-R1: Qwen2.5 Instruction-Tuned Model
RLMIA/GRPO-3B-SEARCH-R1 is an instruction-tuned variant of the Qwen2.5 series, a family of large language models developed by Qwen. This specific model has 3.09 billion parameters and supports a substantial context length of 32,768 tokens, with a generation capacity of up to 8,192 tokens.
Key Capabilities and Improvements
This model builds upon the Qwen2 architecture, offering significant enhancements:
- Expanded Knowledge & Specialized Skills: Features substantially more knowledge and improved performance in coding and mathematics, leveraging specialized expert models.
- Instruction Following: Demonstrates significant improvements in adhering to instructions and generating structured outputs, particularly JSON.
- Long Text Generation: Excels at generating extended texts, capable of producing outputs over 8,000 tokens.
- Structured Data Understanding: Enhanced ability to comprehend and process structured data, such as tables.
- Multilingual Support: Provides robust support for over 29 languages, including major global languages like Chinese, English, French, Spanish, German, and Japanese.
- Resilience: More resilient to diverse system prompts, which benefits role-play and chatbot condition-setting.
Architecture and Training
The model utilizes a transformer architecture incorporating RoPE, SwiGLU, RMSNorm, Attention QKV bias, and tied word embeddings. It underwent both pretraining and post-training stages to achieve its instruction-tuned capabilities.
When to Use This Model
This model is particularly well-suited for applications requiring strong instruction following, code generation, mathematical problem-solving, and the ability to process and generate long, structured text or multilingual content. Its resilience to varied system prompts also makes it a good choice for complex conversational AI and role-playing scenarios.