SeongryongJung/Qwen3-4B-Chemical-GRPO-TR
SeongryongJung/Qwen3-4B-Chemical-GRPO-TR is a 4 billion parameter Qwen3-based language model developed by SeongryongJung, specifically fine-tuned using the GRPO method on a chemical dataset. This model is optimized for tasks within the chemical domain, demonstrating a validation mean@16 score of 68.81%. It is designed for specialized applications requiring chemical knowledge and reasoning, leveraging a 32768 token context length.
Loading preview...
Overview
SeongryongJung/Qwen3-4B-Chemical-GRPO-TR is a specialized 4 billion parameter language model built upon the Qwen3-4B architecture. It has been fine-tuned using the GRPO (Generative Reinforcement Learning with Policy Optimization) method, specifically targeting the chemical domain. The model achieved a best validation mean@16 score of 68.81% on the Chemical dataset, indicating its proficiency in chemical-related tasks.
Key Capabilities
- Chemical Domain Specialization: Optimized for understanding and generating content related to chemistry, as evidenced by its training on a dedicated chemical dataset.
- GRPO Fine-tuning: Utilizes the GRPO method for enhanced performance, with a training batch size of 32 and 8 rollout samples per prompt.
- Performance Metrics: Achieved a
val_mean16of 0.688095238095 at step 100, demonstrating consistent improvement throughout its 100 training steps. - Qwen3-4B Base: Leverages the robust capabilities of the Qwen3-4B base model, providing a strong foundation for its specialized chemical knowledge.
Good For
- Applications requiring chemical reasoning and data interpretation.
- Research and development in chemistry-related fields.
- Tasks that benefit from a model specifically trained on chemical datasets.
- Developers looking for a compact yet specialized model for chemical domain tasks.