SeongryongJung/Qwen3-4B-Physics-GRPO-TR
SeongryongJung/Qwen3-4B-Physics-GRPO-TR is a 4 billion parameter Qwen3 model fine-tuned using the GRPO method specifically on the Physics dataset from SciKnowEval. This model achieves a validation mean@16 score of 68.28% on physics-related tasks, indicating its specialization in scientific reasoning within the physics domain. It is optimized for tasks requiring knowledge and problem-solving in physics, leveraging a 32768 token context length.
Loading preview...
Model Overview
SeongryongJung/Qwen3-4B-Physics-GRPO-TR is a specialized 4 billion parameter language model based on the Qwen3 architecture. It has been fine-tuned using the GRPO (Generalized Reinforcement Learning from Policy Optimization) method, specifically targeting the Physics dataset from SciKnowEval. This training approach aims to enhance its performance and reasoning capabilities in the domain of physics.
Key Capabilities
- Physics Domain Specialization: Achieves a validation mean@16 score of 68.28% on the Physics dataset, demonstrating proficiency in physics-related questions and problems.
- GRPO Fine-tuning: Utilizes the GRPO method for training, which is designed to optimize policy performance.
- Qwen3-4B Base: Built upon the Qwen3-4B model, providing a strong foundation for language understanding and generation.
- Extended Context Length: Supports a maximum model length of 10240 tokens, with a max response length of 8192 tokens, suitable for complex physics problems requiring extensive context.
When to Use This Model
- Physics-specific tasks: Ideal for applications requiring accurate responses or reasoning in the field of physics.
- Scientific research assistance: Can be used to aid in understanding or generating content related to physics concepts.
- Educational tools: Suitable for developing tools that help students learn or solve physics problems.