SeongryongJung/Qwen3-8B-Physics-GRPO-TR
SeongryongJung/Qwen3-8B-Physics-GRPO-TR is an 8 billion parameter Qwen3-based language model specifically fine-tuned for physics-related tasks. Utilizing the GRPO (Generalized Reinforcement Learning from Policy Optimization) method, this model demonstrates specialized performance on the SciKnowEval physics dataset, achieving a validation mean@16 score of 72.97%. It is optimized for accurate responses in physics domains, making it suitable for scientific question answering and research assistance.
Loading preview...
Model Overview
SeongryongJung/Qwen3-8B-Physics-GRPO-TR is an 8 billion parameter language model built upon the Qwen3 architecture, specifically fine-tuned for physics-related tasks. This model leverages the GRPO (Generalized Reinforcement Learning from Policy Optimization) training method to enhance its performance in scientific domains.
Key Capabilities
- Physics Specialization: Fine-tuned on the
SciKnowEval physicsdataset, demonstrating strong performance in this specific scientific area. - GRPO Training: Utilizes the GRPO method with a batch size of 32, optimizing for improved accuracy in physics problem-solving.
- Performance Metrics: Achieved a peak validation
mean@16score of 72.97% on the SciKnowEval physics test set after 100 training steps. - Context Length: Supports a maximum prompt length of 2048 tokens and a maximum response length of 8192 tokens, with a maximum model length of 10240 tokens.
Good For
- Scientific Question Answering: Excels at answering questions and providing information within the field of physics.
- Physics Research Assistance: Can be used to aid researchers and students in understanding complex physics concepts and solving problems.
- Specialized Applications: Ideal for applications requiring high accuracy and domain-specific knowledge in physics, where general-purpose models might fall short.