SeongryongJung/Qwen3-8B-Physics-RLSD-TR
SeongryongJung/Qwen3-8B-Physics-RLSD-TR is an 8 billion parameter Qwen3-based language model specifically fine-tuned for physics-related tasks using the RLSD_TR method. This model is optimized for performance on the SciKnowEval physics dataset, achieving a best validation mean@16 of 68.05%. It is designed to excel in specialized scientific knowledge evaluation within the domain of physics.
Loading preview...
Overview
SeongryongJung/Qwen3-8B-Physics-RLSD-TR is an 8 billion parameter model built upon the Qwen3 architecture, specifically fine-tuned for physics-related tasks. The model utilizes the RLSD_TR (Reinforcement Learning with Self-Distillation and Trust-Region) method during its training process, focusing on the SciKnowEval physics dataset. This specialized training aims to enhance its performance and accuracy in scientific knowledge evaluation within the physics domain.
Key Capabilities
- Physics Domain Specialization: Fine-tuned on the SciKnowEval physics dataset, making it highly proficient in physics-specific queries and evaluations.
- Performance Metrics: Achieved a best validation
mean@16score of 68.05% at step 90, demonstrating its capability in the targeted domain. - RLSD_TR Method: Incorporates the
RLSD_TRtraining method, which includes policy loss moderlsdand teacher regularizationtrust-region, indicating a sophisticated approach to reinforcement learning for improved performance.
Training Details
- Base Model: Qwen3-8B.
- Dataset: Physics / SciKnowEval physics.
- Training Batch Size: 32.
- Total Training Steps: 100, with validation performed every 10 steps.
- Learning Rate: 1e-6.
Good For
- Scientific Knowledge Evaluation: Ideal for tasks requiring deep understanding and accurate responses in physics.
- Research in Physics AI: Useful for researchers exploring specialized language models for scientific domains.
- Physics Education Tools: Can be integrated into applications for physics problem-solving or knowledge assessment.