SeongryongJung/qwen3-8b-physics-rlsd-ema005
SeongryongJung/qwen3-8b-physics-rlsd-ema005 is an 8 billion parameter Qwen3-based language model fine-tuned with RLSD (Reinforcement Learning from Simulated Data) using an EMA of 0.05. This model is specifically optimized on the 'physics' split of the SciKnowEval dataset, demonstrating enhanced performance in physics-related reasoning tasks. It features a 32K context length and is designed for applications requiring specialized knowledge in physics.
Loading preview...
Model Overview
SeongryongJung/qwen3-8b-physics-rlsd-ema005 is an 8 billion parameter language model built upon the Qwen3 architecture. It has been fine-tuned using Reinforcement Learning from Simulated Data (RLSD) with an Exponential Moving Average (EMA) of 0.05, specifically targeting the 'physics' subset of the SciKnowEval dataset.
Key Capabilities
- Physics Domain Specialization: Optimized for tasks and questions within the physics domain, leveraging its fine-tuning on the
physicssplit of SciKnowEval. - RLSD Fine-tuning: Utilizes RLSD for improved performance, with validation metrics showing a peak
mean@16reward of 68.83% at step 70 during training. - Qwen3 Base: Benefits from the robust capabilities of the Qwen3-8B base model, including a 32K token context length.
Performance Metrics
Validation performance was tracked using the val-aux/sciknoweval/reward/mean@16 metric. The model achieved its best mean@16 of 68.83% at step 70, with a final mean@16 of 67.73% at step 100. This indicates a strong specialization and performance within its target domain.
Good For
- Applications requiring accurate and nuanced understanding of physics concepts.
- Research and development in scientific AI, particularly for physics-related problem-solving.
- Educational tools focused on physics.