SeongryongJung/qwen3-4b-physics-rlsd-ema005
The SeongryongJung/qwen3-4b-physics-rlsd-ema005 model is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B. It was specifically optimized using RLSD (EMA 0.05) on a physics-focused dataset. This model demonstrates a validation performance of 72.11% on the 'val-aux/sciknoweval/reward/mean@16' metric, indicating its specialization in physics-related reasoning tasks. Its 32768 token context length supports processing extensive physics-related texts.
Loading preview...
Model Overview
SeongryongJung/qwen3-4b-physics-rlsd-ema005 is a 4 billion parameter language model derived from the Qwen/Qwen3-4B architecture. It has undergone specialized fine-tuning using the RLSD (Reinforcement Learning from Scientific Data) method with an EMA of 0.05, specifically targeting the physics split of a dataset.
Key Capabilities & Performance
- Physics Specialization: The model is explicitly fine-tuned on physics-related data, making it suitable for tasks within this domain.
- Validation Performance: Achieved a peak validation performance of 72.11% on the
val-aux/sciknoweval/reward/mean@16metric after 100 training steps. This metric is indicative of its ability to handle scientific knowledge evaluation in physics. - Training Details: The fine-tuning process involved a specific checkpoint source (
/mnt/mole/SDPO/L2T/checkpoints/datasets/sciknoweval/physics/qwen3gen-physics-RLSD-Qwen-Qwen3-4B-mbs8-decay0-ema0.05-train64-rollout8-lr1e-6-vllm0.8) and was tracked via a W&B run (run-20260629_190103-rs6wo2bx).
Use Cases
This model is particularly well-suited for applications requiring:
- Physics-related question answering.
- Scientific text analysis within the physics domain.
- Knowledge extraction from physics literature.