SeongryongJung/qwen3-4b-biology-grpo
The SeongryongJung/qwen3-4b-biology-grpo model is a 4 billion parameter language model fine-tuned from Qwen/Qwen3-4B. It was optimized using GRPO specifically on a biology dataset, achieving a peak validation performance of 54.62% mean@16 on a scientific knowledge evaluation metric. This model is specialized for tasks requiring biological domain knowledge, leveraging its targeted fine-tuning for improved accuracy in this field.
Loading preview...
Overview
This model, SeongryongJung/qwen3-4b-biology-grpo, is a 4 billion parameter language model derived from the Qwen/Qwen3-4B architecture. It has undergone specialized fine-tuning using the GRPO (Generalized Reinforcement Learning with Policy Optimization) method, specifically targeting the biology split of a scientific knowledge dataset.
Key Capabilities
- Biology Domain Specialization: Fine-tuned on biological data, making it suitable for tasks requiring specific knowledge in this scientific field.
- Performance Metrics: Achieved a peak validation performance of 54.62% on the
val-aux/sciknoweval/reward/mean@16metric during training, indicating its proficiency in scientific knowledge evaluation. - GRPO Optimization: Utilizes GRPO for enhanced performance within its specialized domain.
Training Details
The model was trained with a focus on scientific knowledge evaluation, with validation metrics tracked over 100 steps. The uploaded weights correspond to the final global_step_100/actor checkpoint. The training process involved converting VERL FSDP shards to the Hugging Face format for deployment.