SeongryongJung/qwen3-4b-biology-grpo

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The SeongryongJung/qwen3-4b-biology-grpo model is a 4 billion parameter language model fine-tuned from Qwen/Qwen3-4B. It was optimized using GRPO specifically on a biology dataset, achieving a peak validation performance of 54.62% mean@16 on a scientific knowledge evaluation metric. This model is specialized for tasks requiring biological domain knowledge, leveraging its targeted fine-tuning for improved accuracy in this field.

Loading preview...

Overview

This model, SeongryongJung/qwen3-4b-biology-grpo, is a 4 billion parameter language model derived from the Qwen/Qwen3-4B architecture. It has undergone specialized fine-tuning using the GRPO (Generalized Reinforcement Learning with Policy Optimization) method, specifically targeting the biology split of a scientific knowledge dataset.

Key Capabilities

  • Biology Domain Specialization: Fine-tuned on biological data, making it suitable for tasks requiring specific knowledge in this scientific field.
  • Performance Metrics: Achieved a peak validation performance of 54.62% on the val-aux/sciknoweval/reward/mean@16 metric during training, indicating its proficiency in scientific knowledge evaluation.
  • GRPO Optimization: Utilizes GRPO for enhanced performance within its specialized domain.

Training Details

The model was trained with a focus on scientific knowledge evaluation, with validation metrics tracked over 100 steps. The uploaded weights correspond to the final global_step_100/actor checkpoint. The training process involved converting VERL FSDP shards to the Hugging Face format for deployment.