SeongryongJung/qwen3-4b-physics-grpo

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

SeongryongJung/qwen3-4b-physics-grpo is a 4 billion parameter language model fine-tuned from Qwen/Qwen3-4B. This model specializes in physics-related tasks, having been optimized using GRPO on the 'physics' split of the SciKnowEval dataset. It demonstrates a peak validation performance of 62.27% on the val-aux/sciknoweval/reward/mean@16 metric, making it suitable for applications requiring specialized physics knowledge.

Loading preview...

Model Overview

This model, SeongryongJung/qwen3-4b-physics-grpo, is a 4 billion parameter language model derived from Qwen/Qwen3-4B. Its primary distinction lies in its fine-tuning process, which utilized the GRPO (Generalized Reinforcement Learning from Policy Optimization) method specifically on the physics split of the SciKnowEval dataset. This optimization aims to enhance its performance in physics-related reasoning and knowledge tasks.

Performance Highlights

The model's validation performance was tracked using the val-aux/sciknoweval/reward/mean@16 metric. Key performance points include:

  • Best Validation Performance: Achieved 62.27% at step 20.
  • Final Validation Performance: Recorded 53.05% at step 100.

These metrics indicate its specialized capability within the physics domain. The uploaded weights correspond to the final global_step_100/actor checkpoint.

Intended Use Cases

This model is particularly well-suited for applications requiring:

  • Physics-specific question answering: Leveraging its specialized training.
  • Scientific text analysis: Especially within the field of physics.
  • Educational tools: For generating or evaluating physics content.