SeongryongJung/qwen3-4b-material-grpo
The SeongryongJung/qwen3-4b-material-grpo model is a 4 billion parameter language model fine-tuned from Qwen/Qwen3-4B. It utilizes the GRPO method on the 'material' split of the sciknoweval dataset, achieving a peak validation performance of 78.32% mean@16. This model is specialized for tasks related to material science, demonstrating optimized performance on scientific knowledge evaluation benchmarks.
Loading preview...
Model Overview
SeongryongJung/qwen3-4b-material-grpo is a 4 billion parameter language model, fine-tuned from the base Qwen/Qwen3-4B architecture. This model has undergone specialized training using the GRPO (Generalized Reinforcement Learning from Policy Optimization) method, specifically on the material split of the sciknoweval dataset.
Key Characteristics
- Base Model: Qwen3-4B, a robust foundation for language understanding.
- Fine-tuning Method: GRPO, indicating a reinforcement learning approach to optimize performance.
- Specialized Domain: Trained on the
materialsplit ofsciknoweval, suggesting expertise in material science-related knowledge and evaluation tasks.
Performance Highlights
Validation performance was tracked using the val-aux/sciknoweval/reward/mean@16 metric. The model achieved a best mean@16 score of 78.32% at step 40 during its training process. The final uploaded weights correspond to global_step_100/actor, which recorded a mean@16 of 66.16%.
Intended Use Cases
This model is particularly suited for applications requiring knowledge and reasoning within the domain of material science. Its fine-tuning on the sciknoweval dataset's material split implies strong capabilities in tasks such as:
- Scientific knowledge evaluation in material science.
- Answering questions related to material properties, synthesis, and applications.
- Assisting with research and analysis in the materials domain.