SeongryongJung/Qwen3-8B-Chemistry-GRPO-TR

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

SeongryongJung/Qwen3-8B-Chemistry-GRPO-TR is an 8 billion parameter Qwen3-based language model fine-tuned using the GRPO method specifically for chemistry-related tasks. This model demonstrates specialized performance on the Chemistry / SciKnowEval chemistry dataset, achieving a validation mean@16 score of 69.02%. It is optimized for applications requiring accurate and nuanced understanding within the domain of chemistry.

Loading preview...

Model Overview

SeongryongJung/Qwen3-8B-Chemistry-GRPO-TR is an 8 billion parameter model built upon the Qwen3-8B architecture, specifically fine-tuned for chemistry applications. The model leverages the GRPO (Generalized Reinforcement Learning from Policy Optimization) method during its training process, focusing on enhancing performance within the chemistry domain.

Key Capabilities and Performance

  • Chemistry Specialization: This model is explicitly trained on the Chemistry / SciKnowEval chemistry dataset, making it highly specialized for tasks within this scientific field.
  • Performance Metrics: It achieved a validation mean@16 score of 69.02% on the Chemistry / SciKnowEval chemistry dataset after 100 training steps.
  • Training Details: The fine-tuning process involved a train batch size of 32, a learning rate of 1e-6, and a total of 100 training steps. The model uses a maximum prompt length of 2048 tokens and a maximum response length of 8192 tokens.

When to Use This Model

This model is particularly well-suited for use cases requiring a deep understanding and generation of content related to chemistry. Developers should consider this model for applications such as:

  • Answering chemistry-specific questions.
  • Analyzing chemical data or literature.
  • Generating text or insights within the field of chemistry.

Its specialized training makes it a strong candidate for tasks where general-purpose LLMs might lack the necessary domain-specific accuracy.