laion/a3-rl-DCAgent_exp_rpt_unitsyn-python-v3-10-8B

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 1, 2026Architecture:Transformer Featherless Exclusive Cold

The laion/a3-rl-DCAgent_exp_rpt_unitsyn-python-v3-10-8B model is an 8 billion parameter language model based on the Qwen3-8B architecture, fine-tuned using Reinforcement Learning (RL) with the DCAgent/exp_rpt_unitsyn-python-v3 dataset. This model is specifically optimized for tasks related to experimental reports and unit synthesis in Python, leveraging a fully-asynchronous SkyRL training approach. Its primary strength lies in generating and understanding Python code within the context of experimental reporting and unit synthesis.

Loading preview...

Model Overview

The laion/a3-rl-DCAgent_exp_rpt_unitsyn-python-v3-10-8B is an 8 billion parameter model built upon the Qwen3-8B architecture. It has been fine-tuned using Reinforcement Learning (RL) with a focus on experimental reports and unit synthesis in Python. The training utilized a fully-asynchronous SkyRL approach, reaching global step 10, with checkpoint selection based on a 5-period reward-EMA.

Key Capabilities

  • Python-centric tasks: Optimized for generating and interpreting Python code, particularly in the domains of experimental reporting and unit synthesis.
  • Reinforcement Learning fine-tuning: Benefits from an RLOO (Reinforcement Learning with Offline Optimization) training methodology, enhancing its performance on specific task trajectories.
  • Dataset specific training: Trained on the DCAgent/exp_rpt_unitsyn-python-v3 dataset for 2 epochs, ensuring specialized knowledge in its target applications.

Training Details

The model's training traces, including the last episode of each trial, are available as a companion dataset: open-athena/a3-rl-DCAgent_exp_rpt_unitsyn-python-v3. This dataset contains the rollouts the policy was trained on after rollback/truncation, providing insight into the training process.