Xirui1208/readtwice-v7-RL-step-20

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026Architecture:Transformer Featherless Exclusive Cold

The Xirui1208/readtwice-v7-RL-step-20 is a 7.6 billion parameter Qwen2ForCausalLM architecture model, representing a reinforcement learning checkpoint at logical step 20. This model was developed by Xirui1208 and is derived from a physical checkpoint after 60 training steps, including initial pilot steps and subsequent full-scale training. It is designed as a causal language model, indicating its primary use for text generation and understanding tasks.

Loading preview...

Model Overview

The Xirui1208/readtwice-v7-RL-step-20 is a 7.6 billion parameter causal language model based on the Qwen2ForCausalLM architecture. This specific release is a reinforcement learning (RL) checkpoint, marking logical step 20 in its training progression.

Training Details

The model's development involved a multi-stage training process:

  • Initial Pilot Steps: The first 50 logical steps utilized a batch size of 32 with 8 rollouts, mapping to logical step 10.
  • Full-Scale Training: Subsequent 10 steps were conducted with a larger batch size of 128 and 16 rollouts.

This checkpoint was exported from a physical checkpoint at global_step_60, indicating a significant amount of training has been performed to reach this stage.

Key Characteristics

  • Architecture: Qwen2ForCausalLM, a robust architecture known for its performance in various language tasks.
  • Parameter Count: 7.6 billion parameters, placing it in the medium-large category for language models.
  • Context Length: Supports a context length of 32768 tokens, allowing for processing and generating longer sequences of text.

Potential Use Cases

Given its causal language model nature and RL training, this model is suitable for applications requiring advanced text generation, understanding, and potentially tasks where reinforcement learning has optimized its behavior for specific objectives. Developers can leverage its capabilities for tasks such as content creation, summarization, question answering, and more, especially where the RL fine-tuning might offer specialized performance.