Stage-org/4b-A-200-luna-8k-epoch3
Stage-org/4b-A-200-luna-8k-epoch3 is a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, fine-tuned using a reinforcement learning (RL) method over 3 epochs. This model was trained with a sequence length of 300,000 tokens and utilizes flash attention 2 for efficient processing. It is specifically optimized for tasks requiring advanced reasoning and generation capabilities, leveraging a sophisticated RL setup with an open-ended judge model (gpt-5.6-luna) for evaluation.
Loading preview...
Overview
Stage-org/4b-A-200-luna-8k-epoch3 is a 4.5 billion parameter language model derived from the Qwen3.5-4B base architecture. It has undergone 3 epochs of reinforcement learning (RL) fine-tuning, utilizing a dataset identified as Stage-org/4b-A-200-luna-8k. The training process incorporated a substantial sequence length of 300,000 tokens and leveraged flash attention 2 for optimized performance.
Key Capabilities
- Reinforcement Learning (RL) Fine-tuning: The model was trained using an advanced RL methodology, including a sophisticated open-ended judge model (gpt-5.6-luna) for evaluating generated responses, indicating a focus on quality and alignment.
- Efficient Architecture: Built upon the Qwen3.5-4B model, it benefits from a robust base and incorporates
flash_attention_2for enhanced computational efficiency during inference and training. - Advanced Generation Parameters: The RL setup includes specific generation parameters such as
temperature=0.9,max_tokens=4096, andtop_p=1.0, suggesting a design for diverse and high-quality output generation.
Good for
- Applications requiring advanced reasoning: The use of a powerful judge model during RL training implies a strong emphasis on developing sophisticated reasoning abilities.
- Tasks benefiting from long context: With a training sequence length of 300,000 tokens, the model is likely well-suited for processing and generating content over extended contexts.
- Research and development in RL-tuned LLMs: This model serves as an example of a language model fine-tuned with a complex RL pipeline, making it relevant for researchers exploring similar methodologies.