Stage-org/4b-A-200-luna-8k-epoch3

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 18, 2026Architecture:Transformer Featherless Exclusive Cold

Stage-org/4b-A-200-luna-8k-epoch3 is a 4.5 billion parameter language model based on the Qwen3.5-4B architecture, fine-tuned using a reinforcement learning (RL) method over 3 epochs. This model was trained with a sequence length of 300,000 tokens and utilizes flash attention 2 for efficient processing. It is specifically optimized for tasks requiring advanced reasoning and generation capabilities, leveraging a sophisticated RL setup with an open-ended judge model (gpt-5.6-luna) for evaluation.

Loading preview...

Overview

Stage-org/4b-A-200-luna-8k-epoch3 is a 4.5 billion parameter language model derived from the Qwen3.5-4B base architecture. It has undergone 3 epochs of reinforcement learning (RL) fine-tuning, utilizing a dataset identified as Stage-org/4b-A-200-luna-8k. The training process incorporated a substantial sequence length of 300,000 tokens and leveraged flash attention 2 for optimized performance.

Key Capabilities

  • Reinforcement Learning (RL) Fine-tuning: The model was trained using an advanced RL methodology, including a sophisticated open-ended judge model (gpt-5.6-luna) for evaluating generated responses, indicating a focus on quality and alignment.
  • Efficient Architecture: Built upon the Qwen3.5-4B model, it benefits from a robust base and incorporates flash_attention_2 for enhanced computational efficiency during inference and training.
  • Advanced Generation Parameters: The RL setup includes specific generation parameters such as temperature=0.9, max_tokens=4096, and top_p=1.0, suggesting a design for diverse and high-quality output generation.

Good for

  • Applications requiring advanced reasoning: The use of a powerful judge model during RL training implies a strong emphasis on developing sophisticated reasoning abilities.
  • Tasks benefiting from long context: With a training sequence length of 300,000 tokens, the model is likely well-suited for processing and generating content over extended contexts.
  • Research and development in RL-tuned LLMs: This model serves as an example of a language model fine-tuned with a complex RL pipeline, making it relevant for researchers exploring similar methodologies.