SelectiveDOPD/JustRL-Qwen3-4b-FKLSelective-Top10pct

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 11, 2026Architecture:Transformer Featherless Exclusive Cold

JustRL-Qwen3-4b-FKLSelective-Top10pct is a 4 billion parameter language model based on the Qwen3 architecture, developed by SelectiveDOPD. This model is derived from BiDirect-OPD experiments, specifically from the `justrl_qwen3_4b_fkl_ladder_90_100_kl` training run. It features a substantial 32768 token context length, making it suitable for tasks requiring extensive contextual understanding. The model's primary differentiation lies in its origin from specific reinforcement learning experiments, suggesting potential optimizations for particular interactive or policy-based language generation tasks.

Loading preview...

JustRL-Qwen3-4b-FKLSelective-Top10pct Overview

This model, developed by SelectiveDOPD, is a 4 billion parameter variant of the Qwen3 architecture, distinguished by its origin from the BiDirect-OPD experimental framework. Specifically, it was uploaded from the justrl_qwen3_4b_fkl_ladder_90_100_kl training run, indicating a focus on reinforcement learning (RL) based optimization techniques. The model supports a large context window of 32768 tokens, enabling it to process and generate longer sequences of text while maintaining coherence.

Key Characteristics

  • Architecture: Based on the Qwen3 model family.
  • Parameter Count: 4 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Features a 32768 token context window, beneficial for tasks requiring extensive input or generating lengthy outputs.
  • Training Origin: Derived from specific BiDirect-OPD experiments, suggesting specialized fine-tuning or training methodologies related to reinforcement learning.
  • Checkpoints: Multiple checkpoints are available, ranging from global_step_20 to global_step_280, with global_step_300 being the main branch, allowing users to experiment with different stages of its training progression.

Potential Use Cases

Given its experimental origin in reinforcement learning, this model may be particularly suited for:

  • Research in RL-driven language generation: Exploring the effects of specific RL training strategies on language model behavior.
  • Applications requiring long context understanding: Its 32768 token context window makes it viable for summarization of long documents, complex dialogue systems, or code analysis.
  • Tasks benefiting from specific FKL-ladder optimization: Users interested in models trained with FKLSelective and Top10pct methodologies might find this model relevant for their specific needs.