ElvisWang111/Qwen3-4B-OutsideTheBox-RL

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ElvisWang111/Qwen3-4B-OutsideTheBox-RL is a 4 billion parameter Qwen3-based causal language model developed by ElvisWang111. It is fine-tuned using outcome-based reinforcement learning on workflow-conditioned mathematical reasoning prompts. This model is specifically designed to research selective reliance on external guidance, particularly in mathematical problem-solving, by training on data mixing helpful and misleading workflows. Its primary strength lies in understanding and utilizing external workflows while maintaining robustness against incorrect guidance.

Loading preview...

Model Overview

ElvisWang111/Qwen3-4B-OutsideTheBox-RL is a 4 billion parameter language model, initialized from Qwen3-4B-OutsideTheBox-SFT and further refined using outcome-based reinforcement learning (RL). This model's development is detailed in the paper "Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?".

Key Capabilities

  • Workflow-conditioned Mathematical Reasoning: Specialized in solving mathematical problems by processing and utilizing external workflow guidance.
  • Selective Reliance on Guidance: Trained to discern between helpful and misleading external workflows, improving its robustness to incorrect information.
  • Research on External Guidance: Primarily intended for research into how language models can selectively rely on external information and handle conflicting guidance.

Training Details

The model's training lineage begins with Qwen/Qwen3-4B as its base, followed by supervised fine-tuning (SFT) to create ElvisWang111/Qwen3-4B-OutsideTheBox-SFT. The final stage involved outcome-based reinforcement learning, using a dataset that combines prompts with both helpful and misleading workflows to enhance its ability to selectively use external information.

Intended Use Cases

This model is a research checkpoint, ideal for studies focusing on:

  • The effectiveness of external workflow utilization in mathematical reasoning.
  • Evaluating model robustness when presented with potentially incorrect or misleading guidance.
  • Advancing understanding of selective information processing in large language models.