ElvisWang111/Qwen3-4B-OutsideTheBox-RL
ElvisWang111/Qwen3-4B-OutsideTheBox-RL is a 4 billion parameter Qwen3-based causal language model developed by ElvisWang111. It is fine-tuned using outcome-based reinforcement learning on workflow-conditioned mathematical reasoning prompts. This model is specifically designed to research selective reliance on external guidance, particularly in mathematical problem-solving, by training on data mixing helpful and misleading workflows. Its primary strength lies in understanding and utilizing external workflows while maintaining robustness against incorrect guidance.
Loading preview...
Model Overview
ElvisWang111/Qwen3-4B-OutsideTheBox-RL is a 4 billion parameter language model, initialized from Qwen3-4B-OutsideTheBox-SFT and further refined using outcome-based reinforcement learning (RL). This model's development is detailed in the paper "Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?".
Key Capabilities
- Workflow-conditioned Mathematical Reasoning: Specialized in solving mathematical problems by processing and utilizing external workflow guidance.
- Selective Reliance on Guidance: Trained to discern between helpful and misleading external workflows, improving its robustness to incorrect information.
- Research on External Guidance: Primarily intended for research into how language models can selectively rely on external information and handle conflicting guidance.
Training Details
The model's training lineage begins with Qwen/Qwen3-4B as its base, followed by supervised fine-tuning (SFT) to create ElvisWang111/Qwen3-4B-OutsideTheBox-SFT. The final stage involved outcome-based reinforcement learning, using a dataset that combines prompts with both helpful and misleading workflows to enhance its ability to selectively use external information.
Intended Use Cases
This model is a research checkpoint, ideal for studies focusing on:
- The effectiveness of external workflow utilization in mathematical reasoning.
- Evaluating model robustness when presented with potentially incorrect or misleading guidance.
- Advancing understanding of selective information processing in large language models.