metacognitive-behavioral-tuning/Qwen3-4B-MBT-R
The metacognitive-behavioral-tuning/Qwen3-4B-MBT-R is a 4 billion parameter Qwen3-based causal language model developed by metacognitive-behavioral-tuning, fine-tuned for multi-hop question answering. This model utilizes Metacognitive Behavioral Tuning (MBT-R) with a 5-phase reasoning structure and GRPO, excelling in complex reasoning tasks. It is specifically optimized for question answering benchmarks like HotpotQA, MuSiQue, and 2WikiMultiHopQA, demonstrating strong performance in both in-domain and out-of-domain scenarios.
Loading preview...
Qwen3-4B-MBT-R: Metacognitive Behavioral Tuning for Multi-Hop QA
This model, Qwen3-4B-MBT-R, is a 4 billion parameter language model built upon the Qwen/Qwen3-4B base, specifically engineered for advanced multi-hop question answering. Its core innovation lies in the application of Metacognitive Behavioral Tuning (MBT-R), a method that refines the model's reasoning processes.
Key Capabilities & Training:
- Enhanced Reasoning: MBT-R involves rewriting the model's own reasoning traces into a structured 5-phase format for Supervised Fine-Tuning (SFT), followed by Guided Reinforcement Learning with Policy Optimization (GRPO).
- Multi-Hop QA Specialization: The model is explicitly trained and optimized for complex question answering tasks that require synthesizing information from multiple sources or steps.
- Benchmark Performance: Evaluated on challenging datasets such as HotpotQA (in-domain) and MuSiQue / 2WikiMultiHopQA (out-of-domain), indicating robust performance across various multi-hop QA scenarios.
Should You Use This Model?
- Complex QA: Ideal for applications requiring the model to answer questions that demand multi-step reasoning or information retrieval from disparate facts.
- Research in Reasoning: Suitable for researchers exploring advanced fine-tuning techniques for improving LLM reasoning capabilities, particularly in metacognitive approaches.
- Qwen3 Ecosystem: If you are already working with Qwen3 models and need a specialized variant for complex question answering, this model offers a targeted solution.