metacognitive-behavioral-tuning/Qwen3-4B-MBT-R

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The metacognitive-behavioral-tuning/Qwen3-4B-MBT-R is a 4 billion parameter Qwen3-based causal language model developed by metacognitive-behavioral-tuning, fine-tuned for multi-hop question answering. This model utilizes Metacognitive Behavioral Tuning (MBT-R) with a 5-phase reasoning structure and GRPO, excelling in complex reasoning tasks. It is specifically optimized for question answering benchmarks like HotpotQA, MuSiQue, and 2WikiMultiHopQA, demonstrating strong performance in both in-domain and out-of-domain scenarios.

Loading preview...

Qwen3-4B-MBT-R: Metacognitive Behavioral Tuning for Multi-Hop QA

This model, Qwen3-4B-MBT-R, is a 4 billion parameter language model built upon the Qwen/Qwen3-4B base, specifically engineered for advanced multi-hop question answering. Its core innovation lies in the application of Metacognitive Behavioral Tuning (MBT-R), a method that refines the model's reasoning processes.

Key Capabilities & Training:

  • Enhanced Reasoning: MBT-R involves rewriting the model's own reasoning traces into a structured 5-phase format for Supervised Fine-Tuning (SFT), followed by Guided Reinforcement Learning with Policy Optimization (GRPO).
  • Multi-Hop QA Specialization: The model is explicitly trained and optimized for complex question answering tasks that require synthesizing information from multiple sources or steps.
  • Benchmark Performance: Evaluated on challenging datasets such as HotpotQA (in-domain) and MuSiQue / 2WikiMultiHopQA (out-of-domain), indicating robust performance across various multi-hop QA scenarios.

Should You Use This Model?

  • Complex QA: Ideal for applications requiring the model to answer questions that demand multi-step reasoning or information retrieval from disparate facts.
  • Research in Reasoning: Suitable for researchers exploring advanced fine-tuning techniques for improving LLM reasoning capabilities, particularly in metacognitive approaches.
  • Qwen3 Ecosystem: If you are already working with Qwen3 models and need a specialized variant for complex question answering, this model offers a targeted solution.