xuechengliu/qa-harness-rl-b

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 21, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The xuechengliu/qa-harness-rl-b model is a 4 billion parameter Qwen3-based language model, fine-tuned using reinforcement learning. It specializes in revising executable harnesses for question-answering solvers based on execution feedback. This model maintains the same architecture, tokenizer, and chat template as the base Qwen3-4B, making it suitable for integration into existing Qwen3 workflows.

Loading preview...

Overview

xuechengliu/qa-harness-rl-b is a 4 billion parameter model built upon the Qwen3 architecture. Its core innovation lies in its fine-tuning approach, which utilizes reinforcement learning to enhance its ability to revise executable harnesses for question-answering (QA) solvers. This process leverages execution feedback, allowing the model to iteratively improve its output based on real-world performance.

Key Capabilities

  • Reinforcement Learning Fine-tuning: Specifically trained to refine executable QA harnesses using feedback from their execution.
  • Qwen3 Compatibility: Retains the original Qwen3-4B architecture, tokenizer, and chat template, ensuring seamless integration and deployment.
  • Specialized Revision: Focuses on improving the correctness and efficiency of code or scripts that interact with QA solvers.

Good For

  • Automated QA System Improvement: Ideal for developers working on systems where a QA solver's harness needs dynamic, feedback-driven revision.
  • Research in RL for Code Generation: Useful for exploring reinforcement learning applications in code generation and refinement, particularly in the context of question answering.
  • Integrating with Qwen3 Ecosystem: Can be served and utilized like any standard Qwen3-4B checkpoint, simplifying deployment for users already familiar with Qwen3.