yuxuanw8/qwen3b-racpo-v3-hotpot-checkpoint-30

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 8, 2026Architecture:Transformer Featherless Exclusive Cold

The yuxuanw8/qwen3b-racpo-v3-hotpot-checkpoint-30 is a 3.1 billion parameter language model based on the Qwen architecture, developed by yuxuanw8. This model is fine-tuned using RACPO (Reinforcement Learning from AI Feedback with Contrastive Preference Optimization) specifically for the HotpotQA dataset, indicating its specialization in multi-hop question answering. Its primary use case is to provide accurate and comprehensive answers to complex questions requiring information synthesis from multiple sources, making it suitable for advanced QA systems.

Loading preview...

Overview

This model, yuxuanw8/qwen3b-racpo-v3-hotpot-checkpoint-30, is a 3.1 billion parameter language model built upon the Qwen architecture. It has been specifically fine-tuned using the RACPO (Reinforcement Learning from AI Feedback with Contrastive Preference Optimization) method, targeting the HotpotQA dataset. This specialization suggests its strength in handling complex, multi-hop question answering tasks.

Key Capabilities

  • Multi-hop Question Answering: Optimized for questions that require synthesizing information from multiple passages or facts.
  • Reinforcement Learning from AI Feedback (RACPO): Benefits from advanced fine-tuning techniques designed to improve answer quality and relevance.
  • Qwen Architecture: Leverages the foundational capabilities of the Qwen model family.

Good for

  • Developing advanced question answering systems that need to process and combine information from various sources.
  • Research into reinforcement learning from AI feedback methods for improving language model performance on complex reasoning tasks.
  • Applications requiring robust factual recall and synthesis for intricate queries.