BeastyZ/Qwen2.5-3B-ConvSearch-R1-TopiOCQA
TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 21, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
BeastyZ/Qwen2.5-3B-ConvSearch-R1-TopiOCQA is a 3.1 billion parameter language model based on Qwen2.5-3B-Instruct, fine-tuned with a 32,768 token context length. It utilizes the ConvSearch-R1 training method on the TopiOCQA dataset. This model is specifically optimized for conversational search and question answering tasks.
Loading preview...
Model Overview
BeastyZ/Qwen2.5-3B-ConvSearch-R1-TopiOCQA is a specialized language model built upon the Qwen2.5-3B-Instruct architecture. With 3.1 billion parameters and a substantial context length of 32,768 tokens, it is designed for advanced conversational applications.
Key Capabilities
- Conversational Search: The model is fine-tuned using the ConvSearch-R1 training method, which focuses on improving performance in conversational search scenarios.
- Question Answering: Its training on the TopiOCQA dataset specifically targets complex open-domain conversational question answering, enabling it to handle multi-turn interactions and retrieve relevant information effectively.
- Robust Base Model: Leveraging the Qwen2.5-3B-Instruct base provides a strong foundation for general language understanding and generation, which is then specialized for search and QA.
When to Use This Model
This model is particularly well-suited for applications requiring:
- Interactive Search Systems: Ideal for building chatbots or virtual assistants that can understand and respond to follow-up questions in a search context.
- Complex QA Systems: Effective for scenarios where users ask intricate questions that might require understanding conversational history to provide accurate answers.
- Research in Conversational AI: Provides a strong baseline for further experimentation and development in conversational search and question answering. The associated code and paper are available for deeper technical insight here and here.