voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning
The voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning is a 9 billion parameter causal language model based on Qwen/Qwen3.5-9B, with a 32768 token context length. It is specifically fine-tuned to enhance reasoning-oriented multiple-choice performance, demonstrating improved scores on benchmarks like ARC-Challenge, ARC-Easy, and BoolQ. This model excels in structured reasoning and calibrated answer selection, making it suitable for tasks requiring strong analytical capabilities.
Loading preview...
Model Overview
voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning is a 9 billion parameter causal language model built upon the Qwen/Qwen3.5-9B architecture. Its primary focus is to significantly improve performance on reasoning-oriented multiple-choice tasks while maintaining robust general capabilities. The model was fine-tuned using the voidful/gemini-3.1-opus-4.6-reasoning-merged dataset.
Key Capabilities & Performance
- Enhanced Reasoning: Demonstrates notable improvements in structured reasoning and calibrated answer selection.
- Benchmark Leadership: Achieves the best overall aggregate performance in zero-shot evaluations against
Qwen/Qwen3.5-9BandDavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT. - Specific Strengths: Shows significant gains on ARC-Challenge, ARC-Easy, and BoolQ benchmarks, indicating strong performance in science and reading-style reasoning tasks.
- Competitive General Capability: Maintains competitive performance on other zero-shot commonsense tasks, though it may be slightly behind some baselines on HellaSwag and OpenBookQA.
Ideal Use Cases
This model is particularly well-suited for applications requiring:
- Reasoning-oriented zero-shot performance: Excels in scenarios where logical deduction and precise answer selection are critical.
- Multiple-choice question answering: Optimized for tasks involving structured reasoning and selecting the best option from a set.
- Analytical tasks: Benefits use cases that demand strong analytical capabilities and calibrated responses.