bigatuna/Qwen3.5-9b-Sushi-Coder

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 25, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

bigatuna/Qwen3.5-9b-Sushi-Coder is a 9 billion parameter language model, fine-tuned from unsloth/qwen3.5-9b. This model specializes in reasoning tasks, having undergone continuation training on the nohurry/Opus-4.6-Reasoning-3000x-filtered dataset. It is optimized for code-related applications and complex reasoning, leveraging its 32768 token context length.

Loading preview...

Model Overview

bigatuna/Qwen3.5-9b-Sushi-Coder is a 9 billion parameter language model, fine-tuned by bigatuna from the unsloth/qwen3.5-9b base model. It has been specifically developed for enhanced reasoning capabilities, building upon an earlier training lineage that included open-r1/codeforces-cots.

Key Characteristics

  • Base Model: Derived from unsloth/qwen3.5-9b.
  • Specialized Training: Underwent continuation training using the nohurry/Opus-4.6-Reasoning-3000x-filtered dataset, focusing on reasoning tasks.
  • Training Method: Utilizes LoRA continuation from an adapter-only Unsloth Studio output, with 16-bit LoRA and bf16 precision.
  • Context Length: Features a 32768 token context window, suitable for handling extensive inputs.
  • Development Tools: Built with Unsloth and TRL, ensuring efficient fine-tuning and deployment.

Ideal Use Cases

This model is particularly well-suited for applications requiring strong reasoning abilities and code-related tasks, given its specialized training data. Its large context window also makes it effective for processing and generating longer sequences of text or code.