ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 25, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ermiaazarkhalili/Qwen3-8B-SFT-Claude-Opus-Reasoning-Unsloth model is an 8 billion parameter Qwen3-based language model, fine-tuned by ermiaazarkhalili. It is specifically optimized for reasoning distillation and chain-of-thought learning, leveraging a dataset of Claude's reasoning traces. This model utilizes Unsloth for efficient training, resulting in faster fine-tuning and reduced VRAM usage, making it suitable for tasks requiring detailed, step-by-step problem-solving.

Loading preview...

Overview

This model, developed by ermiaazarkhalili, is a fine-tuned version of the 8 billion parameter Qwen3-8B base model, enhanced for reasoning capabilities. It leverages Unsloth for efficient training, achieving 2x faster training and 60% less VRAM consumption compared to traditional methods.

Key Capabilities

  • Reasoning Distillation: Optimized for chain-of-thought learning using a dataset of Claude's reasoning traces, including <think> blocks.
  • Efficient Fine-tuning: Trained with Unsloth and QLoRA (4-bit) on an NVIDIA H100 GPU, demonstrating a final training loss of 0.8753 in just over 40 minutes.
  • Context Length: Fine-tuned with a 2,048 token context window.
  • Accessibility: Available for use with Hugging Face Transformers, Unsloth, and also in GGUF formats for CPU and edge inference (e.g., Ollama, llama.cpp).

Use Cases

This model is particularly well-suited for applications requiring detailed, step-by-step problem-solving and complex reasoning. Its training on Claude's reasoning distillation dataset makes it effective for tasks where explicit thought processes are beneficial. It is primarily trained on English data and has a knowledge cutoff limited to its base model's training data.