Hutgaecha/Qwen3-0.6B-JSON-SFT-GRPO

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Hutgaecha/Qwen3-0.6B-JSON-SFT-GRPO is a 0.8 billion parameter Qwen3 model developed by Hutgaecha, fine-tuned from Hutgaecha/Qwen3-0.6B-JSON-SFT. This model was trained using Unsloth and Huggingface's TRL library, achieving 2x faster training speeds. It is designed for tasks requiring efficient processing within a 32768 token context length.

Loading preview...

Model Overview

Hutgaecha/Qwen3-0.6B-JSON-SFT-GRPO is a compact 0.8 billion parameter language model based on the Qwen3 architecture, developed by Hutgaecha. It is a fine-tuned version of the Hutgaecha/Qwen3-0.6B-JSON-SFT model, optimized for efficiency and performance.

Key Characteristics

  • Architecture: Qwen3 base model.
  • Parameter Count: 0.8 billion parameters, making it suitable for resource-constrained environments or applications requiring faster inference.
  • Training Efficiency: This model was trained with significant speed improvements, achieving 2x faster training times by leveraging the Unsloth library in conjunction with Huggingface's TRL library.
  • Context Length: Supports a substantial context window of 32768 tokens, allowing for processing longer inputs and maintaining conversational coherence over extended interactions.

Intended Use Cases

This model is well-suited for applications where a balance between model size, performance, and training efficiency is crucial. Its optimized training process suggests it could be beneficial for developers looking to quickly iterate on fine-tuning tasks or deploy smaller, yet capable, language models.