ggbetz/qwen3-4b-think-s1-full-sft

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ggbetz/qwen3-4b-think-s1-full-sft model is a 4 billion parameter Qwen3-based language model, developed by modrill, specifically fine-tuned for code reasoning tasks. It represents Stage 1 of the 'think' curriculum, focusing on easy and medium code problems. This model utilizes a 'thinking mode' during inference and is optimized for scenarios requiring robust code generation and problem-solving capabilities within its 32768 token context window.

Loading preview...

Model Overview

The ggbetz/qwen3-4b-think-s1-full-sft is a 4 billion parameter language model based on the Qwen3 architecture, developed by modrill. It has undergone full-parameter supervised fine-tuning (SFT) using the think_s1 curriculum stage, which comprises 72,555 samples of easy and medium code reasoning data. This model is specifically designed to enhance code problem-solving abilities.

Key Capabilities

  • Code Reasoning: Optimized for easy and medium-difficulty code reasoning tasks, as evidenced by its training on the think_s1 dataset.
  • Thinking Mode: Configured to operate with enable_thinking=true using the Qwen3 chat template, suggesting an internal reasoning process during generation.
  • Extended Context: Supports a cutoff length of 16384 tokens during training, and is recommended for inference with max_tokens=16384.

Performance Highlights

Evaluations on EvalScope (release_latest / AIME) show notable performance in code-related benchmarks:

  • LiveCodeBench: Achieved 36.06% pass@1.
  • AIME24: Scored 16.67%.

Good For

  • Code Generation and Completion: Particularly for tasks involving logical code reasoning.
  • Educational Tools: Developing applications that assist with learning or practicing coding problems.
  • Automated Code Review: Identifying and suggesting improvements for code segments based on reasoning patterns.