ggbetz/qwen3-4b-think-s1-full-sft
The ggbetz/qwen3-4b-think-s1-full-sft model is a 4 billion parameter Qwen3-based language model, developed by modrill, specifically fine-tuned for code reasoning tasks. It represents Stage 1 of the 'think' curriculum, focusing on easy and medium code problems. This model utilizes a 'thinking mode' during inference and is optimized for scenarios requiring robust code generation and problem-solving capabilities within its 32768 token context window.
Loading preview...
Model Overview
The ggbetz/qwen3-4b-think-s1-full-sft is a 4 billion parameter language model based on the Qwen3 architecture, developed by modrill. It has undergone full-parameter supervised fine-tuning (SFT) using the think_s1 curriculum stage, which comprises 72,555 samples of easy and medium code reasoning data. This model is specifically designed to enhance code problem-solving abilities.
Key Capabilities
- Code Reasoning: Optimized for easy and medium-difficulty code reasoning tasks, as evidenced by its training on the
think_s1dataset. - Thinking Mode: Configured to operate with
enable_thinking=trueusing the Qwen3 chat template, suggesting an internal reasoning process during generation. - Extended Context: Supports a cutoff length of 16384 tokens during training, and is recommended for inference with
max_tokens=16384.
Performance Highlights
Evaluations on EvalScope (release_latest / AIME) show notable performance in code-related benchmarks:
- LiveCodeBench: Achieved 36.06% pass@1.
- AIME24: Scored 16.67%.
Good For
- Code Generation and Completion: Particularly for tasks involving logical code reasoning.
- Educational Tools: Developing applications that assist with learning or practicing coding problems.
- Automated Code Review: Identifying and suggesting improvements for code segments based on reasoning patterns.