motokonltk/3b-sft-base

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 12, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The motokonltk/3b-sft-base model is an instruction-tuned 3.09 billion parameter causal language model from the Qwen2.5-Coder series, developed by Qwen. It features a 32,768 token context length and is specifically designed for enhanced code generation, code reasoning, and code fixing. This model maintains strong performance in mathematics and general competencies, making it suitable for real-world applications like Code Agents.

Loading preview...

Overview

The motokonltk/3b-sft-base model is an instruction-tuned variant of the Qwen2.5-Coder-3B model, part of the latest Qwen2.5-Coder series developed by Qwen. This series, formerly known as CodeQwen, focuses on code-specific large language models. The 3.09 billion parameter model is built on the Qwen2.5 architecture, incorporating features like RoPE, SwiGLU, RMSNorm, and Attention QKV bias, and supports a substantial context length of 32,768 tokens.

Key Capabilities

  • Significantly improved code generation, reasoning, and fixing: Leveraging a massive training dataset of 5.5 trillion tokens, including source code and text-code grounding.
  • Foundation for Code Agents: Designed to enhance coding capabilities while retaining strengths in mathematics and general competencies.
  • Causal Language Model: Utilizes a transformer architecture for sequential text generation.

Good For

  • Code-centric applications: Ideal for tasks requiring robust code generation, debugging, and understanding.
  • Developers needing a compact yet powerful model: The 3B parameter size offers a balance between performance and computational efficiency.
  • Building intelligent coding assistants: Its enhanced coding and general reasoning abilities make it suitable for agent-based systems.