Montalte/qwen4b-code-think-taskvector
Montalte/qwen4b-code-think-taskvector is a 4 billion parameter Qwen3-based language model developed by Montalte, specifically engineered for directional math-to-code transfer experiments. This model utilizes a 'taskvector' method, applying full SFT weights from a code-specialist source model to enhance its code-thinking capabilities. It is primarily designed for advanced code-related reasoning and problem-solving tasks, leveraging its specialized training for improved performance in the domain.
Loading preview...
Model Overview
Montalte/qwen4b-code-think-taskvector is a 4 billion parameter language model built upon the Qwen/Qwen3-4B-Base architecture. Its primary purpose is to facilitate directional transfer experiments between mathematical and coding domains, specifically enhancing its code-thinking abilities.
Key Characteristics
- Base Model: Derived from
Qwen/Qwen3-4B-Base. - Specialization: Features a source specialist (
modrill/code-think-q4b-20260908) focused on the code domain. - Methodology: Employs a taskvector method, which involves applying the complete Supervised Fine-Tuning (SFT) weights from a dense source specialist. This is equivalent to integrating the full task vector (difference between specialist and base model weights).
- Context Length: Supports a substantial context length of 32768 tokens, beneficial for complex code analysis and generation tasks.
Use Cases
This model is particularly well-suited for:
- Code Reasoning: Tasks requiring deep understanding and logical inference within code.
- Code Generation: Generating code based on complex prompts or specifications.
- Mathematical-to-Code Transfer: Experiments and applications that bridge mathematical concepts with their coding implementations.
- Research: Ideal for researchers exploring advanced fine-tuning techniques like task vectors for domain adaptation in large language models.