Montalte/qwen4b-code-think-localize
TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 11, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
Montalte/qwen4b-code-think-localize is a 4 billion parameter Qwen3-based model developed by Montalte, specifically engineered for directional math-to-code transfer experiments. This model utilizes a "localize" method for merging, focusing on code-related tasks. It is optimized for code generation and reasoning, making it suitable for specialized programming applications.
Loading preview...
Model Overview
Montalte/qwen4b-code-think-localize is a 4 billion parameter model built upon the Qwen3-4B-Base architecture. It represents a unified merge artifact designed for experimental research into directional transfer between mathematical and coding domains.
Key Characteristics
- Base Model: Derived from
Qwen/Qwen3-4B-Base. - Specialization: Focuses on code-related tasks, leveraging a source specialist model
modrill/code-think-q4b-20260908. - Methodology: Employs a unique localize method, described as "Plan B Localize-and-Stitch," validated on source-only MergeBench.
- Training Details: The localization process involved specific parameters including a sparsity target of 0.1, a learning rate of 1e7, 10 epochs, 64 n-shots, and a seed of 42. The mask profile used was
plan_b, with the task specified ascoding. - Architecture Focus: The embedding and language model head layers were skipped during the mask application, indicating a body-only stitch protocol consistent with prior Plan B methods.
Intended Use Cases
This model is particularly suited for research and development in:
- Code Generation: Its specialization in the code domain suggests strong performance in generating programming language constructs.
- Code Reasoning: The "think" mode and experimental nature imply capabilities in understanding and processing code logic.
- Experimental AI: Ideal for exploring advanced merging techniques and transfer learning in specialized domains like code.