alirezaaminzadeh/minizinc-codegen-coder
The alirezaaminzadeh/minizinc-codegen-coder is a 1.5 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-Coder-1.5B-Instruct. This model specializes in code generation tasks, leveraging its base architecture and fine-tuning process. With a context length of 32768 tokens, it is designed for generating code, particularly in scenarios requiring a deep understanding of programming logic.
Loading preview...
Overview
This model, alirezaaminzadeh/minizinc-codegen-coder, is a specialized 1.5 billion parameter instruction-tuned language model. It is a fine-tuned version of the Qwen/Qwen2.5-Coder-1.5B-Instruct base model, developed by alirezaaminzadeh. The fine-tuning process utilized the TRL library, indicating a focus on reinforcement learning from human feedback or similar techniques to enhance its performance.
Key Capabilities
- Code Generation: Optimized for generating code, building upon the capabilities of its Qwen2.5-Coder base.
- Instruction Following: Designed to respond to instructions effectively due to its instruction-tuned nature.
- Large Context Window: Features a substantial context length of 32768 tokens, allowing it to process and generate longer code snippets or complex programming logic.
Training Details
The model was trained using Supervised Fine-Tuning (SFT) methods. The training leveraged specific versions of key frameworks:
- TRL: 1.9.2
- Transformers: 5.14.1
- Pytorch: 2.13.0
- Datasets: 5.0.1
- Tokenizers: 0.22.2
Good For
- Developers seeking a specialized model for code generation tasks.
- Applications requiring a model with strong instruction-following capabilities in a coding context.
- Scenarios where a large context window is beneficial for handling extensive codebases or complex programming problems.