yueqis/full_sft_non_web-qwen-coder-7b-3epochs-30k-5e-5
The yueqis/full_sft_non_web-qwen-coder-7b-3epochs-30k-5e-5 model is a 7.6 billion parameter language model, fine-tuned from Qwen/Qwen2.5-Coder-7B-Instruct. This model is specifically optimized for code-related tasks, leveraging its base architecture's capabilities. It was trained for 3 epochs on a non-web dataset, focusing on improving its performance in coding contexts. The model is suitable for applications requiring robust code generation and understanding.
Loading preview...
Model Overview
This model, yueqis/full_sft_non_web-qwen-coder-7b-3epochs-30k-5e-5, is a fine-tuned variant of the Qwen2.5-Coder-7B-Instruct base model, developed by yueqis. It has 7.6 billion parameters and a context length of 32768 tokens, making it suitable for handling substantial code inputs.
Key Capabilities
- Code-centric Fine-tuning: Specifically trained on a
full_sft_non_webdataset, indicating a focus on code-related tasks and potentially excluding general web data. - Optimized for Instruction Following: Inherits instruction-following capabilities from its base
Qwen2.5-Coder-7B-Instructmodel, making it adept at responding to coding prompts. - Training Details: Underwent 3 epochs of training with a learning rate of 5e-05 and a total batch size of 128, utilizing a cosine learning rate scheduler with a 0.05 warmup ratio.
Good For
- Code Generation: Its fine-tuning on a code-specific dataset suggests strong performance in generating programming code.
- Code Understanding and Analysis: Likely capable of assisting with code comprehension, debugging, and refactoring tasks.
- Instruction-based Coding Tasks: Effective for scenarios where clear, instruction-based coding assistance is required.