bcoding/deepseek-llvm-sft-merged-7B
The bcoding/deepseek-llvm-sft-merged-7B is a 7.6 billion parameter Qwen2 model developed by bcoding, fine-tuned from unsloth/DeepSeek-R1-Distill-Qwen-7B-unsloth-bnb-4bit. This model was trained using Unsloth and Huggingface's TRL library, enabling faster training. With a context length of 32768 tokens, it is optimized for specific tasks related to its fine-tuning, likely in code generation or understanding given its 'llvm-sft' designation.
Loading preview...
Model Overview
The bcoding/deepseek-llvm-sft-merged-7B is a 7.6 billion parameter Qwen2 model, developed by bcoding. It was fine-tuned from the unsloth/DeepSeek-R1-Distill-Qwen-7B-unsloth-bnb-4bit base model.
Key Characteristics
- Architecture: Qwen2
- Parameters: 7.6 billion
- Context Length: 32768 tokens
- Training: Fine-tuned using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.
- License: Apache-2.0
Potential Use Cases
Given its 'llvm-sft' designation and fine-tuning, this model is likely specialized for tasks involving code generation, code understanding, or other applications within the LLVM ecosystem. Its substantial context window makes it suitable for processing larger codebases or complex programming instructions.