bcoding/deepseek-llvm-sft-merged-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 26, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The bcoding/deepseek-llvm-sft-merged-7B is a 7.6 billion parameter Qwen2 model developed by bcoding, fine-tuned from unsloth/DeepSeek-R1-Distill-Qwen-7B-unsloth-bnb-4bit. This model was trained using Unsloth and Huggingface's TRL library, enabling faster training. With a context length of 32768 tokens, it is optimized for specific tasks related to its fine-tuning, likely in code generation or understanding given its 'llvm-sft' designation.

Loading preview...

Model Overview

The bcoding/deepseek-llvm-sft-merged-7B is a 7.6 billion parameter Qwen2 model, developed by bcoding. It was fine-tuned from the unsloth/DeepSeek-R1-Distill-Qwen-7B-unsloth-bnb-4bit base model.

Key Characteristics

  • Architecture: Qwen2
  • Parameters: 7.6 billion
  • Context Length: 32768 tokens
  • Training: Fine-tuned using Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.
  • License: Apache-2.0

Potential Use Cases

Given its 'llvm-sft' designation and fine-tuning, this model is likely specialized for tasks involving code generation, code understanding, or other applications within the LLVM ecosystem. Its substantial context window makes it suitable for processing larger codebases or complex programming instructions.