rageshvr222/rakesh-deepseek-r1

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jun 24, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The rageshvr222/rakesh-deepseek-r1 is an 8 billion parameter Llama-based model developed by rageshvr222. This model was fine-tuned using Unsloth and Huggingface's TRL library, enabling faster training. It is designed for general language tasks, leveraging the DeepSeek-R1-Distill-Llama architecture for efficient performance.

Loading preview...

Overview

rageshvr222/rakesh-deepseek-r1 is an 8 billion parameter language model, fine-tuned by rageshvr222. It is based on the unsloth/DeepSeek-R1-Distill-Llama-8B-bnb-4bit architecture, indicating its foundation in the Llama family of models. A key aspect of its development is the use of Unsloth and Huggingface's TRL library, which facilitated a significantly faster training process.

Key Capabilities

  • Efficient Training: Leverages Unsloth for 2x faster fine-tuning.
  • Llama-based Architecture: Built upon the DeepSeek-R1-Distill-Llama foundation.
  • General Purpose: Suitable for a wide range of language understanding and generation tasks.

Good For

  • Developers seeking a Llama-based model with optimized training.
  • Applications requiring an 8 billion parameter model for various NLP tasks.
  • Experimentation with models fine-tuned using Unsloth's acceleration techniques.