rageshvr222/rakesh-deepseek-r1
TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jun 24, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold
The rageshvr222/rakesh-deepseek-r1 is an 8 billion parameter Llama-based model developed by rageshvr222. This model was fine-tuned using Unsloth and Huggingface's TRL library, enabling faster training. It is designed for general language tasks, leveraging the DeepSeek-R1-Distill-Llama architecture for efficient performance.
Loading preview...
Overview
rageshvr222/rakesh-deepseek-r1 is an 8 billion parameter language model, fine-tuned by rageshvr222. It is based on the unsloth/DeepSeek-R1-Distill-Llama-8B-bnb-4bit architecture, indicating its foundation in the Llama family of models. A key aspect of its development is the use of Unsloth and Huggingface's TRL library, which facilitated a significantly faster training process.
Key Capabilities
- Efficient Training: Leverages Unsloth for 2x faster fine-tuning.
- Llama-based Architecture: Built upon the DeepSeek-R1-Distill-Llama foundation.
- General Purpose: Suitable for a wide range of language understanding and generation tasks.
Good For
- Developers seeking a Llama-based model with optimized training.
- Applications requiring an 8 billion parameter model for various NLP tasks.
- Experimentation with models fine-tuned using Unsloth's acceleration techniques.