zhengj01/DeepSeek-R1-Medical-COT

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 24, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The zhengj01/DeepSeek-R1-Medical-COT is an 8 billion parameter Llama-based causal language model, fine-tuned from unsloth/deepseek-r1-distill-llama-8b-unsloth-bnb-4bit. Developed by zhengj01, this model was trained using Unsloth and Huggingface's TRL library, enabling 2x faster training. It is designed for medical-related tasks, leveraging its Llama architecture and fine-tuning for specialized applications.

Loading preview...

Model Overview

The zhengj01/DeepSeek-R1-Medical-COT is an 8 billion parameter Llama-based model, fine-tuned by zhengj01. It originates from the unsloth/deepseek-r1-distill-llama-8b-unsloth-bnb-4bit base model.

Key Characteristics

  • Architecture: Llama-based, leveraging the DeepSeek-R1 distillation.
  • Training Efficiency: This model was trained with Unsloth and Huggingface's TRL library, which facilitated a 2x faster training process.
  • Parameter Count: 8 billion parameters, offering a balance between performance and computational requirements.
  • Context Length: Supports a context length of 32768 tokens, suitable for processing longer inputs.

Potential Use Cases

This model is specifically fine-tuned for medical applications, indicated by "Medical-COT" in its name. Developers can consider it for tasks requiring specialized understanding and generation within the medical domain, benefiting from its efficient training and Llama foundation.