nlee-208/limo_S-dsr7b_T-dsr32b_10

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 13, 2025Architecture:Transformer Featherless Exclusive Cold

nlee-208/limo_S-dsr7b_T-dsr32b_10 is a 7.6 billion parameter language model, fine-tuned from deepseek-ai/DeepSeek-R1-Distill-Qwen-7B. This model was trained using Supervised Fine-Tuning (SFT) with the TRL framework, and features a context length of 32768 tokens. It is designed for general text generation tasks, leveraging its base architecture for robust language understanding and generation capabilities.

Loading preview...

Model Overview

nlee-208/limo_S-dsr7b_T-dsr32b_10 is a 7.6 billion parameter language model, fine-tuned from the deepseek-ai/DeepSeek-R1-Distill-Qwen-7B base model. It was developed using the TRL (Transformer Reinforcement Learning) framework, specifically through Supervised Fine-Tuning (SFT).

Key Capabilities

  • Text Generation: Capable of generating coherent and contextually relevant text based on user prompts.
  • Extended Context Window: Supports a substantial context length of 32768 tokens, allowing for processing and generating longer sequences of text.
  • Fine-tuned Performance: Benefits from SFT to enhance its performance on various language tasks.

Training Details

The model's training utilized the TRL framework (version 0.19.1) for supervised fine-tuning. Other framework versions involved include Transformers 4.53.3, Pytorch 2.7.1, Datasets 4.0.0, and Tokenizers 0.21.2. The training process can be visualized via Weights & Biases.

Use Cases

This model is suitable for general text generation applications where a robust language model with a significant context window is beneficial. Developers can integrate it using the Hugging Face transformers pipeline for tasks such as question answering, creative writing, or conversational AI.