nlee-208/limo_S-dsr1b_T-dsr32b_50
The nlee-208/limo_S-dsr1b_T-dsr32b_50 model is a 1.5 billion parameter language model, fine-tuned from deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. It features a context length of 32768 tokens and was trained using the TRL framework. This model is designed for general text generation tasks, leveraging its fine-tuned architecture for improved performance.
Loading preview...
Model Overview
nlee-208/limo_S-dsr1b_T-dsr32b_50 is a 1.5 billion parameter language model, fine-tuned from the deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B base model. It was developed by nlee-208 and trained using the TRL (Transformer Reinforcement Learning) framework, specifically employing Supervised Fine-Tuning (SFT).
Key Capabilities
- Text Generation: Excels at generating coherent and contextually relevant text based on given prompts.
- Long Context Handling: Supports a substantial context window of 32768 tokens, allowing for processing and generating longer sequences of text.
- Fine-tuned Performance: Benefits from SFT, which typically enhances the model's ability to follow instructions and produce high-quality outputs for various language tasks.
Training Details
The model's training process utilized TRL version 0.19.1, with Transformers 4.53.3, Pytorch 2.7.1, Datasets 4.0.0, and Tokenizers 0.21.2. The training run can be visualized on Weights & Biases.
Good For
- General Purpose Text Generation: Suitable for a wide range of applications requiring text completion, creative writing, or conversational responses.
- Applications Requiring Long Context: Ideal for tasks where understanding and generating text over extended inputs is crucial, such as summarizing long documents or maintaining context in lengthy dialogues.