nlee-208/limo_S-dsr1b_T-qwq_25

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 11, 2025Architecture:Transformer Featherless Exclusive Cold

The nlee-208/limo_S-dsr1b_T-qwq_25 model is a 1.5 billion parameter language model fine-tuned from deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. Trained using TRL, this model is optimized for general text generation tasks. It leverages a 32768 token context length, making it suitable for processing longer inputs and generating coherent, extended responses. This model is designed for developers seeking a compact yet capable language model for various applications.

Loading preview...

Model Overview

This model, nlee-208/limo_S-dsr1b_T-qwq_25, is a fine-tuned variant of the DeepSeek-R1-Distill-Qwen-1.5B architecture, featuring 1.5 billion parameters. It was developed using the TRL (Transformer Reinforcement Learning) library, indicating a focus on instruction-following or response generation capabilities through supervised fine-tuning (SFT).

Key Capabilities

  • General Text Generation: Capable of generating human-like text based on given prompts.
  • Extended Context Handling: Benefits from a substantial 32768 token context window, allowing for more comprehensive understanding and generation over longer inputs.
  • Fine-tuned Performance: The SFT training process aims to enhance its ability to follow instructions and produce relevant outputs.

Training Details

The model underwent supervised fine-tuning (SFT) using the TRL framework. The training environment utilized specific versions of key libraries including TRL 0.18.1, Transformers 4.52.4, Pytorch 2.7.1, Datasets 4.0.0, and Tokenizers 0.21.1. Further details on the training run can be visualized via its Weights & Biases project.

When to Use This Model

This model is suitable for developers who require a relatively small yet effective language model for tasks such as:

  • Question answering
  • Content creation
  • Dialogue systems
  • Summarization, especially where longer input contexts are beneficial.