phongdq/Qwen1.5_1.8B_SFT_Dolly

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 12, 2026Architecture:Transformer Featherless Exclusive Cold

phongdq/Qwen1.5_1.8B_SFT_Dolly is a 1.8 billion parameter language model based on the Qwen1.5 architecture. This model has been fine-tuned using Supervised Fine-Tuning (SFT) with the Dolly dataset. It is designed for general language understanding and generation tasks, leveraging its compact size for efficient deployment while benefiting from instruction-following capabilities derived from the Dolly dataset.

Loading preview...

Model Overview

This model, phongdq/Qwen1.5_1.8B_SFT_Dolly, is a compact language model with 1.8 billion parameters built upon the Qwen1.5 architecture. It has undergone Supervised Fine-Tuning (SFT) using the Dolly dataset, which typically enhances a model's ability to follow instructions and perform a variety of general-purpose tasks.

Key Characteristics

  • Architecture: Based on the Qwen1.5 family, known for its performance in various language tasks.
  • Parameter Count: At 1.8 billion parameters, it offers a balance between capability and computational efficiency, making it suitable for environments with resource constraints.
  • Context Length: Supports a substantial context window of 32768 tokens, allowing it to process and generate longer sequences of text.
  • Fine-tuning: Utilizes the Dolly dataset for SFT, which aims to improve instruction-following and conversational abilities.

Potential Use Cases

Given its size and fine-tuning approach, this model is potentially suitable for:

  • General text generation: Creating coherent and contextually relevant text.
  • Instruction-following tasks: Responding to prompts and commands in a structured manner.
  • Lightweight deployment: Its smaller size makes it a candidate for applications where larger models are impractical.
  • Experimentation: A good starting point for further fine-tuning on specific downstream tasks due to its foundational SFT.