JackHsieh/sft_with_kl_qwen-4B_NR-short-32k-16-1k-8_lr-3e-06-constant-bs-512_steps-128

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Apr 22, 2026Architecture:Transformer Featherless Exclusive Cold

The JackHsieh/sft_with_kl_qwen-4B_NR-short-32k-16-1k-8_lr-3e-06-constant-bs-512_steps-128 model is a 4 billion parameter language model based on the Qwen architecture, developed by JackHsieh. This model was checkpointed after 128 steps of training, indicating a specific stage in its development. With a context length of 32768 tokens, it is designed for tasks requiring extensive contextual understanding. Its primary application is likely in research or specialized applications where a Qwen-based model with a large context window is beneficial.

Loading preview...

Model Overview

This model, JackHsieh/sft_with_kl_qwen-4B_NR-short-32k-16-1k-8_lr-3e-06-constant-bs-512_steps-128, is a 4 billion parameter language model built upon the Qwen architecture. It was developed by JackHsieh and represents a specific checkpoint from a training run, saved at step 128.

Key Characteristics

  • Architecture: Based on the Qwen model family.
  • Parameter Count: Features 4 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32,768 tokens, enabling it to process and generate longer sequences of text.
  • Training Stage: This particular version is a checkpoint from a training process, specifically from a run involving SFT (Supervised Fine-Tuning) with KL divergence regularization, as indicated by its name.

Potential Use Cases

Given its Qwen foundation and large context window, this model is potentially suitable for:

  • Research and Development: Exploring the effects of SFT with KL regularization on Qwen-based models.
  • Long-form Text Processing: Tasks such as document summarization, detailed content generation, or complex question answering where extensive context is crucial.
  • Specialized Applications: Any domain requiring a robust language model with a significant memory capacity for contextual information.