jahyungu/Qwen2.5-1.5B-Instruct_ifeval-like-data_random

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 5, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Warm

The jahyungu/Qwen2.5-1.5B-Instruct_ifeval-like-data_random model is a 1.5 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct. It features a substantial 32768-token context length, making it suitable for processing longer sequences of text. This model is a specialized iteration, having undergone further fine-tuning on an unspecified dataset, suggesting potential for niche applications beyond its base instruction-following capabilities.

Loading preview...

Model Overview

This model, jahyungu/Qwen2.5-1.5B-Instruct_ifeval-like-data_random, is a fine-tuned variant of the Qwen2.5-1.5B-Instruct base model, developed by Qwen. It features 1.5 billion parameters and supports a 32768-token context length, enabling it to handle extensive textual inputs and generate coherent, long-form responses.

Key Characteristics

  • Base Model: Built upon the robust Qwen2.5-1.5B-Instruct architecture.
  • Fine-tuning: The model has undergone additional fine-tuning on an undisclosed dataset, indicating a specialization or adaptation for particular tasks or data distributions.
  • Context Window: Its large context window is beneficial for tasks requiring deep understanding of lengthy documents or conversations.

Training Details

The fine-tuning process utilized specific hyperparameters:

  • Learning Rate: 2e-05
  • Batch Size: 1 (train), 8 (eval)
  • Gradient Accumulation: 16 steps, resulting in a total effective batch size of 16.
  • Optimizer: AdamW with default betas and epsilon.
  • Scheduler: Linear learning rate scheduler with 200 warmup steps.
  • Epochs: Trained for 3 epochs.

Intended Use

While specific intended uses and limitations are not detailed in the provided information, its instruction-tuned nature combined with further fine-tuning suggests potential for specialized instruction-following tasks where the additional training data has provided a performance boost or adapted it to a particular domain.