minsu0567/IAD-X1-SFT-answer-first

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 1, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The minsu0567/IAD-X1-SFT-answer-first model is a 4.5 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen3.5-4B. It was trained on the PA_SFT_2_reordered dataset. This model is designed for general language generation tasks, leveraging its 32768 token context length for processing longer inputs.

Loading preview...

Model Overview

minsu0567/IAD-X1-SFT-answer-first is a 4.5 billion parameter language model, fine-tuned by minsu0567. It is based on the Qwen/Qwen3.5-4B architecture and has been specifically adapted through supervised fine-tuning (SFT) on the PA_SFT_2_reordered dataset. This model is designed to provide enhanced performance for various language generation and understanding tasks.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3.5-4B.
  • Parameter Count: Features 4.5 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, enabling it to handle longer and more complex inputs.
  • Training Data: Fine-tuned on the PA_SFT_2_reordered dataset, which suggests a focus on specific answer-first or reordered response generation.

Training Details

The model was trained with a learning rate of 1e-05, using the AdamW_BNB optimizer. A cosine learning rate scheduler with 100 warmup steps was employed over 1 epoch. The training utilized a batch size of 1 with 2 gradient accumulation steps, resulting in a total effective batch size of 2.

Intended Use Cases

Given its fine-tuning on an 'answer-first' dataset, this model is likely optimized for applications requiring direct and concise responses, potentially in question-answering systems or conversational AI where immediate answers are prioritized.