minsu0567/IAD-X1-SFT-answer-last2

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 2, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The minsu0567/IAD-X1-SFT-answer-last2 is a 4.5 billion parameter causal language model, fine-tuned from Qwen/Qwen3.5-4B. This model is specifically optimized for generating answers based on the PA_SFT_2_answer_last2 dataset, making it suitable for question-answering tasks. It features a context length of 32768 tokens, enabling it to process extensive input for detailed responses. The model's fine-tuning focuses on improving its ability to provide relevant answers, distinguishing it from general-purpose LLMs.

Loading preview...

Model Overview

The minsu0567/IAD-X1-SFT-answer-last2 is a 4.5 billion parameter language model, fine-tuned from the base model Qwen/Qwen3.5-4B. This model has been specifically trained on the PA_SFT_2_answer_last2 dataset, indicating a specialization in generating answers, likely for specific question-answering or response generation tasks.

Key Characteristics

  • Base Model: Qwen3.5-4B architecture.
  • Parameter Count: 4.5 billion parameters.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Fine-tuning Focus: Optimized for answer generation based on a specialized dataset.

Training Details

The model was trained with the following hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: 1 (train), 8 (eval)
  • Gradient Accumulation: 2 steps
  • Optimizer: ADAMW_BNB
  • LR Scheduler: Cosine type with 100 warmup steps
  • Epochs: 1.0

Intended Use Cases

This model is particularly suited for applications requiring precise and contextually relevant answers, especially within domains covered by its training dataset. Its large context window allows for processing complex queries and generating comprehensive responses.