OP12138/qwen3-4b-star1

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 26, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

OP12138/qwen3-4b-star1 is a 4 billion parameter causal language model, fine-tuned from the Qwen3-4B architecture. This model is specifically adapted using the 'star1' dataset, indicating a specialized application or domain. It is designed for tasks aligned with its fine-tuning data, offering focused performance within that context.

Loading preview...

Model Overview

OP12138/qwen3-4b-star1 is a 4 billion parameter language model, fine-tuned from the base Qwen/Qwen3-4B architecture. This model has undergone specific adaptation using the 'star1' dataset, suggesting a specialization for particular tasks or data distributions.

Training Details

The fine-tuning process involved the following key hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: A total effective batch size of 16 (train_batch_size: 2, gradient_accumulation_steps: 8)
  • Optimizer: Paged AdamW 8-bit with default betas and epsilon
  • Scheduler: Cosine learning rate scheduler with a 0.05 warmup ratio
  • Epochs: Trained for 5.0 epochs

This configuration indicates a focused fine-tuning approach to adapt the base Qwen3-4B model to the characteristics of the 'star1' dataset. Developers should consider the nature of the 'star1' dataset when evaluating its suitability for their specific use cases.