OP12138/qwen3-4b-thinking-star1

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 20, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

OP12138/qwen3-4b-thinking-star1 is a 4 billion parameter language model fine-tuned by OP12138 from the Qwen3-4B-Thinking base model. It has a context length of 32768 tokens and is specifically fine-tuned on the 'star1' dataset. This model is intended for tasks aligned with its specialized 'star1' dataset training, offering enhanced performance in that domain.

Loading preview...

Model Overview

OP12138/qwen3-4b-thinking-star1 is a 4 billion parameter language model, fine-tuned by OP12138. It is based on the Qwen3-4B-Thinking architecture and has been specifically adapted through further training on the 'star1' dataset. This fine-tuning process aims to optimize its performance for tasks and data distributions characteristic of the 'star1' dataset.

Training Details

The model was trained with a learning rate of 1e-05, a batch size of 2, and a gradient accumulation of 8, resulting in an effective total batch size of 16. It utilized the Paged AdamW 8-bit optimizer with cosine learning rate scheduling over 5 epochs. The training leveraged Transformers 4.57.6, Pytorch 2.11.0+cu128, Datasets 4.0.0, and Tokenizers 0.22.2.

Intended Use Cases

Given its specialized fine-tuning on the 'star1' dataset, this model is best suited for applications and tasks that align with the characteristics and content of that specific dataset. Developers should consider its training data when evaluating its suitability for their particular use case.