OP12138/qwen3-4b-thinking-safechain

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 20, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

OP12138/qwen3-4b-thinking-safechain is a 4 billion parameter causal language model, fine-tuned from Qwen3-4B-Thinking. This model is specifically optimized for tasks related to the 'safechain' dataset, making it suitable for applications requiring specialized knowledge or generation within that domain. It features a substantial context length of 32768 tokens, enhancing its ability to process and generate longer sequences of text.

Loading preview...

Model Overview

This model, OP12138/qwen3-4b-thinking-safechain, is a 4 billion parameter large language model. It is a fine-tuned variant of the Qwen3-4B-Thinking base model, specifically adapted for tasks related to the 'safechain' dataset. The model supports a context length of 32768 tokens, allowing for extensive input and output sequences.

Key Characteristics

  • Base Model: Fine-tuned from Qwen3-4B-Thinking.
  • Parameter Count: 4 billion parameters.
  • Context Length: Supports a substantial 32768 tokens.
  • Specialization: Optimized through fine-tuning on the 'safechain' dataset.

Training Details

The model was trained with a learning rate of 1e-05, a batch size of 2 (accumulated to 16), and utilized the Paged AdamW 8-bit optimizer. Training was conducted for 2 epochs with a cosine learning rate scheduler and a warmup ratio of 0.1. The training environment included Transformers 4.57.6, Pytorch 2.11.0+cu128, Datasets 4.0.0, and Tokenizers 0.22.2.

Intended Use Cases

This model is particularly well-suited for applications that require understanding or generation of content within the domain covered by the 'safechain' dataset. Its fine-tuned nature suggests improved performance on tasks aligned with this specific data distribution compared to its base model.