AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b2000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b2000_s0 is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Base. This model is specifically adapted using the capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_cap_b2000_s0 dataset. It is designed for tasks related to scientific domains, leveraging its base architecture and specialized training data.

Loading preview...

Model Overview

This model, named capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b2000_s0, is a specialized version of the Qwen3-4B-Base architecture. It features 4 billion parameters and a context length of 32768 tokens, making it suitable for processing substantial amounts of text.

Key Capabilities

  • Fine-tuned for Scientific Content: The model has undergone fine-tuning on a specific dataset, capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_cap_b2000_s0, indicating a focus on scientific and technical text from sources like StackExchange.
  • Qwen3-4B-Base Foundation: Built upon the robust Qwen3-4B-Base model, it inherits its general language understanding and generation capabilities.

Training Details

The fine-tuning process utilized a learning rate of 1e-05, a total training batch size of 64, and a cosine learning rate scheduler with 0.03 warmup steps. Training was conducted for 1 epoch across 4 devices with a gradient accumulation of 8 steps. The training environment included Transformers 5.8.0 and Pytorch 2.13.0+cu130.

Good For

  • Applications requiring understanding or generation of scientific and technical content.
  • Tasks involving data from StackExchange or similar Q&A platforms in scientific domains.