AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b2000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b2000_s0 model is a fine-tuned Qwen3-4B-Base, a 4 billion parameter language model, specifically adapted for scientific web and StackExchange content. It was trained with a learning rate of 1e-05 over one epoch on a specialized dataset. This model is optimized for tasks related to scientific text processing and understanding, leveraging its base architecture's 32768 token context length.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b2000_s0, is a specialized fine-tuned version of the Qwen3-4B-Base model. It leverages the 4 billion parameter architecture of Qwen3 and its 32768 token context length.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3-4B-Base.
  • Specialized Training: The model was fine-tuned on the capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_ppl_b2000_s0 dataset, indicating a focus on scientific web content and StackExchange data.
  • Training Configuration: Training involved a learning rate of 1e-05, a train_batch_size of 2, and a gradient_accumulation_steps of 8, resulting in a total_train_batch_size of 64. It was trained for 1 epoch using the AdamW optimizer with a cosine learning rate scheduler.

Potential Use Cases

Given its fine-tuning on scientific and StackExchange data, this model is likely suitable for:

  • Processing and generating text related to scientific domains.
  • Answering questions or summarizing content from scientific articles or forums.
  • Applications requiring understanding of technical discussions found on platforms like StackExchange.