AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b1000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b1000_s0 model is a 4 billion parameter Qwen3-Base architecture, fine-tuned specifically on a scientific web and StackExchange dataset. This model is optimized for understanding and generating content related to scientific topics and technical discussions. With a context length of 32768 tokens, it is designed for specialized applications requiring deep comprehension of scientific and technical language.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_cap_b1000_s0, is a specialized variant of the Qwen3-4B-Base architecture, developed by AmberYifan. It has been fine-tuned on a unique dataset, capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_cap_b1000_s0, which combines scientific web content and StackExchange discussions.

Key Characteristics

  • Base Model: Qwen3-4B-Base, a 4 billion parameter language model.
  • Specialized Fine-tuning: Enhanced for scientific and technical domains through training on a mixed dataset of scientific web pages and StackExchange data.
  • Context Length: Supports a substantial context window of 32768 tokens, suitable for processing longer scientific articles or complex technical queries.

Training Details

The model underwent a fine-tuning process with specific hyperparameters:

  • Learning Rate: 1e-05
  • Batch Size: A total training batch size of 64 (2 per device across 4 GPUs with 8 gradient accumulation steps).
  • Optimizer: ADAMW_TORCH with standard betas and epsilon.
  • Scheduler: Cosine learning rate scheduler with 0.03 warmup steps over 1 epoch.

Intended Use Cases

This model is particularly well-suited for applications requiring a strong understanding of scientific concepts, technical problem-solving, and generating responses in a scientific or technical context, leveraging its specialized training data.