AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b8000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b8000_s0 model is a fine-tuned 4 billion parameter Qwen3-4B-Base model, optimized for scientific text processing. It was trained on a specialized dataset combining scientific web content and StackExchange data, enhancing its performance on science-related natural language tasks. This model is designed for applications requiring robust understanding and generation within scientific domains, leveraging its 32768 token context length.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b8000_s0, is a specialized variant of the Qwen3-4B-Base architecture, featuring 4 billion parameters and a 32768 token context length. It has undergone fine-tuning to enhance its capabilities specifically within scientific domains.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen3-4B-Base.
  • Specialized Training Data: The model was fine-tuned on the capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_ppl_b8000_s0 dataset, indicating a focus on scientific web content and StackExchange data.
  • Training Hyperparameters: Training involved a learning rate of 1e-05, a total batch size of 64, and a cosine learning rate scheduler over 1 epoch.

Intended Use Cases

This model is particularly well-suited for applications that involve processing, understanding, or generating text within scientific fields. Its fine-tuning on relevant datasets suggests improved performance on tasks such as:

  • Scientific document analysis.
  • Question answering in science and technology.
  • Summarization of research papers.
  • Content generation for scientific explanations.