AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_random_b8000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_random_b8000_s0 is a 4 billion parameter Qwen3-based language model fine-tuned by AmberYifan. This model is specifically adapted from Qwen/Qwen3-4B-Base using a dataset focused on scientific web and StackExchange content. It is designed for tasks requiring knowledge and understanding within scientific domains, leveraging its 32768 token context length.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_random_b8000_s0, is a fine-tuned variant of the Qwen3-4B-Base architecture, developed by AmberYifan. It leverages a 4 billion parameter base model with a substantial context length of 32768 tokens.

Key Characteristics

  • Base Model: Qwen/Qwen3-4B-Base
  • Parameter Count: 4 billion
  • Context Length: 32768 tokens
  • Fine-tuning Dataset: capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_random_b8000_s0, indicating a specialization in scientific web content and StackExchange data.

Training Details

The model was trained with the following hyperparameters:

  • Learning Rate: 1e-05
  • Optimizer: ADAMW_TORCH
  • Batch Size: 64 (total train batch size)
  • Epochs: 1
  • Frameworks: Transformers 5.8.0, Pytorch 2.13.0+cu130, Datasets 4.0.0, Tokenizers 0.22.2.

Intended Use Cases

Given its fine-tuning on scientific web and StackExchange data, this model is likely best suited for applications requiring specialized knowledge in scientific fields, such as:

  • Scientific text analysis
  • Question answering in science domains
  • Information extraction from scientific articles or discussions.