AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b1000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b1000_s0 is a 4 billion parameter Qwen3-based causal language model fine-tuned by AmberYifan. This model is specifically adapted from Qwen/Qwen3-4B-Base using the capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_ppl_b1000_s0 dataset. It is intended for applications requiring specialized knowledge, likely in scientific or technical domains, given its fine-tuning data. The model has a context length of 32768 tokens, making it suitable for processing longer inputs in its target domain.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-sciweb-stackexchange-Qwen3-4B-Base-science_ppl_b1000_s0, is a fine-tuned variant of the Qwen3-4B-Base architecture. It has been adapted by AmberYifan through further training on a specialized dataset, capsd_Qwen3-4B-Base-n80000-sciweb-stackexchange__mix_science_ppl_b1000_s0.

Key Training Details

  • Base Model: Qwen/Qwen3-4B-Base
  • Learning Rate: 1e-05
  • Batch Size: 2 (train), 8 (eval)
  • Gradient Accumulation: 8 steps, resulting in a total effective batch size of 64
  • Optimizer: AdamW with betas=(0.9, 0.999) and epsilon=1e-08
  • Scheduler: Cosine learning rate scheduler with 0.03 warmup steps
  • Epochs: 1

Intended Use Cases

While specific intended uses and limitations require more information, the fine-tuning on a dataset incorporating "sciweb" and "stackexchange" suggests a specialization in scientific, technical, or question-answering domains. Developers might consider this model for tasks requiring knowledge retrieval or generation within these specific areas, leveraging its 32768-token context window for detailed inputs.