AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b1000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b1000_s0 is a 4 billion parameter language model, fine-tuned from Qwen/Qwen3-4B-Base. This model is specifically optimized for mathematical reasoning tasks, leveraging a specialized dataset for its training. It is designed to enhance performance in numerical and logical problem-solving contexts, making it suitable for applications requiring robust mathematical capabilities.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b1000_s0, is a fine-tuned variant of the Qwen3-4B-Base architecture. It has been specifically adapted using the capsd_Qwen3-4B-Base-n80000-numina__mix_math_random_b1000_s0 dataset, indicating a focus on mathematical reasoning and problem-solving.

Key Training Details

  • Base Model: Qwen/Qwen3-4B-Base
  • Learning Rate: 1e-05
  • Batch Size: 2 (train), 8 (eval)
  • Gradient Accumulation Steps: 8
  • Optimizer: ADAMW_TORCH
  • LR Scheduler: Cosine with 0.03 warmup steps
  • Epochs: 1

Potential Use Cases

Given its fine-tuning on a math-related dataset, this model is likely best suited for:

  • Mathematical Problem Solving: Tasks involving arithmetic, algebra, or other numerical reasoning.
  • Data Analysis: Generating insights or performing calculations based on structured data.
  • Educational Tools: Assisting with math homework or generating practice problems.

Further details on specific intended uses and limitations would require more information from the original developers.