AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b8000_s0

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b8000_s0 is a 4 billion parameter Qwen3-Base model fine-tuned by AmberYifan. This model is specifically optimized for mathematical and random number generation tasks, having been trained on the capsd_Qwen3-4B-Base-n80000-numina__mix_math_random_b8000_s0 dataset. It features a 32768 token context length and is intended for applications requiring specialized numerical reasoning.

Loading preview...

Model Overview

This model, AmberYifan/capsd-qwen3-numina-Qwen3-4B-Base-math_random_b8000_s0, is a fine-tuned variant of the Qwen3-4B-Base architecture. It has been specifically adapted using the capsd_Qwen3-4B-Base-n80000-numina__mix_math_random_b8000_s0 dataset, indicating a specialization in mathematical and random number generation contexts.

Key Training Details

The model underwent a focused training regimen with the following hyperparameters:

  • Learning Rate: 1e-05
  • Optimizer: ADAMW_TORCH
  • Batch Size: 64 (total train batch size)
  • Epochs: 1
  • Distributed Training: Multi-GPU setup with 4 devices

Intended Use Cases

Given its specialized fine-tuning, this model is likely best suited for applications that require:

  • Mathematical reasoning: Tasks involving numerical operations, problem-solving, or data analysis.
  • Random number generation: Scenarios where the model needs to produce or understand sequences with random properties.

Further information regarding specific intended uses, limitations, and detailed training/evaluation data is noted as needing more documentation.