akshay-sked/qwen3-14b-svamp-sft

TEXT GENERATIONPricing:Input $0.48 / Output $0.96Concurrent Unit Cost:1Model Size:14BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 4, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The akshay-sked/qwen3-14b-svamp-sft model is a 14 billion parameter language model, fine-tuned from Qwen/Qwen3-14B on the svamp_final_sft_train dataset. This model is specifically optimized for mathematical reasoning tasks, demonstrating a validation loss of 0.2875 after one epoch of training. It is designed to enhance performance in problem-solving scenarios requiring numerical and logical understanding.

Loading preview...

Model Overview

The akshay-sked/qwen3-14b-svamp-sft is a 14 billion parameter language model, derived from the Qwen/Qwen3-14B architecture. This model has undergone specific fine-tuning on the svamp_final_sft_train dataset, indicating an optimization focus on mathematical word problems and similar reasoning tasks.

Key Training Details

  • Base Model: Qwen/Qwen3-14B
  • Dataset: svamp_final_sft_train
  • Training Epochs: 1.0
  • Validation Loss: Achieved 0.2875 at the end of training.
  • Learning Rate: 5e-06
  • Optimizer: ADAMW_TORCH_FUSED
  • Batch Size: A total train batch size of 8 (with gradient accumulation steps of 8).

Potential Use Cases

Given its fine-tuning on a dataset related to mathematical problem-solving, this model is likely best suited for:

  • Mathematical Reasoning: Tasks involving numerical understanding, arithmetic, and logical deduction from problem statements.
  • Educational Applications: Generating solutions or explanations for math problems.
  • Data Analysis Support: Assisting with tasks that require interpreting quantitative information.

Limitations

The model card indicates that more information is needed regarding its full description, intended uses, limitations, and specific training/evaluation data. Users should exercise caution and conduct thorough testing for critical applications, as detailed performance characteristics beyond the reported loss are not yet available.