formalmathatepfl/qwen3-8b-classic-one-shot

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 4, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The formalmathatepfl/qwen3-8b-classic-one-shot is an 8 billion parameter language model fine-tuned from formalmathatepfl/qwen3-cpt. This model is specifically adapted for one-shot learning scenarios, building upon its base Qwen3 architecture. It is intended for tasks requiring rapid adaptation from a single example, leveraging its specialized fine-tuning. Further details on its specific capabilities and intended uses are not extensively documented.

Loading preview...

Model Overview

The formalmathatepfl/qwen3-8b-classic-one-shot is an 8 billion parameter language model derived from formalmathatepfl/qwen3-cpt. This model has undergone a specific fine-tuning process using an sft dataset, indicating an optimization for particular supervised fine-tuning tasks, likely related to one-shot learning as suggested by its name.

Training Details

The model was trained with a learning rate of 2e-05 over 1.0 epoch. Key training hyperparameters include a train_batch_size of 1 and an eval_batch_size of 8, utilizing a multi-GPU distributed setup across 8 devices. The optimizer used was ADAMW_TORCH with standard betas and epsilon, and a cosine learning rate scheduler with a warmup ratio of 0.05 was applied. The training environment leveraged Transformers 4.57.3, Pytorch 2.9.0+cu128, Datasets 4.0.0, and Tokenizers 0.22.2.

Current Limitations

Detailed information regarding the model's specific capabilities, intended uses, limitations, and the exact nature of the training and evaluation data is not provided in the current documentation. Users should exercise caution and conduct further evaluation to determine its suitability for specific applications.