EhDa24/MNLP_M2_mcqa_model_full_ft1

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.8BQuant:BF16Context Size:32kTool Calling:SupportedPublished:May 29, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Loading

EhDa24/MNLP_M2_mcqa_model_full_ft1 is a fine-tuned language model based on Qwen/Qwen3-0.6B-Base. This model was trained with a learning rate of 3e-05 over 2 epochs, utilizing an AdamW optimizer. Its specific fine-tuning dataset and primary use cases are not detailed in the available information, suggesting it's a specialized adaptation of the Qwen3-0.6B-Base architecture.

Loading preview...

Model Overview

EhDa24/MNLP_M2_mcqa_model_full_ft1 is a specialized language model derived from the Qwen/Qwen3-0.6B-Base architecture. It has undergone a fine-tuning process, though the specific dataset used for this training is not disclosed in the available documentation.

Training Details

The model was trained using the following key hyperparameters:

  • Base Model: Qwen/Qwen3-0.6B-Base
  • Learning Rate: 3e-05
  • Optimizer: AdamW_TORCH with betas=(0.9, 0.999) and epsilon=1e-08
  • Batch Size: 3 (train and eval), with a total effective batch size of 12 due to gradient accumulation
  • Epochs: 2
  • LR Scheduler: Linear type with 200 warmup steps

Current Limitations

Detailed information regarding the model's specific intended uses, limitations, and the training and evaluation data is currently unavailable. Users should be aware that its precise capabilities and optimal applications are not fully documented, requiring further investigation or experimentation for specific use cases.