cutelemonlili/Qwen2.5-Coder-1.5B-Instruct_MATH_training_Qwen_QwQ_32B_Preview

Hugging Face
TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Dec 29, 2024License:otherArchitecture:Transformer0.0K Featherless Exclusive Warm

cutelemonlili/Qwen2.5-Coder-1.5B-Instruct_MATH_training_Qwen_QwQ_32B_Preview is a 1.5 billion parameter language model, fine-tuned from Qwen/Qwen2.5-Coder-1.5B-Instruct. This model is specifically optimized for mathematical tasks, leveraging a 32K context length. It is designed to enhance performance in mathematical reasoning and problem-solving, building upon the Qwen2.5-Coder architecture.

Loading preview...

Model Overview

This model, cutelemonlili/Qwen2.5-Coder-1.5B-Instruct_MATH_training_Qwen_QwQ_32B_Preview, is a specialized fine-tuned version of the Qwen/Qwen2.5-Coder-1.5B-Instruct base model. With 1.5 billion parameters and a 32,768 token context length, it is specifically adapted for mathematical applications.

Key Characteristics

  • Base Model: Fine-tuned from Qwen/Qwen2.5-Coder-1.5B-Instruct, indicating a foundation in code generation and instruction following.
  • Specialization: The model's name and training dataset (MATH_training_Qwen_QwQ_32B_Preview) strongly suggest an optimization for mathematical reasoning and problem-solving tasks.
  • Training Details: Trained for 2 epochs with a learning rate of 1e-05, using a multi-GPU setup. The training loss reached 0.2983, with a validation loss of 0.3738.

Intended Use Cases

Given its fine-tuning on a math-specific dataset, this model is likely best suited for:

  • Mathematical Problem Solving: Assisting with various mathematical challenges.
  • Educational Tools: Potentially useful in applications requiring mathematical explanations or solutions.
  • Code Generation for Math: Leveraging its coder base for generating code related to mathematical computations or algorithms.