dolphinnlp/DeepSeek-R1-Distill-Qwen-1.5B-Math_ADAP

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Feb 6, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

The dolphinnlp/DeepSeek-R1-Distill-Qwen-1.5B-Math_ADAP is a 1.5 billion parameter language model with a 32768-token context length. This model is a distilled version of DeepSeek-R1, fine-tuned from Qwen-1.5B, and specifically adapted for mathematical reasoning tasks. Its primary strength lies in its optimized performance for solving complex mathematical problems.

Loading preview...

Model Overview

The dolphinnlp/DeepSeek-R1-Distill-Qwen-1.5B-Math_ADAP is a 1.5 billion parameter language model designed for mathematical applications. It features a substantial context length of 32768 tokens, allowing it to process and understand lengthy mathematical problems and contexts.

Key Characteristics

  • Parameter Count: 1.5 billion parameters, offering a balance between performance and computational efficiency.
  • Context Length: Supports a 32768-token context window, beneficial for complex, multi-step mathematical reasoning.
  • Mathematical Adaptation: This model is a distilled version of DeepSeek-R1, fine-tuned from the Qwen-1.5B architecture, with specific adaptations to enhance its capabilities in mathematical problem-solving.

Use Cases

Given its specialized training and architecture, this model is particularly well-suited for:

  • Mathematical Reasoning: Solving arithmetic, algebra, geometry, and other mathematical problems.
  • Educational Tools: Assisting in generating explanations or solutions for math-related queries.
  • Research in Math AI: Serving as a base model for further fine-tuning on specific mathematical domains or datasets.

Limitations

The provided model card indicates that much information regarding its development, training data, evaluation, biases, and specific use cases is currently "More Information Needed." Users should be aware of these gaps and exercise caution, especially regarding potential biases or limitations not yet documented. Further details are required to provide comprehensive recommendations for its use.