aminediroHF/async-grpo-pr5911

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 19, 2026Architecture:Transformer Featherless Exclusive Cold

The aminediroHF/async-grpo-pr5911 model is a 1.5 billion parameter language model, fine-tuned from deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. It was trained using the AsyncGRPO method, which is designed to enhance mathematical reasoning capabilities. With a context length of 32768 tokens, this model is optimized for tasks requiring advanced reasoning, particularly in mathematical domains.

Loading preview...

Overview

aminediroHF/async-grpo-pr5911 is a 1.5 billion parameter language model, fine-tuned from the deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B base model. It leverages a substantial context length of 32768 tokens, making it suitable for processing longer inputs and maintaining context over extended interactions.

Key Training Method

This model was specifically trained using AsyncGRPO, a method detailed in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This training approach suggests an optimization for tasks that involve complex reasoning, particularly in mathematical contexts.

Use Cases

Given its training methodology and base architecture, this model is well-suited for:

  • Mathematical Reasoning: Tasks requiring logical deduction and problem-solving in mathematical domains.
  • Complex Question Answering: Handling intricate questions that benefit from a deep understanding of context and reasoning.
  • Text Generation: Generating coherent and contextually relevant text, especially in scenarios where logical consistency is important.

Technical Details

The model was trained using the TRL framework (version 1.11.0.dev0) with Transformers 5.15.0 and Pytorch 2.13.0+cu130. Its fine-tuning process aims to enhance capabilities beyond the base model, focusing on the strengths imparted by the AsyncGRPO method.