aminediroHF/async-grpo-ckpt-smoke-r1d-1.5b

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026Architecture:Transformer Featherless Exclusive Cold

The aminediroHF/async-grpo-ckpt-smoke-r1d-1.5b is a 1.5 billion parameter language model fine-tuned from deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. It was trained using the AsyncGRPO method, which is designed to enhance mathematical reasoning capabilities. This model is optimized for tasks requiring advanced reasoning, leveraging its base architecture and specialized training for improved performance in complex problem-solving.

Loading preview...

Model Overview

This model, aminediroHF/async-grpo-ckpt-smoke-r1d-1.5b, is a 1.5 billion parameter language model built upon the deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B base architecture. It has been specifically fine-tuned using the TRL library and incorporates the AsyncGRPO training method.

Key Differentiator: AsyncGRPO Training

The core distinction of this model lies in its training methodology. It utilizes AsyncGRPO, a technique introduced in the research paper "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models" (arXiv:2402.03300). This suggests an optimization for tasks that benefit from enhanced reasoning, particularly in mathematical contexts.

Technical Details

  • Base Model: deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
  • Parameters: 1.5 billion
  • Context Length: 32768 tokens
  • Training Frameworks: TRL (version 1.11.0.dev0), Transformers (version 5.15.0), Pytorch (version 2.13.0+cu130), Datasets (version 5.0.1), Tokenizers (version 0.22.2).

Potential Use Cases

Given its specialized training with AsyncGRPO, this model is likely well-suited for:

  • Reasoning-intensive tasks: Especially those involving logical deduction or problem-solving.
  • Mathematical applications: Where the AsyncGRPO method is designed to improve performance.
  • Applications requiring a compact yet capable model: Its 1.5B parameter count makes it relatively efficient while benefiting from advanced training techniques.