stellalisy/rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step50

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 12, 2025Architecture:Transformer Featherless Exclusive Cold

The stellalisy/rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step50 model is a 7.6 billion parameter language model with a 32768 token context length. This model is based on the Qwen2.5 architecture and is specifically fine-tuned for mathematical reasoning tasks. Its primary differentiator is its optimization for ground truth reproduction in mathematical contexts, making it suitable for applications requiring precise numerical and logical outputs.

Loading preview...

Model Overview

This model, stellalisy/rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step50, is a 7.6 billion parameter language model built upon the Qwen2.5 architecture. It features a substantial context length of 32768 tokens, enabling it to process and understand extensive inputs for complex tasks. The model's core focus is on mathematical reasoning and the reproduction of ground truth, indicating a specialized fine-tuning process aimed at accuracy in numerical and logical problem-solving.

Key Capabilities

  • Mathematical Reasoning: Optimized for handling mathematical problems and generating accurate solutions.
  • Ground Truth Reproduction: Designed to reproduce precise and verifiable outputs, particularly in quantitative domains.
  • Large Context Window: Supports a 32768 token context, allowing for the processing of detailed and lengthy mathematical or logical queries.

Good For

  • Applications requiring high accuracy in mathematical computations and problem-solving.
  • Tasks where the reproduction of exact, verifiable answers is critical.
  • Research and development in AI for quantitative analysis and logical inference.