Thrillcrazyer/TACReward7B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 20, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

Thrillcrazyer/TACReward7B is a 7.6 billion parameter reasoning-aware proxy reward model developed by BAELAB at Pusan National University. It utilizes process mining techniques to aggregate stepwise structural deviations between teacher and policy reasoning, producing a scalar reward in the range of [0, 1]. This model is designed to improve the structural quality of reasoning in sparse reward reinforcement learning frameworks, particularly for mathematical problem-solving tasks.

Loading preview...

Overview

Thrillcrazyer/TACReward7B is a 7.6 billion parameter reasoning-aware proxy reward model developed by BAELAB at Pusan National University. This model addresses the limitations of binarized outcome rewards in sparse reward policy gradient methods, especially for complex reasoning tasks like mathematical problem-solving. It provides more granular feedback on intermediate reasoning steps by considering reasoning as a structured process.

Key Capabilities

  • Reasoning-Aware Reward Generation: Produces a scalar reward between 0 and 1 by aggregating stepwise structural deviations between teacher and policy reasoning.
  • Process Mining Integration: Leverages process mining techniques to analyze and compare reasoning structures.
  • Seamless Integration: Designed to be integrated into existing sparse reward frameworks without requiring additional human annotation or architectural modifications.
  • Improved Reasoning Quality: Encourages policy models to enhance the structural quality of their reasoning, leading to consistent performance improvements.

Good For

  • Mathematical Reasoning Tasks: Demonstrates effectiveness in improving performance on multiple mathematical reasoning benchmarks.
  • Reinforcement Learning Fine-tuning: Enhancing sparse reward policy gradient methods for post-training language models.
  • Structured Reasoning Feedback: Providing more detailed feedback on intermediate reasoning steps beyond simple outcome rewards.