HerrHruby/RCT-4B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Nov 23, 2025Architecture:Transformer Featherless Exclusive Cold

The HerrHruby/offline_acemath_rl_4b_inst_hard_with_dishsoap_16k_no_summ_curr_step_120 model, developed by Ian Wu, Yuxiao Qu, Amrith Setlur, and Aviral Kumar, is a 4 billion parameter Large Language Model (LLM) trained with the Reasoning Cache (RC) iterative decoding algorithm. This model is designed to continually improve reasoning capabilities over long horizons by exploiting an asymmetry between response generation and summarization. It demonstrates substantial performance gains on challenging benchmarks like HMMT 2025, outperforming comparably sized and many larger reasoning LLMs by effectively scaling test-time compute. Its primary use case is advanced reasoning tasks requiring iterative improvement and extrapolation.

Loading preview...

Overview

This repository hosts the RCT-4B model, a 4 billion parameter Large Language Model (LLM) developed by Ian Wu, Yuxiao Qu, Amrith Setlur, and Aviral Kumar. It is distinguished by its training with the Reasoning Cache (RC) iterative decoding algorithm, detailed in the paper Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL.

Key Capabilities

  • Iterative Reasoning Improvement: RC replaces standard autoregressive decoding during both training and inference, enabling the model to construct reasoning chains that consistently improve across iterations.
  • Long-Horizon Extrapolation: Models trained with RC can extrapolate and continually enhance performance over reasoning horizons significantly longer than those encountered during training.
  • Enhanced Benchmark Performance: RCT-4B has shown substantial performance gains on complex benchmarks such as HMMT 2025, surpassing both similarly sized models and many larger reasoning LLMs by efficiently scaling test-time compute.

Usage Considerations

While the model can be loaded using Hugging Face transformers, leveraging its unique iterative RC-decoding algorithm requires specific inference logic. Users should refer to the official GitHub repository for detailed instructions on RC-decoding for inference (which supports vLLM) and for accessing the training code.