juihuichung/awakening-rl-replay100

TEXT GENERATIONPricing:Input $0.408 / Cached $0.0816 / Output $1.972Concurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 10, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The juihuichung/awakening-rl-replay100 is a 32 billion parameter language model, derived from the awakening-goedel-v2-32b-rl base model. It was fine-tuned using 100 behavior-tilted Lean agentic traces, specifically designed to recover tool use capabilities in formal-math LLMs. This model demonstrates a BFCL Non-Live score of 85.8, indicating its effectiveness in maintaining and recovering agentic reasoning after reinforcement learning. It is primarily intended for research in formal mathematics and agentic AI behavior.

Loading preview...

Model Overview

The juihuichung/awakening-rl-replay100 is a 32 billion parameter language model developed by juihuichung. It represents the "RL-base +100 agentic replay" arm of the research presented in the paper "Minimal Agentic Replay Recovers Tool Use in Formal-Math Fine-Tuned LLMs" (COLM 2026). This model is a fine-tuned version of the awakening-goedel-v2-32b-rl base model.

Key Characteristics

  • Base Model: Derived from awakening-goedel-v2-32b-rl.
  • Fine-tuning Data: Trained on 100 behavior-tilted Lean agentic traces, identical to those used for the SFT arm of the project.
  • Training Details: Fine-tuned over 128 epochs with a learning rate of 1e-5 using a cosine schedule.
  • Performance: Achieves a BFCL Non-Live score of 85.8 (referred to as "Goedel-RL-RAG-BT100" in the associated paper's score tables). This demonstrates that the minimal-replay recovery method is effective for RL checkpoints, similar to SFT checkpoints.

Purpose and Use Cases

This model is specifically designed to investigate and demonstrate the recovery of tool use capabilities in formal-math fine-tuned LLMs through minimal agentic replay. It is suitable for:

  • Research in formal mathematics and automated theorem proving.
  • Experiments involving agentic behavior and tool use in LLMs.
  • Studying the impact of replay mechanisms on RL-trained models in specialized domains.

Further resources, including code, data manifests, training recipes, and evaluation scripts, are available on the juihuichung/awakening GitHub repository.