TianHongZXY/CHIMERA-4B-RL

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 2, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

CHIMERA-4B-RL is a 4 billion parameter language model developed by Xinyu Zhu et al., fine-tuned with reinforcement learning on the CHIMERA synthetic reasoning dataset. This model specializes in generalizable cross-domain reasoning, leveraging a compact dataset of 9K samples with rich Chain-of-Thought trajectories across 8 scientific disciplines. It demonstrates reasoning performance comparable to or exceeding significantly larger models, making it suitable for complex analytical tasks.

Loading preview...

CHIMERA-4B-RL Overview

CHIMERA-4B-RL is a 4 billion parameter language model, developed by Xinyu Zhu et al., that has been further trained using reinforcement learning (RL) on the specialized CHIMERA dataset. This model builds upon its base, CHIMERA-4B-SFT, by incorporating RL to enhance its reasoning capabilities.

Key Capabilities

  • Generalizable Cross-Domain Reasoning: The model is specifically designed to excel in reasoning tasks across 8 major scientific disciplines, leveraging a compact yet rich synthetic dataset.
  • Efficient Performance: Despite its modest 4B parameter size, CHIMERA-4B-RL achieves reasoning performance that approaches or matches that of much larger models, such as DeepSeek-R1 and Qwen3-235B, as evidenced by its strong results on benchmarks like GPQA-D and AIME.
  • Reinforcement Learning Enhanced: The application of reinforcement learning on the CHIMERA dataset, which features long Chain-of-Thought (CoT) trajectories, significantly boosts its analytical and problem-solving prowess.

Good For

  • Complex Reasoning Tasks: Ideal for applications requiring advanced logical deduction and problem-solving in scientific or academic contexts.
  • Resource-Constrained Environments: Offers high reasoning performance within a smaller model footprint, making it suitable for deployment where computational resources are limited.
  • Benchmarking and Research: Provides a strong baseline for research into efficient reasoning models and synthetic data training methodologies, as detailed in the paper CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning.