phenixace/Chem-R-Faithful

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:cc-by-4.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The phenixace/Chem-R-Faithful is an 8 billion parameter language model, a continuation of Chem-R-8B, specifically fine-tuned using a verification-grounded process reward (GRPO) to significantly reduce hallucination in chemical reasoning. It excels at generating faithful reasoning traces for chemical tasks, achieving a five-fold reduction in per-claim fabrication rate and nearly doubling the clean-trace rate compared to its base model. This model is optimized for applications requiring high fidelity and grounded reasoning in chemistry, such as molecule captioning, retrosynthesis, and S²-Bench subtasks, while maintaining task performance.

Loading preview...

Chem-R-Faithful: Enhanced Chemical Reasoning with Reduced Hallucination

phenixace/Chem-R-Faithful is an 8 billion parameter model derived from weidawang/Chem-R-8B, fine-tuned using a novel verification-grounded process reward optimization (GRPO). This approach specifically targets and penalizes unfaithful reasoning traces, where claims are not supported by the input, predicted molecule, or reference data.

Key Differentiators & Performance

Unlike models rewarded solely on answer accuracy, Chem-R-Faithful's reward mechanism gates accuracy on the cleanliness of the reasoning trace, ensuring that correct answers are reached through verifiable steps. This results in significant improvements in reasoning fidelity across twelve generative chemical task variants (including ChEBI-20 caption↔molecule, USPTO-50k retrosynthesis, and S²-Bench subtasks):

  • Per-claim fabrication rate: Reduced from 22.56% to 4.35%.
  • Mean ER (fabrication score): Decreased from 10.63 to 2.39.
  • Clean-trace rate (ER = 0): Increased from 45.45% to 87.58%.
  • Task performance: Slightly improved from 58.91 to 61.48.

This demonstrates a substantial reduction in fabrication without trading away task performance.

Training and Usage

The model was trained for 936 steps using GRPO, incorporating a reward function that balances format, accuracy, hallucination reduction, and groundedness. It generates responses in a <think>…</think><answer>…</answer> format, allowing for auditing of the reasoning trace. The underlying code, data, and the structural claim detector are available in the MolReHallu repository.

Limitations

The verifier focuses on structurally decidable claims (functional groups, ring systems, molecular classes) and does not certify a complete chemical argument. Fabrication is reduced but not entirely eliminated, as ER is a fabrication rate over explicit, structurally verifiable claims rather than a recall-complete audit of all reasoning.