Jackwang111/M2RL-RL_Math

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 7, 2026Architecture:Transformer Featherless Exclusive Cold

Jackwang111/M2RL-RL_Math is a 4 billion parameter model developed by Haoqing Wang et al. from Samsung Research and Peking University, focusing on multi-domain reinforcement learning for large language models. This model is specifically designed for research into weight merging and multi-teacher on-policy distillation techniques. It offers a foundation for exploring advanced RL methods in LLMs, particularly for post-training research.

Loading preview...

M2RL-RL_Math: Multi-Domain Reinforcement Learning for LLMs

Jackwang111/M2RL-RL_Math is a 4 billion parameter model developed by Haoqing Wang et al. from Samsung Research and Peking University. This model is a product of research into "Multi-Domain Reinforcement Learning for Large Language Models," as detailed in their paper accepted to COLM 2026.

Key Capabilities & Purpose

  • Research Foundation: Provides model checkpoints specifically for post-training research.
  • Advanced RL Techniques: Designed to facilitate studies in weight merging and multi-teacher on-policy distillation.
  • Academic Contribution: Represents a contribution to the field of language modeling, particularly in the application of reinforcement learning.

Good For

  • Academic Researchers: Ideal for researchers exploring novel reinforcement learning methods for LLMs.
  • Post-Training Experimentation: Suitable for experiments involving weight merging, model distillation, and multi-teacher learning strategies.
  • Understanding Multi-Domain RL: Offers a practical resource for those studying how to apply multi-domain RL to improve LLM performance and adaptability.