Jackwang111/M2RL-RL_Math
TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 7, 2026Architecture:Transformer Featherless Exclusive Cold
Jackwang111/M2RL-RL_Math is a 4 billion parameter model developed by Haoqing Wang et al. from Samsung Research and Peking University, focusing on multi-domain reinforcement learning for large language models. This model is specifically designed for research into weight merging and multi-teacher on-policy distillation techniques. It offers a foundation for exploring advanced RL methods in LLMs, particularly for post-training research.
Loading preview...
M2RL-RL_Math: Multi-Domain Reinforcement Learning for LLMs
Jackwang111/M2RL-RL_Math is a 4 billion parameter model developed by Haoqing Wang et al. from Samsung Research and Peking University. This model is a product of research into "Multi-Domain Reinforcement Learning for Large Language Models," as detailed in their paper accepted to COLM 2026.
Key Capabilities & Purpose
- Research Foundation: Provides model checkpoints specifically for post-training research.
- Advanced RL Techniques: Designed to facilitate studies in weight merging and multi-teacher on-policy distillation.
- Academic Contribution: Represents a contribution to the field of language modeling, particularly in the application of reinforcement learning.
Good For
- Academic Researchers: Ideal for researchers exploring novel reinforcement learning methods for LLMs.
- Post-Training Experimentation: Suitable for experiments involving weight merging, model distillation, and multi-teacher learning strategies.
- Understanding Multi-Domain RL: Offers a practical resource for those studying how to apply multi-domain RL to improve LLM performance and adaptability.