Jackwang111/M2RL-RL_Agent

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 7, 2026Architecture:Transformer0.0K Featherless Exclusive Cold

Jackwang111/M2RL-RL_Agent is a 4 billion parameter reinforcement learning agent developed by Haoqing Wang et al. from Samsung Research and Peking University, designed for multi-domain reinforcement learning in large language models. This model focuses on exploring strategies for mixing or merging models to enhance performance across diverse domains. It is particularly suited for research in post-training techniques like weight merging and multi-teacher on-policy distillation.

Loading preview...

M2RL-RL_Agent: Multi-Domain Reinforcement Learning for LLMs

Jackwang111/M2RL-RL_Agent is a 4 billion parameter model developed by Haoqing Wang et al. from Samsung Research and Peking University, presented at COLM 2026. This model explores novel approaches to multi-domain reinforcement learning for large language models, specifically investigating whether to "mix" or "merge" different models to achieve better generalization and performance across various tasks.

Key Capabilities & Focus Areas

  • Multi-Domain Reinforcement Learning: Designed to address challenges in applying RL to LLMs across diverse domains.
  • Model Mixing and Merging Strategies: Investigates techniques for combining or integrating models to improve overall capabilities.
  • Research Checkpoints: The model checkpoints are open-sourced to facilitate further research in post-training methods.

Good For

  • Post-training Research: Ideal for researchers exploring advanced techniques such as weight merging, where different model components are combined.
  • Multi-Teacher On-Policy Distillation: Useful for experiments involving knowledge transfer from multiple teacher models to a student model.
  • Understanding LLM Generalization: Provides a foundation for studying how LLMs can be made more robust and adaptable across various domains through RL-based fine-tuning.