Jackwang111/M2RL-RL_Science

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 7, 2026Architecture:Transformer Featherless Exclusive Cold

Jackwang111/M2RL-RL_Science is a 4 billion parameter language model developed by Haoqing Wang, Xiang Long, Ziheng Li, Yilong Xu, Tingguang Li, and Yehui Tang from Samsung Research and Peking University. This model focuses on multi-domain reinforcement learning for large language models, exploring methods for mixing or merging. It is designed for post-training research, particularly in areas like weight merging and multi-teacher on-policy distillation, and has a context length of 32768 tokens.

Loading preview...

Overview

Jackwang111/M2RL-RL_Science is a 4 billion parameter language model developed by researchers from Samsung Research and Peking University. This model is a result of the paper "To Mix or To Merge? Toward Multi-Domain Reinforcement Learning for Large Language Models," which was accepted to COLM 2026. It explores advanced techniques in multi-domain reinforcement learning for large language models, specifically focusing on strategies for combining or integrating different model components.

Key Capabilities

  • Multi-Domain Reinforcement Learning: Investigates methods for effective reinforcement learning across multiple domains for LLMs.
  • Post-Training Research: Designed to support further research in areas such as weight merging and multi-teacher on-policy distillation.
  • High Context Length: Features a substantial context window of 32768 tokens, enabling processing of longer sequences.

Good For

  • Academic Research: Ideal for researchers exploring novel approaches in multi-domain reinforcement learning for LLMs.
  • Model Merging Experiments: Provides a foundation for experiments involving the merging of model weights.
  • Distillation Techniques: Useful for developing and testing multi-teacher on-policy distillation methods.