Jackwang111/M2RL-MT_OPD

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 7, 2026Architecture:Transformer Featherless Exclusive Cold

Jackwang111/M2RL-MT_OPD is a 4 billion parameter language model developed by Samsung Research and Peking University, focusing on multi-domain reinforcement learning for large language models. This model is specifically designed for post-training research, including weight merging and multi-teacher on-policy distillation. It offers a 32768 token context length, making it suitable for advanced research in RLHF and model adaptation.

Loading preview...

Overview

Jackwang111/M2RL-MT_OPD is a 4 billion parameter language model developed by a collaboration between Samsung Research and Peking University. It is a product of research presented at COLM 2026, focusing on "Multi-Domain Reinforcement Learning for Large Language Models." The model is released with a substantial context length of 32768 tokens, providing ample capacity for complex tasks.

Key Capabilities

  • Multi-Domain Reinforcement Learning: Designed to explore techniques for combining or merging knowledge from various domains within reinforcement learning frameworks for LLMs.
  • Post-Training Research: Specifically intended as a base for further research in areas like weight merging and multi-teacher on-policy distillation.
  • High Context Length: Features a 32768 token context window, enabling the processing of extensive inputs and facilitating more complex reasoning or generation tasks.

Good For

  • Academic Research: Ideal for researchers investigating advanced reinforcement learning techniques for large language models.
  • Model Adaptation Studies: Useful for experiments involving the adaptation of LLMs to new domains or tasks through methods like weight merging.
  • On-Policy Distillation: Provides a foundation for exploring multi-teacher on-policy distillation strategies to improve model performance or efficiency.