EdwinYue/Mem-T-4B
EdwinYue/Mem-T-4B is a 4 billion parameter language model derived from Qwen3-4B-Instruct, specifically trained using the MoT-GRPO method within the Mem-T framework. This model is designed for applications requiring long-horizon memory agents, leveraging its specialized training for enhanced performance in such tasks. It features a 32768 token context length, making it suitable for processing extensive inputs.
Loading preview...
Model Overview
EdwinYue/Mem-T-4B is a 4 billion parameter language model built upon the Qwen3-4B-Instruct architecture. Its core distinction lies in its training methodology, utilizing MoT-GRPO (Memory-augmented Trajectory-based Gradient Policy Optimization) within the Mem-T framework. This specialized training aims to enhance the model's capabilities as a long-horizon memory agent.
Key Characteristics
- Base Model: Derived from Qwen3-4B-Instruct.
- Parameter Count: 4 billion parameters.
- Context Length: Supports a substantial context window of 32768 tokens.
- Training Framework: Employs the Mem-T framework with MoT-GRPO for optimizing memory-intensive tasks.
Primary Differentiator
The primary differentiator of Mem-T-4B is its focus on densifying rewards for long-horizon memory agents. This suggests an optimization for tasks that require maintaining and utilizing information over extended sequences or interactions, which is crucial for complex reasoning and planning.
Usage and Further Information
For detailed implementation and usage instructions within the Mem-T framework, users are directed to the main Mem-T GitHub repository. This model is particularly suited for research and applications exploring advanced memory mechanisms in large language models.
Citation
This work is associated with the paper "Mem-T: Densifying Rewards for Long-Horizon Memory Agents" by Yue et al. (2026), available on arXiv.