MemTensor/MemReader-4B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 11, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

MemTensor/MemReader-4B is a 4 billion parameter causal language model developed by MemTensor, fine-tuned from Qwen3-4B. It is specifically designed as a Memory Operator sub-model for MemOS, excelling at memory extraction tasks from conversations and documents. This model supports local deployment, offers faster and more accurate memory extraction, and provides multilingual support for English and Chinese, outperforming GPT-4o-mini in memory extraction performance while being significantly more resource-efficient.

Loading preview...

MemReader-4B: A Specialized Memory Extraction Model

MemReader-4B, developed by MemTensor, is a 4 billion parameter causal language model derived from Qwen3-4B. It functions as a dedicated Memory Operator sub-model within the MemOS ecosystem, specifically engineered for efficient and accurate memory extraction from both conversational data and documents. This model is fine-tuned using a combination of human-annotated and model-generated data to optimize its performance in this specialized task.

Key Capabilities and Features

  • Optimized Memory Extraction: Excels at extracting high-quality conversation summaries and document snippet summaries.
  • Resource Efficiency: At 4 billion parameters, it supports local deployment on most machines, offering significant resource savings (over 70% compared to 14B models) while maintaining strong performance.
  • Performance: On the Locomo evaluation, MemReader-4B achieves an overall score of 0.7348, outperforming GPT-4o-mini (0.7253) in memory extraction tasks.
  • Multilingual Support: Capable of performing memory extraction in both English and Chinese.
  • Integration: Designed for seamless integration with MemOS, but also usable directly via Hugging Face Transformers, vLLM, or SGLang with preset templates.

Ideal Use Cases

  • Local-only AI Deployments: Suitable for environments with restricted internet access or where data privacy requires on-premise processing.
  • Cost-Sensitive Applications: Provides high performance for memory operations at a lower computational cost and higher speed.
  • Conversational AI and Document Processing: Enhances systems requiring precise and efficient summarization of interactions and textual content.