swadeshb/comem-qwen3-4b-regret-gate
TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026Architecture:Transformer Featherless Exclusive Cold
The swadeshb/comem-qwen3-4b-regret-gate is a 4 billion parameter Qwen3-based model, developed by swadeshb, specifically designed as a CoMem manager. It is trained for summary generation using GRPO with DeepSWE next-action preservation and bounded context length advantages. This model uniquely employs regret-weighted soft targets and action-balanced replay for its KEEP/SUM selector, making it suitable for tasks requiring intelligent content summarization and management within a constrained token budget.
Loading preview...
CoMem Qwen3-4B Regret-Gate Manager
This model is a merged, directly loadable checkpoint from an experimental 310-trajectory regret-weighted CoMem manager. It is based on YWZBrandon/summary-sft-qwen3-4b and features a 4 billion parameter Qwen3 architecture.
Key Capabilities & Training:
- Summary Generation: Trained with GRPO (Generalized Reinforcement Learning with Policy Optimization) using DeepSWE next-action preservation.
- Bounded Context Length Advantages: Optimized for scenarios with limited token budgets.
- KEEP/SUM Selector: Utilizes regret-weighted soft targets and action-balanced replay for intelligent content selection.
- Context Handling: Manages eight summaries when above an 8,192-token budget, and one virtual KEEP reference with seven summaries at or below budget.
Experimental Caveats:
- Gate Evaluation: The current one-token, unrestricted text-generation gate evaluation may emit a third token instead of the intended
KEEPorSUMverbalizers. For accurate gate measurement, it is recommended to compare next-token logits ofKEEPandSUMor use constrained decoding. - Validation Status: This checkpoint is experimental and has not yet undergone full validation via a complete SWE-bench Verified harness run.
Good For:
- Experimental research in regret-weighted content management and summarization.
- Applications requiring intelligent selection and summarization of text within specific token constraints.
- Developers interested in exploring advanced reinforcement learning techniques for language model control.