swadeshb/comem-qwen3-4b-regret-gate

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 17, 2026Architecture:Transformer Featherless Exclusive Cold

The swadeshb/comem-qwen3-4b-regret-gate is a 4 billion parameter Qwen3-based model, developed by swadeshb, specifically designed as a CoMem manager. It is trained for summary generation using GRPO with DeepSWE next-action preservation and bounded context length advantages. This model uniquely employs regret-weighted soft targets and action-balanced replay for its KEEP/SUM selector, making it suitable for tasks requiring intelligent content summarization and management within a constrained token budget.

Loading preview...

CoMem Qwen3-4B Regret-Gate Manager

This model is a merged, directly loadable checkpoint from an experimental 310-trajectory regret-weighted CoMem manager. It is based on YWZBrandon/summary-sft-qwen3-4b and features a 4 billion parameter Qwen3 architecture.

Key Capabilities & Training:

  • Summary Generation: Trained with GRPO (Generalized Reinforcement Learning with Policy Optimization) using DeepSWE next-action preservation.
  • Bounded Context Length Advantages: Optimized for scenarios with limited token budgets.
  • KEEP/SUM Selector: Utilizes regret-weighted soft targets and action-balanced replay for intelligent content selection.
  • Context Handling: Manages eight summaries when above an 8,192-token budget, and one virtual KEEP reference with seven summaries at or below budget.

Experimental Caveats:

  • Gate Evaluation: The current one-token, unrestricted text-generation gate evaluation may emit a third token instead of the intended KEEP or SUM verbalizers. For accurate gate measurement, it is recommended to compare next-token logits of KEEP and SUM or use constrained decoding.
  • Validation Status: This checkpoint is experimental and has not yet undergone full validation via a complete SWE-bench Verified harness run.

Good For:

  • Experimental research in regret-weighted content management and summarization.
  • Applications requiring intelligent selection and summarization of text within specific token constraints.
  • Developers interested in exploring advanced reinforcement learning techniques for language model control.