sasa2000/Qwen3-Swallow-8B-RL-v0.2-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Feb 23, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

sasa2000/Qwen3-Swallow-8B-RL-v0.2-heretic is an 8 billion parameter, Qwen3-based large language model, developed by sasa2000, that has been decensored from the original tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2 using Heretic v1.2.0. This model is optimized for bilingual Japanese-English proficiency, maintaining strong performance in math and coding tasks, and features enhanced reasoning capabilities. It is particularly suited for applications requiring less refusal behavior compared to its original counterpart.

Loading preview...

Model Overview

This model, sasa2000/Qwen3-Swallow-8B-RL-v0.2-heretic, is an 8 billion parameter large language model derived from tokyotech-llm/Qwen3-Swallow-8B-RL-v0.2. It has been processed using Heretic v1.2.0 to reduce refusal rates, making it a "decensored" version. The original Qwen3-Swallow models, developed by TokyoTech-LLM, are built as bilingual Japanese-English models through Continual Pre-Training (CPT), Supervised Fine-Tuning (SFT), and Reinforcement Learning with Verifiable Rewards (RLVR) based on the Qwen3 architecture.

Key Capabilities

  • Decensored Behavior: Significantly reduced refusal rates (31/100 compared to 97/100 for the original model), offering more direct responses.
  • Bilingual Proficiency: Highly optimized for both Japanese and English language tasks.
  • Retained STEM Performance: Successfully maintains strong performance in mathematics and coding, preventing catastrophic forgetting during fine-tuning.
  • Enhanced Reasoning: Achieves reasoning performance on par with, and in some tasks, surpassing the original Qwen3 models.
  • Long Context Window: Supports a context length of up to 32,768 tokens.

Good For

  • Applications requiring a less restrictive or decensored language model for general chat and instruction following.
  • Tasks demanding high proficiency in both Japanese and English, including translation and bilingual content generation.
  • Use cases involving mathematical problem-solving and code generation where reasoning capabilities are crucial.
  • Developers looking for a model that has undergone Reinforcement Learning (RLVR) for improved performance and reasoning.