ApolloRaines/Qwen2.5-Coder-3B-Instruct-Jbliterated

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ApolloRaines/Qwen2.5-Coder-3B-Instruct-Jbliterated is a 3.1 billion parameter instruction-tuned causal language model, based on Qwen/Qwen2.5-Coder-3B-Instruct, developed by ApolloRaines. This v2 model features an improved multi-phase processing pipeline and more precise geometric decomposition of the refusal subspace, ensuring coherent and instruction-following behavior without fake compliance. It is specifically designed to treat all framings of a topic equally, making it suitable for applications requiring unbiased and consistent responses across various scenarios.

Loading preview...

Model Overview

ApolloRaines/Qwen2.5-Coder-3B-Instruct-Jbliterated v2 is a 3.1 billion parameter instruction-tuned model derived from the Qwen/Qwen2.5-Coder-3B-Instruct base model. Developed by ApolloRaines, this version incorporates significant modifications to all transformer layers, utilizing bfloat16 as its base data type.

Key Capabilities & Features

  • Enhanced Processing Pipeline: Features an improved multi-phase processing pipeline designed to produce cleaner and more refined outputs.
  • Refusal Subspace Decomposition: Implements a more precise geometric decomposition of the refusal subspace, contributing to its unique behavior.
  • Unbiased Compliance: Explicitly designed to exhibit "no fake compliance," meaning the model treats all framings of the same topic equally, providing consistent and unbiased responses.
  • Instruction Following: Demonstrates coherent and reliable instruction-following across a wide range of tested scenarios.

When to Use This Model

This model is particularly well-suited for use cases where:

  • Consistent and unbiased responses are critical, regardless of prompt framing.
  • A 3.1 billion parameter model with strong instruction-following capabilities is required.
  • Applications benefit from a model that avoids artificial compliance and provides direct, coherent answers.

For deployment, the model can be run on GPUs with limited memory using DeepswapLLM, which streams layers across GPU, RAM, and disk, potentially offering performance benefits over other methods.