ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated is a 7.6 billion parameter instruction-tuned causal language model, developed by ApolloRaines, based on the Qwen2.5-Coder-7B-Instruct architecture. This model features a 'Jbliterated' v2 processing pipeline designed for cleaner output and improved instruction-following across various scenarios. It is specifically optimized to run efficiently on GPUs with limited memory by streaming layers across GPU, RAM, and disk, offering up to 4x faster performance than AirLLM.

Loading preview...

Model Overview

ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated is a 7.6 billion parameter instruction-tuned model, building upon the Qwen/Qwen2.5-Coder-7B-Instruct base. This version, designated 'v2', incorporates a unique 'Jbliteration' process, which refers to a specialized multi-phase processing pipeline designed to enhance output quality and instruction adherence.

Key Enhancements and Features

This v2 iteration introduces several significant improvements:

  • Improved Output Quality: Features a refined multi-phase processing pipeline that aims for cleaner and more precise output.
  • Enhanced Refusal Handling: Utilizes a more precise geometric decomposition of the refusal subspace, ensuring consistent treatment of different framings of the same topic without exhibiting fake compliance.
  • Robust Instruction Following: Demonstrates coherent and reliable instruction-following capabilities across a wide range of tested scenarios.
  • Memory-Efficient Execution: Designed to run on GPUs with insufficient memory by streaming layers across GPU, RAM, and disk, leveraging the DeepswapLLM framework. This allows for full precision (bfloat16) operation without quantization, achieving up to 4x faster performance compared to AirLLM.

Technical Details

The model's 'Jbliteration' process involves modifications across all transformer layers, maintaining a base dtype of bfloat16. It is derived from the robust Qwen/Qwen2.5-Coder-7B-Instruct model.

Ideal Use Cases

This model is particularly well-suited for developers and researchers who require:

  • A 7.6B parameter instruction-following model with enhanced output clarity and consistent behavior.
  • Efficient execution of large models on hardware with limited GPU memory, without compromising on precision.
  • Applications demanding reliable instruction adherence and unbiased handling of diverse prompts.