plawanrath/qwen2.5-7b-instruct-bf16-mlx-cba

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The plawanrath/qwen2.5-7b-instruct-bf16-mlx-cba model is an MLX-format BF16 (uncompressed baseline) variant of the Qwen2.5-7B-Instruct model, developed by Plawan Kumar Rath. With 7.6 billion parameters and a 32768-token context length, this model serves as a reference artifact for studying bias emergence in compressed LLMs. It is specifically designed for research into how quantization aggressiveness impacts stereotypical behavior, particularly on fairness-sensitive tasks, and is optimized for use on Apple Silicon via the MLX framework.

Loading preview...

Overview

This model, plawanrath/qwen2.5-7b-instruct-bf16-mlx-cba, is an MLX-format BF16 (uncompressed baseline) variant of the Qwen2.5-7B-Instruct model. It was created by Plawan Kumar Rath and Rahul Maliakkal as one of 15 model artifacts for their paper, "Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels." This specific version serves as the uncompressed reference for their research.

Key Capabilities

  • Uncompressed Baseline: Represents the full BF16 precision of the Qwen2.5-7B-Instruct model, making it suitable as a reference for comparison with quantized versions.
  • MLX Optimized: Pre-converted to the MLX format, enabling direct loading and efficient inference on Apple Silicon devices without additional conversion steps.
  • Research Artifact: Crucial for studying the impact of quantization on model bias, particularly in fairness-sensitive tasks, as detailed in the associated research paper.

Good For

  • Bias Research: Ideal for researchers investigating how quantization affects model alignment and emergent stereotypical behaviors, especially when analyzing the "dose-response" relationship between quantization aggressiveness and bias.
  • MLX Development: Developers working with Apple Silicon can leverage this model for applications requiring a high-fidelity, unquantized Qwen2.5-7B-Instruct variant.
  • Performance Benchmarking: Useful for establishing baseline performance and quality metrics before applying various quantization techniques, particularly in contexts where bias and fairness are critical considerations.