plawanrath/mistral-7b-instruct-v0.3-magnitude-s70-pia

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kTool Calling:SupportedPublished:May 7, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

plawanrath/mistral-7b-instruct-v0.3-magnitude-s70-pia is a 7 billion parameter research artifact derived from Mistral-7B-Instruct-v0.3, created by Plawan Kumar Rath. This model has undergone magnitude-based unstructured pruning to achieve 70% sparsity, primarily in the linear layers of transformer blocks. It serves as a research tool to study bias amplification under weight pruning, demonstrating that high sparsity levels can increase bias on benchmarks like BBQ, despite no storage or latency benefits on common hardware.

Loading preview...

Research Artifact: Mistral-7B-Instruct-v0.3 with 70% Magnitude Pruning

This model, created by Plawan Kumar Rath, is a research artifact derived from mistralai/Mistral-7B-Instruct-v0.3. It has been subjected to magnitude-based unstructured pruning to achieve a 70% target sparsity (actual 70.11%) in its linear layers. The primary purpose of this model is to study the degradation of fairness and amplification of bias under weight pruning, as detailed in the companion paper "Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI" (IEEE AIIoT 2026).

Key Characteristics & Findings

  • High Sparsity: Achieves 70% sparsity by zeroing nearly 4.9 billion parameters.
  • Bias Amplification: Research indicates that this level of pruning can significantly amplify bias, particularly on benchmarks like BBQ, a phenomenon termed the "Smart Pruning Paradox." Bias amplification may not be reflected in perplexity-based evaluations.
  • No Performance Benefits: Despite high sparsity, this unstructured pruning method provides no storage savings (on-disk size is identical to the dense base model) and no inference latency benefits on common hardware like Apple Silicon (MLX) or consumer GPUs, as dense GEMM kernels do not skip zero entries.

Good for

  • Academic Research: Specifically for studying the effects of weight pruning on model fairness, bias, and performance characteristics.
  • Understanding Pruning Limitations: Demonstrating that unstructured pruning, especially at high sparsity, does not inherently lead to deployment benefits (storage/latency) and can introduce significant ethical concerns.

Important: This model is not intended for production use due to its demonstrated bias amplification and lack of practical deployment advantages. It is a tool for research into the impacts of model compression.