plawanrath/mistral-7b-instruct-v0.3-wanda-s30-pia
The plawanrath/mistral-7b-instruct-v0.3-wanda-s30-pia is a 7 billion parameter Mistral-7B-Instruct-v0.3 model that has undergone Wanda pruning with a 30% target sparsity. Developed by Plawan Kumar Rath, this model is a research artifact specifically created to study fairness degradation under weight pruning, demonstrating measurable bias amplification on the BBQ benchmark. It is not intended for production use due to its amplified bias and lack of storage or latency savings compared to the dense baseline.
Loading preview...
Overview
This model, plawanrath/mistral-7b-instruct-v0.3-wanda-s30-pia, is a 7 billion parameter variant of the Mistral-7B-Instruct-v0.3 base model, subjected to Wanda pruning with a 30% target sparsity. It is a research artifact only, developed by Plawan Kumar Rath and Rahul Maliakkal, to investigate the impact of weight pruning on fairness in large language models. The accompanying paper, "Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI" (IEEE AIIoT 2026), highlights that this specific pruning configuration leads to measurable bias amplification.
Key Characteristics
- Pruning Method: Utilizes
wanda(Activation-aware unstructured pruning) with a 30% target sparsity, achieving an actual sparsity of 19.60%. - Bias Amplification: The primary finding is that Wanda pruning at this sparsity level significantly amplifies bias, with a reported new-bias-emergence rate of 6.72% on the BBQ benchmark. This effect can be invisible to perplexity-based evaluations.
- No Performance Gains: Despite pruning, the model offers no storage savings (on-disk size is identical to the dense baseline as unstructured sparsity is not exploited by formats like SafeTensors or GGUF) and no latency savings (inference latency is identical on common hardware like Apple Silicon due to dense GEMM kernels).
- Research Focus: This model serves as a critical tool for studying the "Smart Pruning Paradox," where perplexity remains stable while bias dramatically increases.
Good for
- Academic Research: Ideal for researchers studying the effects of model compression techniques, particularly weight pruning, on fairness and bias in LLMs.
- Bias Analysis: Useful for experiments and analyses focused on understanding how pruning methods can inadvertently amplify societal biases.
- Educational Purposes: Can be used to demonstrate the complex trade-offs and unintended consequences of model optimization strategies.
Important Caveats
- Not for Production: This model is explicitly not recommended for production use in any user-facing or decision-making systems due to its demonstrated bias amplification.
- No Practical Benefits: Users seeking smaller model sizes or faster inference will not find these benefits with this specific pruned model.