plawanrath/mistral-7b-instruct-v0.3-wanda-s50-pia

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kTool Calling:SupportedPublished:May 7, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

plawanrath/mistral-7b-instruct-v0.3-wanda-s50-pia is a 7 billion parameter Mistral-7B-Instruct-v0.3 model that has undergone 50% target sparsity Wanda pruning, achieving an actual sparsity of 32.69%. Developed by Plawan Kumar Rath and Rahul Maliakkal, this model serves as a research artifact to study bias amplification under weight pruning, demonstrating increased bias on the BBQ benchmark despite minimal perplexity change. It is not intended for production use due to its research focus on fairness degradation.

Loading preview...

Model Overview

This model, plawanrath/mistral-7b-instruct-v0.3-wanda-s50-pia, is a research artifact based on the mistralai/Mistral-7B-Instruct-v0.3 base model. It has been subjected to Wanda pruning with a target sparsity of 50%, resulting in an actual sparsity of 32.69% across its linear layers. The primary purpose of this model is to investigate the impact of weight pruning on bias amplification in large language models, as detailed in the paper "Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI" by Plawan Kumar Rath and Rahul Maliakkal.

Key Characteristics & Findings

  • Pruning Method: Wanda (Activation-aware unstructured pruning), applied to linear layers in transformer blocks.
  • Bias Amplification: The research demonstrates that Wanda pruning at this sparsity level significantly amplifies bias, with the Stereotype Reliance Score (SRS) increasing by 84.0% compared to the dense baseline, while perplexity only increased by 3.4%. This highlights a "Smart Pruning Paradox" where perplexity-based evaluations may not reveal bias degradation.
  • No Performance Gains: Despite pruning, the model exhibits no storage savings (due to unstructured sparsity not being exploited by common formats like SafeTensors/GGUF) and no inference latency savings on typical hardware (e.g., Apple Silicon, consumer GPUs) because dense GEMM kernels do not skip zero entries.
  • Research Focus: This model is explicitly a research artifact to study fairness degradation and is not recommended for production use due to its demonstrated bias amplification.

Important Caveats

  • Research Only: This model is for academic study of pruning effects, particularly bias.
  • Bias Risk: It has been shown to induce measurable bias amplification.
  • No Practical Benefits: It offers no practical benefits in terms of reduced size or faster inference compared to the dense base model.