5456es/last_layer_prune_Llama-3.2-3B-Instruct_prune_0.5-sigmoid

TEXT GENERATIONPricing:Input $0.2036 / Output $1.34Concurrent Unit Cost:1Model Size:3.2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 15, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The 5456es/last_layer_prune_Llama-3.2-3B-Instruct_prune_0.5-sigmoid is a 3.2 billion parameter Llama-3.2-3B-Instruct model fine-tuned using Direct Preference Optimization (DPO) with a last-layer pruning method. This model is designed for instruction-following tasks, leveraging DPO on preference data to enhance its performance. It features a substantial 32768-token context length, making it suitable for applications requiring extensive contextual understanding and generation.

Loading preview...

Model Overview

This model, last_layer_prune_Llama-3.2-3B-Instruct_prune_0.5-sigmoid, is a 3.2 billion parameter instruction-tuned variant of the Llama-3.2-3B-Instruct base model. It has been fine-tuned using Direct Preference Optimization (DPO), specifically employing a "last" method for training and incorporating pruning during this process. The model is designed to excel in instruction-following tasks, benefiting from its DPO-based training on preference data.

Key Characteristics

  • Base Model: Llama-3.2-3B-Instruct
  • Fine-tuning Method: Direct Preference Optimization (DPO) with a "last" method.
  • Pruning: Pruning was applied during the training process.
  • Context Length: Supports a context length of 32768 tokens.

Use Cases

This model is particularly suited for applications requiring a compact yet capable instruction-following language model. Its DPO fine-tuning suggests improved alignment with human preferences, making it potentially useful for:

  • Generating responses based on specific instructions.
  • Tasks where preference-based alignment is beneficial.

Limitations

Users should be aware that this model inherits the limitations of its Llama-3.2-3B-Instruct base model. Additionally, the pruning process may introduce further limitations, which should be considered during deployment and evaluation.