shitshow123/mistral7b_sft_dpo

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jan 11, 2024License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The shitshow123/mistral7b_sft_dpo model is a 7 billion parameter language model based on the Mistral architecture. This model has been fine-tuned using Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) techniques. With an 8192-token context window, it is designed for general-purpose language generation tasks.

Loading preview...

Model Overview

The shitshow123/mistral7b_sft_dpo is a 7 billion parameter language model built upon the Mistral architecture. It has undergone a two-stage fine-tuning process: Supervised Fine-Tuning (SFT) followed by Direct Preference Optimization (DPO). This combination of training methodologies aims to enhance the model's ability to follow instructions and generate high-quality, preferred responses.

Key Characteristics

  • Architecture: Mistral 7B base model.
  • Parameter Count: 7 billion parameters, offering a balance between performance and computational efficiency.
  • Context Window: Supports an 8192-token context length, allowing for processing and generating longer sequences of text.
  • Fine-tuning: Utilizes both Supervised Fine-Tuning (SFT) for initial instruction following and Direct Preference Optimization (DPO) for aligning with human preferences.

Potential Use Cases

Given its architecture and fine-tuning approach, this model is suitable for a variety of general-purpose natural language processing tasks, including:

  • Text generation and completion.
  • Instruction following and conversational AI.
  • Summarization and content creation.

Limitations

As there is no detailed model card or specific benchmark information provided, users should perform their own evaluations to determine its suitability for specific applications. The model's performance characteristics and potential biases are not explicitly documented.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p