Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-Concise-R

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jun 7, 2024License:llama3Architecture:Transformer0.0K Featherless Exclusive Cold

Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-Concise-R is an 8 billion parameter language model from Salesforce, based on the LLaMA 3 architecture. This model is a concise version of SFR-Iterative-DPO-LLaMA-3-8B-R, specifically trained with a concise penalty during iterative DPO. It is designed for research purposes where brevity and conciseness in generated text are prioritized.

Loading preview...

Model Overview

Salesforce/LLaMA-3-8B-SFR-Iterative-DPO-Concise-R is an 8 billion parameter language model developed by Salesforce, built upon the LLaMA 3 architecture. This model is a specialized variant of the Salesforce/SFR-Iterative-DPO-LLaMA-3-8B-R, distinguished by its training methodology. During its iterative Direct Preference Optimization (DPO) process, a specific "concise penalty" was applied.

Key Characteristics

  • Conciseness Optimization: The primary differentiator of this model is its explicit training to generate more concise outputs, achieved through a unique penalty applied during the DPO phase.
  • Research Focus: This model is released strictly for research purposes, as indicated by its ethics disclaimer, supporting academic exploration into conciseness in language models.
  • LLaMA 3 Base: Leverages the foundational capabilities and architecture of the LLaMA 3 8B model.

Intended Use

This model is particularly suited for:

  • Academic Research: Investigating the impact of conciseness penalties in DPO training and their effects on model behavior.
  • Prototyping: Exploring applications where shorter, more direct responses are preferred, provided the research context aligns with the ethical guidelines.

Users are strongly advised to review the provided ethics disclaimer and acceptable use policies, especially for high-risk scenarios, as the model has not been evaluated for all downstream purposes.