jondurbin/bagel-dpo-34b-v0.5

TEXT GENERATIONConcurrent Unit Cost:2Model Size:34BQuant:FP8Context Size:32kPublished:Apr 1, 2024License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

jondurbin/bagel-dpo-34b-v0.5 is a 34 billion parameter language model developed by jondurbin, fine-tuned from the Yi-34B-200K base model using Direct Preference Optimization (DPO). It features enhanced long-context support and is trained on a diverse array of datasets, including those for reasoning, coding, roleplay, and summarization. This model is particularly adept at complex instruction following, function calling, and context-obedient question answering, making it suitable for advanced conversational AI and specialized task execution.

Loading preview...

Model Overview

jondurbin/bagel-dpo-34b-v0.5 is a 34 billion parameter language model, a DPO fine-tune of the Yi-34B-200K base model, featuring improved long-context handling up to 32K tokens. It was trained using Direct Preference Optimization (DPO) on a highly diverse set of data sources, including academic benchmarks, synthetic instructions, roleplay data, and specialized datasets for coding, math, and summarization. A key aspect of its training involved using multiple prompt formats (Vicuna, Llama-2, Alpaca, and a modified ChatML) to enhance generalization across various instruction types.

Key Capabilities

  • Advanced Instruction Following: Trained with a variety of prompt formats to understand and execute complex instructions.
  • Context-Obedient Question Answering (RAG): Specifically tuned to answer questions strictly from provided context, minimizing hallucinations.
  • Function Calling: Supports two primary function-calling formats, enabling integration with external tools and APIs.
  • Specialized Prompting Strategies: Includes methods for chain-of-thought reasoning, reWOO-style function planning, novel writing, character card creation, conversational memory, boolean questions, SQL query generation, and emotion detection.
  • De-censorship: Includes a 'toxic-dpo' dataset for academic and lawful purposes, aimed at de-censoring responses.

Good For

  • Developers requiring a robust model for complex instruction execution and tool use.
  • Applications needing highly accurate, context-bound question answering (RAG).
  • Creative writing, roleplay, and generating structured outputs like character cards or novel chapters.
  • Tasks involving code generation, mathematical problem-solving, and summarization.