jondurbin/bagel-dpo-34b-v0.5
jondurbin/bagel-dpo-34b-v0.5 is a 34 billion parameter language model developed by jondurbin, fine-tuned from the Yi-34B-200K base model using Direct Preference Optimization (DPO). It features enhanced long-context support and is trained on a diverse array of datasets, including those for reasoning, coding, roleplay, and summarization. This model is particularly adept at complex instruction following, function calling, and context-obedient question answering, making it suitable for advanced conversational AI and specialized task execution.
Loading preview...
Model Overview
jondurbin/bagel-dpo-34b-v0.5 is a 34 billion parameter language model, a DPO fine-tune of the Yi-34B-200K base model, featuring improved long-context handling up to 32K tokens. It was trained using Direct Preference Optimization (DPO) on a highly diverse set of data sources, including academic benchmarks, synthetic instructions, roleplay data, and specialized datasets for coding, math, and summarization. A key aspect of its training involved using multiple prompt formats (Vicuna, Llama-2, Alpaca, and a modified ChatML) to enhance generalization across various instruction types.
Key Capabilities
- Advanced Instruction Following: Trained with a variety of prompt formats to understand and execute complex instructions.
- Context-Obedient Question Answering (RAG): Specifically tuned to answer questions strictly from provided context, minimizing hallucinations.
- Function Calling: Supports two primary function-calling formats, enabling integration with external tools and APIs.
- Specialized Prompting Strategies: Includes methods for chain-of-thought reasoning, reWOO-style function planning, novel writing, character card creation, conversational memory, boolean questions, SQL query generation, and emotion detection.
- De-censorship: Includes a 'toxic-dpo' dataset for academic and lawful purposes, aimed at de-censoring responses.
Good For
- Developers requiring a robust model for complex instruction execution and tool use.
- Applications needing highly accurate, context-bound question answering (RAG).
- Creative writing, roleplay, and generating structured outputs like character cards or novel chapters.
- Tasks involving code generation, mathematical problem-solving, and summarization.