Cutyp/dolphin-8b-merged

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Oct 1, 2026License:llama3Architecture:Transformer0.0K Featherless Exclusive Cold

Cutyp/dolphin-8b-merged is an 8 billion parameter causal language model, created by Cutyp, resulting from the merge of dphn/dolphin-2.9-llama3-8b and Cutyp/dolphin-8b-qlora. This model operates without requiring adapter code, offering a streamlined deployment. It is primarily designed for general-purpose language generation tasks, leveraging its merged architecture for enhanced performance.

Loading preview...

Overview

Cutyp/dolphin-8b-merged is an 8 billion parameter language model, created by Cutyp, that combines the base model dphn/dolphin-2.9-llama3-8b with the Cutyp/dolphin-8b-qlora adapter. This merging process was performed using PEFT scaling with an alpha / r ratio of 32 / 16 = 2.0, resulting in a standalone model that does not require separate adapter code for inference.

Key Capabilities

  • Simplified Deployment: The model is fully merged, allowing direct loading and use with standard Hugging Face transformers library calls without needing to manage adapter configurations.
  • Integrated Architecture: Combines a robust base model with a QLoRA adapter targeting key projection layers (q/k/v/o/gate/up/down_proj) for potentially improved performance and efficiency.
  • Standard Inference: Compatible with AutoModelForCausalLM and AutoTokenizer for straightforward integration into existing workflows.

Good for

  • General-purpose text generation: Suitable for a wide range of language tasks where an 8B parameter model is appropriate.
  • Developers seeking ease of use: Ideal for users who prefer a pre-merged model to avoid the complexities of managing and loading separate adapters.
  • Experimentation with merged models: Provides a practical example of a QLoRA adapter merged into a base model for direct deployment.