CultriX/MonaCeption-7B-SLERP-DPO
CultriX/MonaCeption-7B-SLERP-DPO is a 7 billion parameter language model developed by CultriX, fine-tuned using DPO (Direct Preference Optimization) on the CultriX/MonaCeption-7B-SLERP base model. This model was created through a SLERP merge of CultriX/MonaTrix-v4 and CultriX/MergeCeption-7B-v3, offering a 4096-token context window. It is designed for general language generation tasks, leveraging its merged architecture for balanced performance.
Loading preview...
CultriX/MonaCeption-7B-SLERP-DPO Overview
CultriX/MonaCeption-7B-SLERP-DPO is a 7 billion parameter language model developed by CultriX. This model is a result of a Direct Preference Optimization (DPO) finetuning process applied to the base model, CultriX/MonaCeption-7B-SLERP.
Merge Details
The foundational CultriX/MonaCeption-7B-SLERP model was constructed using the SLERP (Spherical Linear Interpolation) merge method via mergekit. It combines two distinct base models:
- CultriX/MonaTrix-v4
- CultriX/MergeCeption-7B-v3
The merge configuration specifically adjusted parameters for self-attention and MLP layers, indicating a tailored approach to blend the characteristics of the constituent models. The model supports a context length of 4096 tokens and is designed for general-purpose language tasks, benefiting from the DPO finetuning for improved alignment and response quality.