v000000/L3.1-Niitorm-8B-DPO-t0.0001
The v000000/L3.1-Niitorm-8B-DPO-t0.0001 model is an 8 billion parameter language model based on the Llama 3.1 architecture, developed by v000000. It is a DPO-trained merge of Llama 3.1-Niitama-v1.1 and Llama 3.1-Storm-8B, specifically fine-tuned on the Gutenberg DPO dataset. This model is optimized for generating human-like prose and story writing, aiming to reduce synthetic-feeling outputs.
Loading preview...
Overview
v000000/L3.1-Niitorm-8B-DPO-t0.0001 is an 8 billion parameter language model built upon the Llama 3.1 architecture. It is a DPO-trained (Direct Preference Optimization) model, resulting from a nearswap merge of two base models: Sao10K/L3.1-8B-Niitama-v1.1 (with an abliteration LoRA) and akjindal53244/Llama-3.1-Storm-8B. The merged model was then fine-tuned for one epoch on the jondurbin/gutenberg-dpo-v0.1 dataset.
Key Capabilities
- Enhanced Prose Generation: The DPO training on the Gutenberg dataset is specifically designed to produce more human-like prose and story writing, significantly reducing synthetic output characteristics.
- Roleplay (RP) Model: Built on a base known for roleplay capabilities, this model is further refined for such applications.
- Optimized DPO Training: Utilizes a higher learning rate and full dataset during DPO training compared to previous iterations, leading to better adaptation to the desired writing style.
Performance
Evaluations on the Open LLM Leaderboard show an average score of 27.89. Specific metrics include:
- IFEval (0-Shot): 76.89
- BBH (3-Shot): 30.51
- MMLU-PRO (5-shot): 31.85
Good for
- Creative Writing: Generating narratives, stories, and descriptive prose.
- Roleplay Scenarios: Creating engaging and natural-sounding dialogue and character interactions.
- Reducing AI-Generated Feel: Producing text that feels less robotic or artificial.
Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.