ledgernova/sn120-40d715805d87
The ledgernova/sn120-40d715805d87 model is a 35.1 billion parameter language model, developed by ledgernova, that has undergone LoRA-merged Online-DPO fine-tuning. This model is derived from the `marsplan0624/affine-5gedzafcvg-queen` base and is optimized for improved performance as indicated by a positive margin against a local n80 baseline. With a context length of 32768 tokens, it is suitable for tasks requiring extensive contextual understanding and generation.
Loading preview...
Model Overview
The ledgernova/sn120-40d715805d87 model is a 35.1 billion parameter language model, developed by ledgernova, that has been fine-tuned using a LoRA-merged Online-DPO (Direct Preference Optimization) approach. This model is based on the marsplan0624/affine-5gedzafcvg-queen architecture and incorporates specific DPO parameters including a beta of 0.1, alpha of 32, rank of 16, gradient accumulation steps of 4, a learning rate of 5e-6, and a temperature of 1.2, trained for 300 maximum steps with a forced prefix.
Key Capabilities
- Online-DPO Fine-tuning: Utilizes an Online-DPO method for preference alignment, indicating a focus on generating responses that align with desired criteria.
- Performance Improvement: Demonstrates a positive performance margin of +0.01182 with a z-score of 2.445 against a local n80 baseline, suggesting an improvement in its capabilities.
- Extended Context Window: Supports a context length of 32768 tokens, enabling the processing and generation of longer and more complex texts.
Good For
- Applications requiring a model with a large parameter count and an extended context window.
- Use cases where a model fine-tuned with Direct Preference Optimization is beneficial for generating high-quality, aligned outputs.
- Scenarios demanding a model that has shown measurable performance improvements over its baseline.