Zhihu-ai/Zhi-Create-DSR1-14B
Zhihu-ai/Zhi-Create-DSR1-14B is a 14.8 billion parameter language model fine-tuned by Zhihu-ai based on DeepSeek-R1-Distill-Qwen-14B, specifically optimized for enhanced creative writing capabilities. It achieves an 8.33 score on the LLM Creative Story-Writing Benchmark and 8.46 on WritingBench, demonstrating improved performance over its base model. The model also shows modest improvements in knowledge, reasoning, and mathematical tasks, making it suitable for general-purpose applications requiring strong creative text generation.
Loading preview...
Model Overview
Zhi-Create-DSR1-14B is a 14.8 billion parameter language model developed by Zhihu-ai, fine-tuned from DeepSeek-R1-Distill-Qwen-14B. Its primary focus is on significantly enhancing creative writing abilities, while also showing improvements in general capabilities.
Key Capabilities & Performance
- Creative Writing: Achieved a score of 8.33 on the LLM Creative Story-Writing Benchmark (up from 7.87) and 8.46 on WritingBench (up from 7.93), indicating substantial improvements in creative text generation across various domains like Literature & Art, and Advertising & Marketing.
- General Reasoning: Demonstrates modest improvements of 2%–5% in knowledge and reasoning tasks (CMMLU, MMLU-Pro).
- Mathematical Reasoning: Shows encouraging progress in mathematical benchmarks such as AIME-2024, AIME-2025, and GSM8K.
- Instruction Following: Improved performance on the ifeval benchmark, from 71.43 to 74.71.
Training Methodology
The model was trained using a combination of Supervised Fine-tuning (SFT) with a curriculum learning strategy and Direct Preference Optimization (DPO). The training data includes rigorously filtered open-source datasets, chain-of-thought reasoning corpora, and curated question-answer pairs from Zhihu, with quality assurance via a Reward Model (RM) filtering pipeline.
Recommended Usage
For optimal performance, it is recommended to set the generation temperature between 0.5-0.7 (0.6 is ideal) and to enforce the model to start its response with "\n" for thorough reasoning, similar to DeepSeek-R1 series models.