Zhihu-ai/Zhi-Create-DSR1-14B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 19, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Zhihu-ai/Zhi-Create-DSR1-14B is a 14.8 billion parameter language model fine-tuned by Zhihu-ai based on DeepSeek-R1-Distill-Qwen-14B, specifically optimized for enhanced creative writing capabilities. It achieves an 8.33 score on the LLM Creative Story-Writing Benchmark and 8.46 on WritingBench, demonstrating improved performance over its base model. The model also shows modest improvements in knowledge, reasoning, and mathematical tasks, making it suitable for general-purpose applications requiring strong creative text generation.

Loading preview...

Model Overview

Zhi-Create-DSR1-14B is a 14.8 billion parameter language model developed by Zhihu-ai, fine-tuned from DeepSeek-R1-Distill-Qwen-14B. Its primary focus is on significantly enhancing creative writing abilities, while also showing improvements in general capabilities.

Key Capabilities & Performance

  • Creative Writing: Achieved a score of 8.33 on the LLM Creative Story-Writing Benchmark (up from 7.87) and 8.46 on WritingBench (up from 7.93), indicating substantial improvements in creative text generation across various domains like Literature & Art, and Advertising & Marketing.
  • General Reasoning: Demonstrates modest improvements of 2%–5% in knowledge and reasoning tasks (CMMLU, MMLU-Pro).
  • Mathematical Reasoning: Shows encouraging progress in mathematical benchmarks such as AIME-2024, AIME-2025, and GSM8K.
  • Instruction Following: Improved performance on the ifeval benchmark, from 71.43 to 74.71.

Training Methodology

The model was trained using a combination of Supervised Fine-tuning (SFT) with a curriculum learning strategy and Direct Preference Optimization (DPO). The training data includes rigorously filtered open-source datasets, chain-of-thought reasoning corpora, and curated question-answer pairs from Zhihu, with quality assurance via a Reward Model (RM) filtering pipeline.

Recommended Usage

For optimal performance, it is recommended to set the generation temperature between 0.5-0.7 (0.6 is ideal) and to enforce the model to start its response with "\n" for thorough reasoning, similar to DeepSeek-R1 series models.