chartreuse-verte/prose-rewriter-4b-v1.2

Hugging Face
TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 21, 2026License:agpl-3.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

The chartreuse-verte/prose-rewriter-4b-v1.2 is a 4 billion parameter Qwen3-4B-Base derivative, specifically designed as a paragraph-level prose rewriter. It takes LLM-generated text and re-renders it to sound more human while preserving original semantics. This model excels at rewriting fictional prose in English, offering modes to match, inflate, or compress text length, and supports a 32768 token context length.

Loading preview...

Model Overview

The chartreuse-verte/prose-rewriter-4b-v1.2 is a specialized 4 billion parameter model built upon Qwen/Qwen3-4B-Base, enhanced with a rank-32 LoRA. Its core function is to act as a paragraph-level prose rewriter, transforming text generated by large language models into more human-like prose while strictly maintaining the original meaning. This model is the larger of two variants, offering improved performance over its 1.7B counterpart, particularly in invention rate and copy excess.

Key Capabilities

  • Prose Rewriting: Rewrites LLM-generated text to sound more natural and human.
  • Semantic Preservation: Designed to preserve the original semantics of the input text.
  • Length Control: Supports three distinct modes via the edit block: match (rewrite in place), inflate (cut from padded input), and compress (expand from shortened input). The match mode is strongly recommended for general rewriting without length modification.
  • Optimized for Fictional Prose: Specifically trained on human-written fictional prose, including content from r/WritingPrompts and AO3.
  • Robustness: Exhibits low self-repetition and consistent performance across various input lengths, with a floor of approximately 15 words before invention increases.
  • VRAM Efficiency: Provides GGUF quantized versions (Q8_0, Q4_K_M) for efficient deployment on consumer GPUs, with detailed VRAM usage metrics provided.

When to Use This Model

  • Humanizing LLM Output: Ideal for post-processing text generated by other LLMs to improve its naturalness and readability.
  • Creative Writing Assistance: Useful for refining fictional narratives, dialogue, and descriptive passages.
  • Specific Length Adjustments: When you need to rewrite a paragraph while maintaining its original length, or subtly adjust it (cut/expand) based on the edit mode.

Limitations

  • Not an Instruct Model: It has a single, specialized function and does not respond to general instructions.
  • Fictional Prose Only: Not suitable for technical documentation or non-narrative text.
  • Single Paragraph Input: Designed for one paragraph per call; longer inputs should be split.
  • English Only: Primarily functions with English text in a narrative register.
  • AI Detector Bypass: Will not bypass AI detection tools as it preserves underlying word choices and sentence structures.
  • Short Input Behavior: Inputs under ~15 words may lead to padding and invention rather than direct rewriting.