Featherless blog hero: Best AI for creative writing, seven open models for fiction and long-form prose

Two years ago, the answer to which AI to write with was only closed-source models. That is no longer true. Open-weight models now sit inside the top twenty of the community-run EQ-Bench Creative Writing v3 leaderboard, hold a whole manuscript in context, and carry licences that let you publish what comes out without asking anyone.

This covers open models for prose rather than writing products. If you want a manuscript-aware editor with a story bible and revision history, Sudowrite and Novelcrafter are good at that. What follows is the layer underneath: seven models, verified against their own cards for parameters, context and licence, and scored against named creative-writing evaluations where those exist.

What separates the best LLM for writing from a good general model

The benchmarks that sell models measure something else. GPQA Diamond, Terminal-Bench and SWE-bench tell you whether a model can reason about physics or repair a repository. None tells you whether it can hold a narrator’s voice through chapter six, or whether it will start every third paragraph with a participial phrase.

Five axes matter for long-form prose, and none of them appear on a model card.

  • Voice persistence: whether the register a model sets in the opening pages survives 30,000 words.
  • Instruction adherence at depth: whether a style note given at turn one still binds at turn sixty.
  • Slop rate: how often the model reaches for phrases language models overuse. The EQ-Bench maintainers measure this directly, and it separates models more reliably than any rubric score.
  • Refusal behaviour: fiction contains violence and sex, and a model tuned to decline those will stall a draft.
  • Context window: how much of your manuscript the model can see at once.

Size predicts none of these well. Google’s Gemma 4 31B, a dense 30.7B model, posts a higher creative-writing rubric score on EQ-Bench v3 than Mistral Large 3 at 675B total parameters. A 30.7B model beating a 675B one on prose is the reason to shortlist on measured writing scores rather than on parameter counts.

Top 7 open models for creative writing in 2026

Kimi K3: the most capable open model for holding a plot across a novella

Moonshot AI released Kimi K3 on 16 July 2026 as a sparse mixture-of-experts transformer with 2.8T total parameters and 104B active per token, using what the card calls Kimi Delta Attention with attention residuals. Native context is 1,048,576 tokens, and Featherless serves it at 256K. It is the largest open-weight model anyone serves at production scale, and self-hosting it is not realistic for an individual writer.

For prose, K3’s value is continuity rather than sentence-level polish. A million-token window carries a 300,000-word manuscript plus notes with room to draft, removing the summarise-and-reload cycle that degrades voice in every long project. Its predecessor Kimi K2.6 holds the highest creative-writing rubric score of any open-weight model on EQ-Bench v3 at 16.67, ahead of GLM-5.2 at 16.44 and DeepSeek-V4-Pro at 16.45. K3 itself is unscored there.

Kimi K3 ships under the Kimi K3 License, a modified MIT licence with an attribution requirement at very large commercial deployment scale. For a novelist or a small press the practical effect matches MIT: you may use the outputs commercially and owe nobody a share.

Best use cases: novel drafting with the full manuscript in context, multi-book series continuity, long-running interactive fiction, developmental editing over a complete draft, screenplay work needing scene consistency, and research-heavy historical fiction.

Performance benchmarks: K3 scores 93.5 on GPQA Diamond and 88.3 on Terminal-Bench 2.1 on its own card, which establishes general capability rather than writing quality. The family’s creative evidence is K2.6’s 16.67 rubric and Elo 1715.5 on EQ-Bench Creative Writing v3, a community-run benchmark judged by another language model. Treat K3’s own creative score as unmeasured.

GLM-5.3: the open line with the strongest creative-writing record behind it

Z.ai’s GLM-5.3 is a 753B-parameter mixture of experts using Deep Sparse Attention, built on the same base as GLM-5.2, and although the card demonstrates up to 1M tokens, Featherless serves it at 256K in FP8. Z.ai’s own headline claim is coding rather than prose, at a 50% improvement over GLM-5.2 on that axis, but the family’s prose is what earns it a place here.

The GLM line has the best creative-writing record among open weights. GLM-5.2 is the highest-ranked open-weight model on EQ-Bench Creative Writing v3 by Elo at 1745.8, fifteenth overall and directly behind GPT-5.2, which puts the family’s prose within reach of the closed mid tier. GLM-5.3 is unscored on that board, so the evidence for it is lineage plus third-party ranking: modelgrep’s September 2026 writing table puts GLM-5.3 first among Z.ai models at 44.9 intelligence, ahead of GLM-5.3-Flash at 41.9 and GLM-5.2 at 34.0.

There is a cheaper sibling worth knowing about. GLM-5.3-Flash is 321B with 18B active, served at the same 256K, and it costs $0.15 per 1M input, $0.03 cached and $0.50 output against GLM-5.3’s $1.40, $0.26 and $4.40. On a full day of drafting that is the difference between about $9 a month and about $82. Draft on Flash, then run the passes that matter through GLM-5.3.

GLM-5.3 ships under a custom GLM-5.3 licence rather than the MIT terms the earlier GLM releases carried, but the substance is MIT with one clause bolted on. Commercial use, fine-tuning and derivative works are all permitted, the copyright notice has to travel with copies, and the weights come as-is with no claim on what you generate. The addition binds only a model-as-a-service business turning over more than $10 billion in any twelve consecutive months, which has to clear commercial deployment with Z.ai first. For a novelist, or for a company building a writing product at any realistic scale, it reads as MIT.

Best use cases: daily drafting at high volume, dialogue-heavy scenes, iterative revision where you regenerate a paragraph twenty times, short fiction, worldbuilding documents, and testing many prompt variants on a budget.

Performance benchmarks: GLM-5.2 scores 16.44 rubric with an Elo of 1745.8 on EQ-Bench Creative Writing v3, the best open-weight Elo on that board. GLM-5.3’s own card reports Terminal Bench 2.1 at 88.2 and DeepSWE v1.1 at 66.9, both coding rather than writing measures. No creative-writing score exists for GLM-5.3 or GLM-5.3-Flash as of this writing.

DeepSeek-V4.1-Flash: a million tokens of manuscript at a small model’s compute cost

DeepSeek released V4.1-Flash in 2026 with an unusual shape: 763B total parameters split into a 552B backbone and a 196B Engram conditional-memory component, arranged as a causal encoder-decoder of 20 encoder and 20 decoder layers. Only 8B parameters are active during prefill and 16B during decode, with six of 384 routed experts firing per token. Native context is 1 million tokens, and Featherless serves it at 256K.

That architecture targets the problem long-form writers have. Prefill is the expensive half of a drafting session, because every turn re-reads the manuscript, and an 8B prefill path makes re-reading a long document cheap in a way dense models cannot match. The V4 family scores 16.45 (Pro) and 16.29 (Flash) on the EQ-Bench creative rubric, both mid-pack, so the prose is competent rather than distinguished. What you are buying is the ability to keep everything in view.

DeepSeek-V4.1-Flash is released under the MIT licence, the same terms as the rest of the line, with no restriction on commercial use of outputs.

Best use cases: whole-manuscript revision passes, continuity checking across a series, long non-fiction with heavy source material in context, retrieval-free work on large document sets, and cost-sensitive drafting at very long context.

Performance benchmarks: the card reports 74.2% resolved on DeepSWE v1.1, 90.6% pass@1 on Terminal-Bench 2.1 and 74.1% on MMLU-Pro. On creative writing the family’s measured points are DeepSeek-V4-Pro at 16.45 rubric and Elo 1553.1, and DeepSeek-V4-Flash at 16.29 and Elo 1563.7, both on EQ-Bench Creative Writing v3. V4.1-Flash is unscored there.

Qwen3.8-27B: the dense model that fits one GPU and still writes long

Alibaba released Qwen3.8-27B in August 2026 as a dense 27B causal language model with a vision encoder, interleaving gated DeltaNet and gated attention blocks across 64 layers. Native context is 262,144 tokens, extensible to 1 million, and Featherless serves it at 256K. Dense matters here: every parameter fires on every token rather than a routed subset, and at 27B it runs on a single high-memory GPU, so self-hosting is available to a writer who wants it.

Qwen models have a reputation in creative work for instruction adherence, the axis most writers feel first. Tell a Qwen model the narrator never uses similes and it tends to still be obeying at turn fifty. The family’s measured creative point is Qwen3.5-397B-A17B at 16.00 rubric and Elo 1486.1 on EQ-Bench v3, mid-pack; against that, Qwen3-235B-A22B-Instruct-2507 tops the normalised Creative Writing v3 table on llm-stats at 0.875. The two boards disagree, which is worth knowing before you trust either.

Qwen3.8-27B is released under the Apache 2.0 licence. Commercial use, modification, redistribution and fine-tuning are permitted, with no acceptable-use policy bolted on.

Best use cases: style-constrained drafting, pastiche and voice imitation, local self-hosted setups, fine-tuning on your own corpus, translating your own fiction, and editing passes with strict rules.

Performance benchmarks: the model card reports 61.7% on SWE-bench Pro, 89.2% on GPQA Diamond and 90.3% on LiveCodeBench v6. On creative writing, the nearest measured family member is Qwen3.5-397B-A17B at 16.00 rubric and Elo 1486.1 on EQ-Bench Creative Writing v3. Qwen3.8-27B has not been scored on that board.

Gemma 4 31B: the small model whose prose scores like a large one

Google published the Gemma 4 family on 30 July 2026. The 31B is a dense 30.7B model interleaving local sliding-window attention with full global attention, at 256K native context, taking text and image input, though Featherless serves it at 32K. Alongside it sit a 26B A4B mixture-of-experts variant with 3.8B active, a 12B unified model, and two E-series models at 4.5B and 2.3B effective parameters.

This is the entry that breaks the size heuristic. Gemma 4 31B scores 16.01 on the EQ-Bench Creative Writing v3 rubric, level with Qwen3.5-397B-A17B at 16.00 despite being one-thirteenth the size, and ahead of Mistral Large 3 at 675B on 15.49. Its Elo of 1407.6 is lower, reflecting how pairwise voting weighs consistency, and that split between rubric and Elo is the clearest illustration in this guide of why one number never settles the question. The practical read is that Gemma 4 31B produces individually good paragraphs and is less reliable across a long run. The 32K served window points the same way: that is roughly 24,000 words of manuscript, so this is a scene-level model here rather than a whole-draft one, whatever its card says.

Gemma 4 is released under the Apache 2.0 licence, a change from the bespoke Gemma Terms of Use that governed earlier generations. That matters if you were avoiding Gemma over its use-restriction clauses.

Best use cases: short fiction and flash, poetry, scene-level drafting, style exercises, local setups on consumer hardware, and line-level rewriting where per-paragraph quality matters more than chapter-scale consistency.

Performance benchmarks: Gemma 4 31B scores 16.01 rubric and Elo 1407.6 on EQ-Bench Creative Writing v3, against 16.03 and 1286.9 for the 26B A4B and 15.71 and 1270.4 for the 12B. Card figures are 85.2% on MMLU Pro, 89.2% on AIME 2026 and 80.0% on LiveCodeBench v6.

Mistral Medium 3.5: the strongest European-language option you can actually call

Mistral released Medium 3.5 on 31 March 2026 as a 128B transformer, and Featherless serves it at $1.00 per 1M input, $0.20 cached and $4.00 output. It replaces Mistral Large 3 in this guide for two reasons: it is four months newer, and Large 3 is not in the Featherless catalogue at all, so a guide that told you to call it would be sending you somewhere else.

What it is for is writing that is not in English. Mistral’s card lists dozens of supported languages and leads on the European set: French, Spanish, German, Italian, Portuguese and Dutch, alongside Chinese, Japanese, Korean and Arabic. For a novelist drafting in French or editing a German translation of their own work, that is a different proposition from an English-first model with multilingual coverage bolted on, and it is the reason this entry exists rather than a claim about English prose quality.

On English prose the family’s measured point is unflattering and worth stating. Mistral Large 3 scores 15.49 rubric with an Elo of 1409.8 on EQ-Bench Creative Writing v3, forty-ninth overall and below Gemma 4 31B on both measures despite being more than five times the size. Medium 3.5 is unscored there. If your work is in English, the models above this entry are better buys.

Mistral Medium 3.5 is released under a Modified MIT licence, which permits commercial use of the model and places no claim on what you generate. Mistral’s earlier Medium releases were API-only, so the open weights here are the change worth noting.

Best use cases: drafting in French, Spanish, German, Italian, Portuguese or Dutch, translating your own work, multilingual dialogue, bilingual editing, and any project where a permissively licensed European model is a procurement requirement.

Performance benchmarks: no creative-writing score exists for Medium 3.5 on EQ-Bench or LMArena. The nearest measured family member is Mistral Large 3 at 15.49 rubric and Elo 1409.8. On context, the model is natively 256K and Featherless serves it at 125K, so it holds roughly 94,000 words at once: enough for a whole novel, but not for a novel plus notes.

Apertus v1.5-70B: the only entry whose training data you can audit

The Swiss AI Initiative’s Apertus v1.5-70B is a dense 70B transformer decoder using the xIELU activation and the AdEMAMix optimiser, at 262,144 tokens of native context, though Featherless serves it at 32K. Its distinguishing property is not architecture. Apertus is trained exclusively on fully open data, and the project publishes weights, data, training details and evaluation protocol together.

For most writers that is a footnote. For some it is the whole decision. Where provenance is contested, by a publisher’s AI disclosure policy, an academic press, or your own position on scraped fiction, Apertus is the only model here that can answer what it was trained on with a document. The cost is capability and context: no entry on any creative-writing leaderboard, a technical report still pending, general evaluation behind every other model here, and a 32K served window that rules out the whole-manuscript work its 262K card figure would suggest. The bucket is adequate rather than strong, and expect to do more editing.

Apertus v1.5-70B is released under the Apache 2.0 licence, covering the weights and, unusually, the training data and recipe alongside them.

Best use cases: work subject to an AI-provenance disclosure, academic and institutional writing, drafting where training-data ethics are a stated position, European data-residency requirements, and fine-tuning where you want to know the base distribution.

Performance benchmarks: Apertus v1.5-70B has no published creative-writing score on EQ-Bench or LMArena, and its technical report has not been released. The model card names evaluations across MMLU, MMLU-Pro, ARC-Challenge, GSM8K, HumanEval and Global-MMLU without publishing the numbers. Treat any capability claim about it as unverified.

Additional notable models

  • MiniMax M3: 428B with 23B active and a 1M context window, under the minimax-community licence. Its predecessor M2.5 scores 15.18 on EQ-Bench v3, so treat the prose as unproven.
  • Kimi K2.6: the highest open-weight rubric score at 16.67, and worth using directly if K3’s cost is a problem.
  • GLM-5.2: the open-weight Elo leader at 1745.8, and the safest pick if you want a measured result rather than a newer model.
  • Qwen3.5-397B-A17B: 16.00 rubric and Elo 1486.1, mid-pack on both.
  • Llama 4 Maverick 17B: 10.5 rubric and Elo 942.4, which puts the Llama line out of contention for prose. Llama 3.1 405B at 10.89 does not rescue it.

Comparing the models: the numbers side by side

ModelParametersTypeContext, native / served hereLicenceCost class
Kimi K32.8T (104B active)MoE1M / 256KKimi K3 License (modified MIT)High
GLM-5.3753BMoE with DSA1M / 256KCustom GLM-5.3 licenceMedium
DeepSeek-V4.1-Flash763B (8B prefill / 16B decode active)MoE encoder-decoder1M / 256KMITLow
Qwen3.8-27B27BDense262K / 256KApache 2.0Low
Gemma 4 31B30.7BDense256K / 32KApache 2.0Very low
Mistral Medium 3.5128BDense transformer256K / 125KModified MITMedium
Apertus v1.5-70B70BDense262K / 32KApache 2.0Low

Creative-writing rubric scores are absent from this table on purpose. Five of the seven have no score on any creative benchmark, and putting a predecessor’s number in a row labelled with the successor’s name is how these tables mislead. The scores that exist sit in each entry above, attached to the model actually measured. Read the context column as two numbers rather than one: what the model can do and what Featherless serves it at are different figures, and on Gemma 4 31B and Apertus the served window is an eighth of the card’s.

Getting started without managing your own GPUs

The models that write best are the ones you cannot run locally. Kimi K3’s weights are measured in terabytes and Moonshot’s deployment guidance assumes a supernode. DeepSeek-V4.1-Flash has the same problem despite its small active path, because the whole 763B parameter set still has to be resident. Nobody compares seven of these in an afternoon by downloading seven sets of weights.

An inference API removes that. Featherless serves 40,000+ open models on one OpenAI-compatible endpoint at https://api.featherless.ai/v1, prices by architecture class rather than per model so new fine-tunes are available the day they land, and keeps no logs of prompts or completions, which matters if what you are drafting is unpublished.

from openai import OpenAI

client = OpenAI(base_url="https://api.featherless.ai/v1", api_key="YOUR_API_KEY")

response = client.chat.completions.create(
    model="zai-org/GLM-5.3",
    messages=[
        {"role": "system", "content": "Continue the scene in the established voice. No similes."},
        {"role": "user", "content": open("chapter-06.txt").read()},
    ],
    temperature=0.9,
)
print(response.choices[0].message.content)

The model string is the only thing that changes between candidates. Swap zai-org/GLM-5.3 for moonshotai/Kimi-K3 or Qwen/Qwen3.8-27B and the rest of the script, including your prompt and evaluation harness, stays identical. That is what makes a real comparison practical: one loop over a list of model IDs.

Temperature is worth tuning deliberately. At 0.9 most of these models produce usable variation between generations, which is what you want when sampling ten continuations of one scene. Below about 0.7 they converge on their median style, the flattening that makes model-written prose recognisable.

On plans, the fit depends on how you work rather than how much you write. Chat at $25 a month is unlimited tokens with context up to 32K and four concurrent units. For a writer in a long interactive drafting session that is the shape that fits, because paying per token on the Developer plan is what makes a long session expensive, and Chat does not count tokens at all.

Price a realistic day against it: sixty turns, 30K tokens of manuscript and notes in context each time, 800 words out per turn. That is 1.8M input and 48K output tokens a day, or 54M input and 1.44M output a month.

RouteMonthly cost
GLM-5.3-Flash on the Developer plan~$9
Chat plan, flat$25
Kimi K3 on the Developer plan, 87% cache hit~$43
GLM-5.3 on the Developer plan~$82
Kimi K3 on the Developer plan, no caching~$122

The Developer plan wins comfortably on flash-class models. A writer shipping an app, or needing the full 256K window for a whole-manuscript pass, needs the Developer plan at $50 in credits a month. Tokenomics 101 covers the underlying levers.

Where open weights are not the answer for writers

Open weights lose at the top of the range, and the gap is measurable. On EQ-Bench Creative Writing v3, Claude Fable 5 sits at an Elo of 2156.3 and GPT-5.6 Sol at 2208, against 1745.8 for GLM-5.2, the best-placed open-weight model. That spread of roughly 410 Elo points is not a rounding error on a pairwise board. If your standard is the best prose available at any price, the answer today is a closed model, and Claude Fable 5.1 at $10.00 per 1M input and $50.00 output is what that costs.

Purpose-built writing products win on a different axis. Sudowrite and Novelcrafter rank on the first page for creative-writing queries because they wrap a model in what a novelist needs: a story bible that stays in sync, per-chapter context assembly, revision history, and an editor rather than a chat box. This guide assumes you will build that layer yourself or work without it.

Featherless is also not always the cheapest route, and the comparison has to be made per model rather than per provider. The flat $25 plan has no equivalent at Together, Fireworks, DeepInfra, Groq, Nebius or Parasail. On per-token pricing, Kimi K3 is a favourable case here at $2.00 per 1M input and $10.00 output against Moonshot’s own published $3.00 and $15.00, but that margin is specific to this model. If your workload is one model at high volume and nothing else, price that model directly. Our LLM API pricing comparison lays out the per-model picture across providers.

Commonly asked questions

Is it legal to publish a novel written by AI? In the United States, yes, and separately the AI-generated portions are not copyrightable. The Copyright Office position is that human authorship is required for protection, so a manuscript you drafted with model assistance and then substantially rewrote is protectable in the parts you wrote. Publishers increasingly require disclosure, so check the rules of the market you submit to.

Which AI is better than ChatGPT for writing? On measured rubric and Elo, several Claude models sit above the GPT line, and among open weights GLM-5.2 and Kimi K2.6 are within striking distance of GPT-5.2. “Better” depends on register, and a model that beats ChatGPT on literary fiction may lose on commercial thriller pacing.

What is the best LLM for creative writing right now? Among open weights with a measured score, GLM-5.2 leads on Elo at 1745.8 and Kimi K2.6 on rubric at 16.67. Among newer models with no creative score, Kimi K3 and GLM-5.3 are the strongest candidates on lineage. Test them on a scene you have already written before committing.

What is the best AI for fiction writing if I need long chapters? Kimi K3 and DeepSeek-V4.1-Flash, both at 1M native context and both served at 256K here, because chapter-scale consistency is mostly a context problem. Degradation across a long generation is measured on the EQ-Bench longform board, which scores eight consecutive 1,000-word chapters.

Can I run these models on my own computer? Gemma 4 31B and Qwen3.8-27B, yes on 24GB to 48GB of VRAM at reasonable quantisation. Apertus v1.5-70B, with effort. The rest, no: GLM-5.3, DeepSeek-V4.1-Flash and Kimi K3 are all datacentre-scale, and Mistral Medium 3.5 at 128B needs more than one card.

How much context do I need for a novel? A 90,000-word manuscript is roughly 120K tokens, so holding a whole novel in context needs a 128K window at minimum and 256K to leave room for drafting. Working chapter by chapter with a synopsis, 32K is workable and is what most writers do. Check the served window rather than the card: Kimi K3, GLM-5.3, DeepSeek-V4.1-Flash and Qwen3.8-27B all run at 256K here, Mistral Medium 3.5 at 125K, and Gemma 4 31B and Apertus v1.5-70B at 32K, which makes those two chapter-by-chapter models whatever their cards claim.

Do open models refuse to write dark or violent scenes? It varies sharply by model, and the pattern tracks safety tuning rather than size or recency. Heavily tuned models return hedged summaries of a violent scene rather than writing it, while models with lighter tuning handle the same prompt without comment. Test this early on a scene you have already written, because it is cheaper to find out now than at chapter twenty. If refusal behaviour is your primary constraint rather than prose quality, our roundup of uncensored models answers that question.

Does an open licence mean I own what the model writes? The licence governs the weights rather than the outputs. MIT, Apache 2.0, the Kimi K3 License and the GLM-5.3 licence all place no claim on what you generate. Copyright in the output is a separate question, answered above, and does not depend on which licence the model carries.

What does it cost to draft a novel with one of these? At sixty turns a day with 30K tokens of context, a month of drafting costs about $9 on the Developer plan with GLM-5.3-Flash, about $82 on GLM-5.3, about $43 on Kimi K3 at an 87% cache-hit rate, or $25 flat on Chat with no token counting and a 32K ceiling. The variable that moves the bill is context length rather than words produced, because every turn re-reads everything.

Choosing the model you will actually finish a draft with

Pick one scene you have already written well, give three candidates the 3,000 words before it and the instruction you would give yourself, and read the continuations a week later without knowing which model produced which. The one you would keep is your model.

For a starting shortlist: GLM-5.3 if you want the best-evidenced open line for prose, GLM-5.3-Flash if cost per drafting hour matters and you work chapter by chapter, Kimi K3 if continuity across a whole manuscript matters most, Gemma 4 31B if you want good paragraphs on your own hardware, Mistral Medium 3.5 if you write in a European language other than English, and Apertus v1.5-70B if provenance is a requirement rather than a preference. Do not pick on size: Gemma 4 31B outscores Mistral Large 3 at more than twenty times its parameter count, and the Llama line sits at 10.5 rubric and Elo 942.4, out of contention for prose entirely.

To test four of them without provisioning any hardware, the Chat plan at $25 a month gives unlimited drafting at 32K context across the whole catalogue, and Developer gives the 256K window and the API access a whole-manuscript pass needs. If you are building a writing product on these models rather than writing with them, talk to us about dedicated capacity before your per-token bill grows faster than your readership.

Last updated: September 17, 2026. Model availability, licences and benchmark figures change; verify the licence for any model you ship on.

Start building under 3 minutes