Benchmarking the whole
writing workflow
One prompt for everyone (a transmigrator wakes as Runtu and trains in traditional martial arts), one set of follow-up constraints (1M characters / traditional martial arts / soul transmigration / hot-blooded power fantasy / late Qing to early Republic), and one identical process: talk through the requirements, produce the plan and beat sheets, write chapter one, stop. Three runs per model. Human editors then scored the chapters blind, out of 100.
One task cannot measure a model's writing ability
Drawing the requirements out of an author, designing a structure, sticking to the outline and writing good prose are four different abilities. Eighteen models make that plain: the four correlate weakly, and some of them actively pull against each other.
Requirement elicitation
Whether the Plan stage helps the author sharpen a vague idea into decisions that can be executed. What counts is whether the questions land, not how many there are. Out of 5.
Structural design
Whether the planning documents are complete, whether the outline holds together end to end, and whether the level of detail stays consistent across three runs. Out of 5.
Instruction following
Whether the prose follows the beat sheet the model itself wrote earlier. This says how controllable a model is, not how good the writing is. Out of 5.
Prose writing
Human editors read chapter one with no model names attached and score it out of 100, averaged over three runs. This is the axis a reader actually experiences.
No model wins at AI writing
Almost every axis orders the field differently. The default ordering here is the mean human blind score for prose — but the models that write well are not necessarily the ones that plan or follow instructions best.
The five blind-review batches ran at different times and may not have shared editors, so a single-digit gap is not a result — Step's 63.3 against Gemini's 63.0, for instance. 23 of the 53 valid chapters reached 60, andnot one reached 80。
Pick two models and see where each is strong
Everything is converted to a 100-point scale for display. Consistency is derived from the widest gap across the three runs — the smaller the gap, the higher the score. The raw scores are kept on the right.
Radial scale is the converted 0 — 100 score
excellent front-end work, prose that doesn't follow
Sonnet 5 tops elicitation at 4.57 and scores 4.6 on structure, yet only 57.3 on prose. GLM-5.2 tops structure at 4.7 and lands at 55.0 on prose. Both are excellent at the front end, and editors said of both that the prose “doesn't read like web fiction”.
The fewest files, and the best prose
Across three runs GPT-5.4 wrote just 7 files (a master plan, an outline of the first 20 chapters, and the prose) and has the field's lowest beat-sheet coverage at 1 of 3 — yet it scores 75 on prose, varying by only 4 points.Across these 54 runs, the number of files bears almost no relation to how good the prose is.
Elicitation scores were not recorded for Hy3, DeepSeek V4 Flash, Step 3.7 Flash and Mimo 2.5 Pro during testing, so the radar draws them as 0 for now.
How far apart are three runs on the same input
Each dot is one human blind score, the bar spans the three runs and the tick is the mean. The steadiest model varies by 4 points across its three runs; the least steady by 35. Hover a dot for the individual score.
Only four models keep their spread under 6 points
GPT-5.4 varies by 4 points, Grok 4.5 and Qwen3.7 Plus by 5, Kimi K2.6 by 6. Steady is not the same as useful, though: GPT-5.4 is steady and high, while the others are steady in the middle and lower range.
Occasionally good is not the same as reliably good
DeepSeek V4 Pro produced the single highest chapter in the field at 70 — and also a 35, a range of 35 points, the widest here. With a model like this, generate several versions and choose; do not expect any one run to be deliverable.
On writing, pricier really doesn't mean better
The x-axis is the bill for one run, the y-axis the mean human blind score. If pricier meant better, the dots would form a diagonal. They don't.
These are the amounts actually paid for the whole chapter-one workflow, planning tokens included — not just the cost of writing prose. Multiplying by chapter count gives a rough budget only, since later chapters do not repeat the planning.
Cheap models now clear 60 too
Hy3 (¥0.27 a run, 61.7) and Step 3.7 Flash (¥0.42 a run, 63.3) lift the cheap tier's prose ceiling to 63 — level with Gemini (63.0 at ¥2.18 a run) at an eighth to a fifth of the price. For churning out standardised first drafts at volume, overseas models no longer hold a clear value advantage.
Spending more tokens doesn't buy a better chapter
Grok 4.5 burns 973k tokens a run, 4.3× Hy3's 225k, and scores 8.4 points lower with human editors. At comparable unit prices,how many tokens a model burns per run matters more to long-run cost. That holds for all 18 models here.
Chinese models offer notably better value for writing — though the newer flagships are creeping up in price: the budget Chinese group runs ¥0.07 — ¥0.13, the premium Chinese group ¥0.29 — ¥0.41, the new flagships ¥0.41 — ¥1.09 and the overseas group ¥1.06 — ¥2.64. Kimi K3's ¥1.09 has already passed Gemini's ¥1.06.
Fast, steady and good is what actually counts
15 of 18 models ran all 54 runs without an error. As models iterate, tool-compatibility problems have become far rarer than they were. What is worth watching now is whether a model finishes the job steadily and efficiently.
Times marked ~ are range estimates. One LongCat 2.0 run was still unfinished past 26 minutes, so it has no comparable completion time.
Four things we noted across 54 runs
Different models, different batches — these kept happening anyway.
The four abilities are not one thing
Step 3.7 Flash comes last on both structure and instruction following, yet second in the field on prose; Mimo 2.5 Pro is creditable on both and last on prose. So pick a model per stage rather than crowning one overall winner.
Steady output, or the occasional flash
Only four models hold their three-run spread inside 6 points. DeepSeek V4 Pro (35), Kimi K3 (17) and Doubao (15) have all produced a high-scoring draft, but there is no guarantee the next run repeats it — generate several and pick.
Following the outline doesn't make it readable
Qwen3.7 Max scores 4.77 on instruction following and was still marked “no conflict”; Qwen 3.8 has the field's lowest at 3.17 and sits mid-table on prose. A high following score only means the model executed the design — not that the design was interesting.
models rush to deliver information
8 of the 18 lost points mainly for cramming later chapters into the first one. The opening pace breaks, and chapter two has nothing left to do.
Every model, every number
Click a column header in the table to sort by it.
The four models with N/A for elicitation kept no Plan transcript (0 of 3 coverage). A * on the following score means beat-sheet coverage was below 3 of 3, so the mean covers only the scorable runs and is less reliable than a fully covered model's.
Everything the 54 runs produced
Master plans, volume and chapter beat sheets, character sheets and the blind-reviewed first chapters — all kept in the folder structure each model built for itself.
Browse all materials →If you have to choose today, pair them like this
Choosing one model for planning and another for prose usually beats asking a single model to do everything.
Planning and front-end work
Co-creation and worldbuilding
First choice Sonnet 5. On a limited budget, switch to GLM-5.2 or Kimi K3. K3 writes the most complete set of files and is also the slowest and priciest, which suits building a story bible once.
the brief is vague and needs advice that doubles as constraint
Usable in the budget tier DeepSeek V4 Pro。
Serial reference, setups and indexes
First choice Qwen3.7 Max. Usable when budget is tight DeepSeek V4 Flash, ¥0.29 a run, and the tidiest set of files.
Prose production
Key chapters (the crucial first three, volume endings, big climaxes)
First choice GPT-5.4. On a limited budget, let Step 3.7 Flash Generate several versions and choose, checking facts, character relationships and chapter boundaries as you go.
Volume and direction checks
Also fine Hy3, ¥0.27 a run. Spell out in the prompta minimum length, a complete conflict and a closing hook, or it tends to write a first chapter so short it reads unfinished.
Unattended bulk drafts
Also fine Qwen 3.8, and all three test runs completed unattended. Before delivery, still check that dates, character ages and seasons line up, and that no stray English crept in.
we suggest building these habits
1 · Beat sheet first
Talk the design through in the Plan stage and save the beat sheet.
2 · Write to the beat sheet
In the Act stage, write prose from the saved beat sheet and nothing else.
3 · Fix structure before wording
If something is off, fix the structure and the beat sheet first and then rewrite, instead of endlessly editing the prose.
4 · Don't spend later chapters early
Use a hard constraint that forbids the model from pulling later plot into the current chapter.
Before switching models, summarise the key decisions so far into soloent.md, then open a fresh window.
Don't lean on the model too hard — people matter.
Editors' criticism of the prose clusters around three things:reaching for profundity, too few payoffs, and not reading like web fiction. No model's prose is deliverable without an editor today, and the author's own voice and judgement still matter most.
How this benchmark was run
Where the scores come from, and what they will not support.
The test
One shared input
One prompt (a transmigrator wakes as Runtu and trains in traditional martial arts), one set of follow-up constraints (1M characters / traditional martial arts / soul transmigration / hot-blooded power fantasy / late Qing to early Republic), one process: Plan → constraints → Act → stop once chapter one is on disk.
What people scored, what machines scored
All five groups of chapters were blind-scored out of 100 by human editors. GPT-5.6 handled the documented non-blind scoring. All five groups used one framework: four abilities, three elicitation criteria and five beat-sheet compliance criteria.
How costs were calculated
All costs are amounts actually paid through SoloEnt's back end. The original budget group used Lingxie discount pricing (25 — 50% of list), the original premium and overseas groups used real platform billing, and the two new batches' dollar spend was converted at ¥6.80/$.
Unit price, consumption and bill are three different things
Unit price says how expensive the model is, tokens per run say how restrained it is, and the bill is the product of the two. The bill is what you budget from; on its own it says nothing about efficiency.
Design flaws to be honest about
Small sample, single genre
Three runs per model, all of them late-Qing martial-arts power fantasy. Urban, mystery or other genres could come out differently.
This scores a workflow, not prose in isolation
Each model wrote prose from its own outline, so what is scored is the output of a whole workflow — not the same outline handed to different models as a pure writing contest.
Across batches, read tiers not places
With this many models we ran five batches; blind review happened at different times and possibly with different editors, so a single-digit gap is not a reliable result. Machine scoring was not blind, though the last two batches were adjusted slightly.
Some data is missing
The supplementary batches did not keep full Plan transcripts, so their elicitation is not scored. Kimi K3, Hy3, GPT-5.4, Doubao, MiniMax and LongCat also lack a scorable beat sheet in all three runs.
Appendix: known error, and the tests still to run
Where the cost data is imprecise
- Both new batches were converted at ¥6.80/$, but the discounts differ (Kimi K3 at 90% of list, Qwen 3.8 at 40%, no discount recorded for Grok), which introduces error when comparing unit prices side by side.
- The cost table does not break out input, output and cached tokens, so the numbers cannot be recomputed from list prices alone.
Other things worth knowing
- Chapter length varies enormously, from Hy3's 1,927 characters on average to Step 3.7 Flash's 5,866. Some files include self-checks, summaries or several merged chapters, so length is no proxy for quality.
- Step 3.7 Flash's first run lost its event stream partway through, so the intermediate interactions and any errors are invisible — “no errors” is only what can be confirmed.
- The old MiMo and Mimo 2.5 Pro are different versions; three clean runs only mean the old problem did not recur.
What's next
- Rerun the supplementary batches, keeping full Plan transcripts and the answers to a standard set of follow-up questions, so elicitation can be scored.
- Compare Kimi K3 / K2.6 and DeepSeek V4 Flash / Pro across urban, xuanhuan and mystery genres, to see whether the newer versions' flat prose is specific to the martial-arts genre.
- Test separately which intermediate files actually improve the prose, focusing on Grok 4.5 at ~973k tokens a run against Hy3 at ~225k.
- Add more blind-review editors, and keep model names, file structures and prices out of their view.
Join the review panel
This issue's model scoring was supplemented by human editors reviewing blind. We are still recruiting readers and editors who know web fiction — if you read a lot and are willing to write considered comments, come and try.
Join the panel →What should we benchmark next?
Tell us which genre, which model or which stage of the workflow you want tested.