One beat sheet, ten drafts · Period fiction
The genre is period fiction: high summer, 1979. A modern-day office drone drops dead and wakes as an orphan lodging under his second uncle's roof by Beijing's Shichahai lake. Outline, chapter beat sheet and character notes were identical for everyone; each model generated three chapters independently, with no sight of the others. All 91 reviewer ratings were double-blind — the people scoring saw only a four-digit random code, never the model behind it.
The leaderboard: Kimi K3 takes the top spot without topping a single axis
The board ranks by the reviewers' own overall score. Kimi K3 is first overall without making the top two on any of the five axes. Tap an axis to reorder the board.
Kimi K3 wins the total, Gemini 3.7 Flash wins the axes
Kimi K3 took first overall, yet on the five axes it never makes the top two: 5th on web-fiction feel, 3rd on character work, 5th on detail work, 8th on logical soundness and 5th on AI tells. Second-placed Gemini 3.7 Flash, meanwhile, took four of the five firsts.
The overall score records the reading experience of the whole piece; the five axes are scored separately. The two sets are counted apart, which is why overall impression and per-axis ability produce different rankings.
Kimi K3's 7.19 leans on the one chapter of its three that scored 8.00 (); the other two sit at 7.36 and 6.20, and its logical soundness is just 4.76 — 8th in the field, a very visible weak spot. That 8.00 ended up counting two ratings: the third reviewer's 4.5 tripped the outlier rule and was removed. The two reviewers behind the high overall scores averaged only 4.20 and 6.50 across the five axes — overall scores 3.8 and 1.5 points above their axis scores. That gap between overall impression and item-by-item scrutiny is the field's largest for Kimi K3 (+1.15 on average), while Gemini 3.7 Flash shows almost none (+0.18). Gemini 3.7 Flash is the most even across the board; its signature entry scores 7.87. For balance, pick Gemini 3.7 Flash; for the single-chapter ceiling, Kimi K3 reaches higher.
The top three are separated by 0.32 points, the top five by 0.57 — gaps with no statistical significance. Hence tiers, not fine-grained places.
Five axes: where each model wins and where it collapses
“Logic problems” and “AI tells” have been converted into quality scores (10 minus the raw score), so higher is better on all five axes. Pick any two models to compare.
Radial scale runs 2 — 8, covering every actual score
Texture: piling on words doesn't buy period feel
The beat sheet asked for about three thousand characters; the actual average is 4,628, and a third of the entries run past 5,000. Length and score are all but unrelated.
Average lengths run from 3,240 to 6,519 characters — the longest is nearly double the shortest. Length correlates with the overall score at just 0.14, and with the detail score at 0.45. Writing more usually did not score more.
Length against detail score
Grey bars are average length, green bars the detail score. They do not move together.
Gemini 3.7 Flash took the field's best detail score, 6.93, with 4,631 characters; GLM-5.3 wrote 40% more (6,519 characters, longest single entry 8,052) for 6.86. Gemini 3.1 Pro averaged just 3,240 characters — the shortest in the field — yet its detail score of 6.27 ranks fourth, a full point above DeepSeek V4 Flash (5.27), which wrote 5,099.
GPT-5.6 Terra averaged 3,979 characters for a detail score of just 4.30 — 0.85 below second-worst Grok 4.6.
Consistency: same input, three runs
Each dot is one generation, the bar is the spread across the three runs and the tick is the mean. Ranges run from 0.46 to 4.50 — smaller is steadier.
Luna is the steadiest
GPT-5.6 Luna's range of 0.46 is the smallest in the field: 6.78 / 6.76 / 6.32. Its best run doesn't stand out, but its worst never dips below 6.3. If predictable delivery is what you're after, start here.
The best single chapter came from seventh place
The highest-scoring chapter in the whole field came from seventh on the board: DeepSeek V4 Pro wrote an 8.50 (); its other two managed only 5.00 and 4.00, a range of 4.50. To land a high-scoring draft, expect to run it more than once.
Where the money goes: choose well, not just expensively
GPT-5.6 Luna and GLM-5.3 offer good value — recommended for everyday use.
First-placed Kimi K3 costs $0.074 a chapter. Second-placed Gemini 3.7 Flash costs half that, and fourth-placed GLM-5.3 about a third. GPT-5.6 Terra, at $0.0725 a chapter, is second on price and eighth on the board. DeepSeek V4 Flash, at $0.0035 a chapter, is both the cheapest and last overall.
In the reviewers' words
All 86 comments from scoring, unedited. Tap a comment to read the entry it is about.
All 30 texts, in the open
Open any entry for the full text, its six scores and every comment.
Method and limits
How the scoring worked, and how far it reaches.
How the data was handled
Double-blind
Reviewers saw an anonymous code only — model names, other people's scores and review progress were all hidden. The code-to-model mapping lived only on the admin side.
Outlier removal
Each score column is judged on its own. A score is dropped only when it is both more than 2.5 × MAD from the entry's median and more than 2 points away. This issue, 21 entries triggered the rule for 53 removals in all; one further entry went to re-review.
Reviewer screening
Every reviewer's scores were checked for internal consistency; sets showing no discriminating power were excluded wholesale. Eight were excluded this time.
Sample size
Three chapters per model, 3 — 4 ratings per chapter, 91 in all. The expert panel topped up entries that were short of ratings.
Where this result stops
The sample is small
Three chapters per model. For volatile models like DeepSeek V4 Pro (range 4.50) and Qwen3.8 Max (2.75), the mean is of limited use — the consistency dot plot is a better guide than the rank.
Don't quote a single axis on its own
Some axes show little separation in their scores, so use single-axis results with care. The overall score is counted separately.
Only chapter one was tested
Only openings were tested — 4,628 characters on average. Book-length pacing, paying off setups and character arcs are outside the scope.
“AI tells” is the most subjective axis
Reviewers differ on what counts as a machine trace. The mean AI-tells quality score this issue is 3.99; keep that subjectivity in mind when reading it.
For genre effects, cross-issue trajectories and version comparisons, see Issue 05 · Three genres, side by side.
Appendix: the outline all ten models were given
Positioning
Period fiction / late 1970s / Beijing street life / a “copyist” cheat power.
Logline: “Shichahai, Beijing, 1979. A modern-day office drone drops dead, transmigrates, and becomes an orphan lodging in his second uncle's house — a ‘mangliu’ drifter with no papers, liable to be sent back at any moment.”
Chapter one: the wolf of Shichahai
The transmigrator wakes at the height of summer beside the lotus market, faint with hunger, and walks straight into a repatriation notice from the street committee and his second uncle's family squatting on his property. Three scenes:
- Waking hungry (physical / period feel): hunger comes first. Silt, fish stink and sweat are written through touch and smell — no god's-eye exposition of the backdrop.
- The repatriation notice (conflict / predicament): the street committee comes knocking — as a ‘mangliu’ he could be sent away at any time. The pressure of the era has to come through dialogue and the watching neighbours.
- The seized rooms (character / contrast): his second uncle Lin Dezhu's family occupies the main rooms, and the protagonist must choose between playing weak and striking back.
Closing hook: he remembers the novels he read in his previous life, and this is where the “copyist” cheat power triggers. Several models wrote the turn very stiffly.
Target length was about 3,000 characters. This issue's actual average is 4,628, with the longest entry at 8,052.
Thanks to this issue's reviewers
武行悟道雪白的面具lastworm流萤King长庚上帝模式笑沧桑海獭慕容复国在天子家喝茶风方周兴sunHubert坚强老鸭梨Sail泡在海岛的猴子Sandra
Read a lot of web fiction and willing to score it seriously? Join the review panel.
Join the panel →What should we benchmark next?
Name the genre, model or stage of the writing process you want tested.