One beat sheet, ten drafts · High-concept
The genre is high-concept: a failing high-school senior stumbles into a WeChat group called the “Azure Dragon Study Group”, whose members are Newton, Einstein and Gauss from different parallel universes. Outline, beat sheet and character notes were identical for everyone; each model generated three chapters independently, with no sight of the others. All 98 reviewer ratings were double-blind — the people scoring saw only a four-digit random code, never the model behind it.
The leaderboard: Gemini 3.1 Pro in a class of its own
The ranking follows the reviewers' own overall score. Tap an axis to reorder the board.
Gemini 3.1 Pro leads the runner-up by 0.66 — the widest gap at the top in three issues. Its range across three chapters is a mere 0.31, also the steadiest in the field. The highest-scoring entry is (7.64 pts).
Five axes: where each model wins and where it collapses
“Logic problems” and “AI tells” have been converted into quality scores (10 minus the raw score) — on all five axes, higher is better. Pick any two models to compare.
Radial scale runs 2 to 8, covering every actual score
AI tells: where high-concept gets hit hardest
The ten models' median is just 3.95, the lowest of the five axes. An absurd premise slips all too easily into “explain the premise → marvel at it → end on an uplift”.
这个一定是国产大模型,满眼的动作+破折号+总结,罗列人名凑字数,AI味太冲了。This one: No.
AI自己发挥的地方逻辑很混乱,尤其是鬼使神差这个词用的直接消解掉了主角的所有动机纯粹沦为推进叙事而进行的动作。This one: No.
ai味太重,然后牛顿和伽利略的说话太文言了……他们不是外国人么?主角的吐槽太过刻意,像个ai。可以这样说:他想写出网文感,但是不会写。This one: No.
The low scores name concrete problems: dashes and a summing-up voice, motives bent to keep the plot moving, strained imitation of web-fiction tone. GPT-5.6 Luna's took the field's highest overall score, 8.33; Luna also has the best mean on AI tells.
Consistency: DeepSeek rewards rerolling
Each dot is one generation, the bar is the spread across the three runs and the tick is the mean. Ranges run from 0.60 to 3.67 — smaller is steadier.
Steady can mean steadily low
GPT-5.6 Terra's range of 0.60 is the smallest in the field, but all three chapters sit at 5.0 to 5.6. Simply rerunning it will not fix much.
Only one is both high and steady
Gemini 3.1 Pro is 1st on the mean with a range of 0.31 — every delivery lands on the same line. Kimi K3's best chapter, 8.00 () would place second in the field; its worst, 4.33 () is dead last. Those two chapters are its ceiling and its floor, in plain sight.
Where the money goes: 21× the price, 1.5× the quality
GPT-5.6 Luna and GLM-5.3 are the value picks — the ones to recommend for everyday use.
Fifth-placed GLM-5.3 costs $0.0233 a chapter; Kimi K3, at 3.2× the price, scores just 0.20 higher, and GPT-5.6 Terra, at 3.1× the price, actually scores 0.75 lower. The priciest model, Kimi K3 ($0.074), ranks 4th; the second-priciest, GPT-5.6 Terra ($0.0725), ranks 8th. Paying more does not reliably buy a better score.
In the reviewers' words
Collected here: 91 verbatim comments from scoring, unedited. Tap a comment to read the entry it scored.
This issue used a human-written beat sheet from a published novel; the previous two issues used AI-written ones.
A human beat sheet leaves more blank space, and reviewers split two ways on what that did.
“Credit where it's due — the human-written outline is actually pretty good.” The blank space gave the models room to work, and the drafts came out more finished.
The blank space also magnified the models' logic problems: “wherever the AI improvises, the logic falls apart”.
All 30 texts, in the open
Open any entry for the full text, its six scores and every comment.
Method and limits
How the scoring worked, and where it applies.
How the data was handled
Where the beat sheet came from
This issue's beat sheet comes from a published novel and was written by a human editor. The previous two issues used AI-written beat sheets.
Double-blind
Reviewers saw an anonymous code and nothing else — model names, other people's scores and review progress were all hidden. The code-to-model mapping existed only on the admin side.
Outlier removal
Judged column by column. A score is dropped only when it is both more than 2.5 × MAD from that entry's median and more than 2 points away in absolute terms. This issue 23 entries triggered the rule, removing 48 values in all.
Reviewer screening
Every reviewer's scores were checked for internal consistency; those showing no discriminating power were excluded wholesale. Eight sets were excluded this time.
Sample size
Three chapters per model, 3 to 4 ratings per entry, 98 in all. The expert group topped up entries that were short of ratings.
Where this result stops
The sample is small
Three chapters per model. For the high-variance models (Kimi K3 at 3.67, GPT-5.6 Luna at 3.11) the mean carries limited weight — the consistency dot plot is a better guide than the rank.
Don't quote a single axis on its own
Some axes show too little separation in the scores, so treat single-axis results with care. The overall score is tallied separately and unaffected.
Only chapter one was tested
Only the opening was tested, at an actual average of 3,912 characters. Book-length pacing, paying off setups and character arcs are outside the scope.
“AI tells” is the most subjective axis
Reviewers differ on what counts as a machine trace. The median AI-tells quality score this issue is 3.95 — keep that subjectivity in mind when reading it.
For the genre effect, cross-issue trajectories and version comparisons, see Issue 05 · Three genres side by side.
Appendix: the outline all ten models were given
Where the beat sheet came from
This issue's beat sheet is taken from a novel already published on Qidian, written by a human editor from the actual prose.
Positioning
High-concept / academic underdog / parallel universes / light comedy.
Logline: “Li Dong, a failing high-school senior, stumbles into a WeChat group called the ‘Azure Dragon Study Group’ — and finds the members are Newton, Einstein and Gauss from different parallel universes.”
Chapter one: into the group of gods by mistake
Li Dong, fresh off 87 on the monthly exam, mistakenly joins a study group at an internet café, where a few avatars are discussing problems he cannot begin to follow. Three scenes:
- The exam blow(everyday / self-mocking): 150 to ace it, 90 to pass — he scored 87, three points short. The brief: show his situation through his deskmate's and his mother's reactions, not through direct interior monologue.
- Into the wrong group(suspense / contrast): the group nicknames are “Isaac”, “Albert” and “Carl”, and they talk strangely. The brief called for restraint here — the protagonist must not guess who they are straight away.
- The first red packet(payoff / setup): he tosses off an answer and receives a “Focus +1” red packet — and the stat actually takes effect.
Closing hook: as the stat boost kicks in, Li Dong looks back at that 87-point paper. Several models dropped this cheat-power beat.
Target length was about 3,000 characters. This issue's actual average was 3,912, with the longest chapter at 7,536.
Thanks to this issue's reviewers
武行悟道雪白的面具lastworm流萤King长庚上帝模式笑沧桑海獭慕容复国在天子家喝茶风方周兴sunHubert坚强老鸭梨Sail荀彧泡在海岛的猴子Sandra
Read a lot of web fiction and willing to score it seriously? Join the review panel.
Join the panel →What should we benchmark next?
Tell us the genre, model or part of the process you want tested.