AI prose · genre by genre
Xianxia, high-concept and period: three benchmarks, 90 works, 306 double-blind ratings.
Genres side by side
The three issues share the same scoring axes, reviewer grouping and weighting — what changes is what each genre tests.
The three issues' means are 5.80, 6.04 and 6.17 — at most 0.37 apart. Scoring severity is close enough that comparing across issues is basically sound.
Where the genres pull the scores apart
Below are the axis means of each issue's ten models.
High-concept: web-fiction feel comes built in
Mean web-fiction feel is 6.53, a good half point above xianxia (5.88) and period (5.69). “Newton pops up in a WeChat group” carries its own payoffs and conflict. Logic is also the lowest of the three at just 5.35 — the wilder the premise, the harder it is to make it hold.
Period: steadiest logic, hardest web-fiction feel
Mean logic of 5.78 is the highest of the three issues; web-fiction feel of 5.69 is the lowest. Period fiction follows the logic of everyday life, so hard flaws are rarer. The difficulty is writing ration coupons and communal housing blocks so they ring true — while keeping the story fun to read.
The quality score for AI tells has slid three issues running, from 4.36 to 4.16 to 3.99, finishing last of the five axes every time. Sounding like a machine is still the most widespread problem.
The correlation between length and score is 0.71 for high-concept and just 0.14 for period. High-concept stories need room to lay the premise out — cut the words and it stops making sense; period fiction runs on picking the right details, and no word count piles up into period feel.
Change what's being tested, and the board reshuffles
The high-concept and period issues entered the pool on the same day and were scored by the same reviewers over the same stretch — the cleanest place to watch what genre alone does.
High-concept ranks on the left, period on the right. Same models, same reviewers.
Steepest fall: DeepSeek V4 Flash
6th on high-concept (5.92) → 10th on period (4.67): down 4 places and 1.25 points. Its web-fiction feel of 7.30 on high-concept (second in the field) shrinks to 5.32 on period, and its AI-tells quality score drops from 4.62 to 2.31.
Sharpest climb: Kimi K3
4th on high-concept (6.28) → 1st on period (7.19). On high-concept it wrote an 8.00 and a 4.33, a range of 3.67; on period its three chapters tighten to 8.00 / 7.36 / 6.20. The moment it steadied, it took the top.
Eight of the ten models change places when the genre changes; only DeepSeek V4 Pro and GPT-5.6 Terra hold perfectly still (7th and 8th both times). The gap between one model's genres can be wider than the gap between models. Choose the genre first, then the model.
Who's good at which genre
Ranks 1 / 1 / 3 across the three issues, with a total-score range of just 0.66. Currently the only model that doesn't mind which genre it's handed.
4th in xianxia is its best rank of the three (9th and 6th after that). An invented world is the question it answers most fluently.
6th on high-concept, against 9th and 10th in the issues either side. In the high-concept issue its web-fiction feel surged to 7.30, second in the field.
Topped the period issue (7.19), with period detail that reads true and stays steady. Ranks of 5 → 4 → 1 across three issues — straight up.
Each model gets only three chapters per issue, so read these as leanings. The full caveats are at the end.
The beat sheet opens gaps too
All three issues had models write prose straight from a beat sheet — and the three beat sheets differ in origin and quality.
Generated by AI, unpolished.
AI-generated, then refined by hand.
From a published work — pacing, conflict and hooks arrive more complete.
A human beat sheet builds the stage better
The high-concept issue's mean web-fiction feel of 6.53 is the highest of the three. The genre helps — but the human beat sheet had also laid out the pacing, conflict and hooks in advance. The models only had to write them down.
Scrubbing the AI flavour starts at the beat sheet
The beat sheet fixes the prose's structure and verbal habits. Editing the prose alone rarely gets it clean — treat the beat sheet and the prose together.
Beat-sheet quality is mixed into the score gaps between issues, so means alone cannot rank the models. In an AI writing pipeline, manage beat-sheet quality alongside prose quality.
The four models that ran all three issues
These four took part in the xianxia, high-concept and period issues back to back.
The second-round list kept the xianxia winner Gemini 3.1 Pro, the top two Chinese models Qwen3.8 Max and Kimi K3, and the then freshly released DeepSeek V4 Flash, as a counterpart for the later-added DeepSeek V4 Pro.
Gemini 3.1 Pro is the only model that stayed in the top tier all three times (ranks 1 / 1 / 3), with a range of just 0.66. Kimi K3 climbed the whole way, from 5th in xianxia to 1st in period. DeepSeek V4 Flash spans 1.25 points across the three issues, the most genre-bound of the four. Qwen3.8 Max caved in once, on high-concept (4th → 9th → 6th).
Each model against itself
GLM and Gemini Flash track version changes; GPT-5.6, three models from one family.
Version upgrades: GLM and Gemini Flash
GLM-5.2 → GLM-5.3 is the most successful upgrade here: +0.63 on high-concept and +1.17 on period, with cost per chapter up from $0.018 to $0.0233 — thirty percent more money for a whole tier of improvement. Gemini 3.6 → 3.7 Flash costs the same and gains +0.20 / +0.42 across the two genres.
Old and new versions were tested on different genres and beat sheets, so treat the gains as indicative. Both new versions gain on both high-concept and period, which at least points one way. The just-released Gemini 3.8 Flash will go into the next issue as soon as we can manage it.
GPT-5.6 Sol / Luna / Terra: three models from one family
Sol, Luna and Terra are three different models in the GPT-5.6 family. Judged by price per chapter, Sol is the flagship, Terra the mid-range and Luna the lightweight, with up to a fivefold gap between them.
Luna and Terra share genres and reviewers, so they can be compared head on: a 1.33-point gap on high-concept and 1.19 on period — while Terra costs 1.5× more per chapter. On these two web-fiction issues, the cheapest, lightweight Luna scores higher and costs less (both boards are in Issue 03 · High-concept and Issue 04 · Period).
Sol only took part in the xianxia issue, with a different genre and different reviewers, so it cannot be set directly against the other two: it placed 3rd, but at $0.145 a chapter it is the priciest of the three.
How far these comparisons hold
Cross-issue comparison is shaped by genre, beat sheets, reviewers and sample size.
- The genre comparison is the most reliable
- The high-concept and period issues entered the pool on the same day and were scored by the same reviewers over the same stretch. The main variables are the genre and the beat sheet that came with it.
- Trajectories show adaptability
- All three issues scored to one rubric, with means at most 0.37 apart. But the genres differ and reviewer headcounts moved, so the trajectories speak to cross-genre performance — not to a model's ability changing over time.
- Version gains are indicative only
- Old and new versions faced different genres and beat sheets, so other factors are mixed into the gap. The same-issue contrast between GPT-5.6 Luna and Terra is the more direct one.
- The beat sheets have different origins
- Xianxia used a plain AI beat sheet, period a hand-refined AI one, and high-concept a human beat sheet from a published work. The three issues' means cannot simply be ranked against one another.
- Sample size
- Three chapters per model per issue. A volatile model's mean carries little weight, and rank changes pick up randomness too.
Each issue's full leaderboard, axis breakdown and complete texts: Issue 02 · Xianxia · Issue 03 · High-concept · Issue 04 · Period