Issue 05 · Three genres side by side · 2026.09.04

AI prose · genre by genre

Xianxia, high-concept and period: three benchmarks, 90 works, 306 double-blind ratings.

0Genres
0Models involved
0Anonymous works
0Valid ratings
01

Genres side by side

The three issues share the same scoring axes, reviewer grouping and weighting — what changes is what each genre tests.

The three issues' means are 5.80, 6.04 and 6.17 — at most 0.37 apart. Scoring severity is close enough that comparing across issues is basically sound.

Where the genres pull the scores apart

Below are the axis means of each issue's ten models.

High-concept: web-fiction feel comes built in

Mean web-fiction feel is 6.53, a good half point above xianxia (5.88) and period (5.69). “Newton pops up in a WeChat group” carries its own payoffs and conflict. Logic is also the lowest of the three at just 5.35 — the wilder the premise, the harder it is to make it hold.

Period: steadiest logic, hardest web-fiction feel

Mean logic of 5.78 is the highest of the three issues; web-fiction feel of 5.69 is the lowest. Period fiction follows the logic of everyday life, so hard flaws are rarer. The difficulty is writing ration coupons and communal housing blocks so they ring true — while keeping the story fun to read.

The quality score for AI tells has slid three issues running, from 4.36 to 4.16 to 3.99, finishing last of the five axes every time. Sounding like a machine is still the most widespread problem.

The correlation between length and score is 0.71 for high-concept and just 0.14 for period. High-concept stories need room to lay the premise out — cut the words and it stops making sense; period fiction runs on picking the right details, and no word count piles up into period feel.

Change what's being tested, and the board reshuffles

The high-concept and period issues entered the pool on the same day and were scored by the same reviewers over the same stretch — the cleanest place to watch what genre alone does.

High-concept ranks on the left, period on the right. Same models, same reviewers.

Steepest fall: DeepSeek V4 Flash

6th on high-concept (5.92) → 10th on period (4.67): down 4 places and 1.25 points. Its web-fiction feel of 7.30 on high-concept (second in the field) shrinks to 5.32 on period, and its AI-tells quality score drops from 4.62 to 2.31.

Sharpest climb: Kimi K3

4th on high-concept (6.28) → 1st on period (7.19). On high-concept it wrote an 8.00 and a 4.33, a range of 3.67; on period its three chapters tighten to 8.00 / 7.36 / 6.20. The moment it steadied, it took the top.

Eight of the ten models change places when the genre changes; only DeepSeek V4 Pro and GPT-5.6 Terra hold perfectly still (7th and 8th both times). The gap between one model's genres can be wider than the gap between models. Choose the genre first, then the model.

Who's good at which genre

Gemini 3.1 Pro
Steady front-runner

Ranks 1 / 1 / 3 across the three issues, with a total-score range of just 0.66. Currently the only model that doesn't mind which genre it's handed.

Qwen3.8 Max
Strongest in xianxia

4th in xianxia is its best rank of the three (9th and 6th after that). An invented world is the question it answers most fluently.

DeepSeek V4 Flash
Strongest in high-concept

6th on high-concept, against 9th and 10th in the issues either side. In the high-concept issue its web-fiction feel surged to 7.30, second in the field.

Kimi K3
Strongest in period

Topped the period issue (7.19), with period detail that reads true and stays steady. Ranks of 5 → 4 → 1 across three issues — straight up.

Each model gets only three chapters per issue, so read these as leanings. The full caveats are at the end.

02

The beat sheet opens gaps too

All three issues had models write prose straight from a beat sheet — and the three beat sheets differ in origin and quality.

Beat-sheet quality
Issue 02 · Xianxia
AI beat sheet · ordinary quality

Generated by AI, unpolished.

Issue 04 · Period
AI beat sheet · improved

AI-generated, then refined by hand.

Issue 03 · High-concept
Human author · real beat sheet

From a published work — pacing, conflict and hooks arrive more complete.

Beat-sheet quality weighs heavily on prose quality. With the same bare, no-prompt generation, a good beat sheet buys visibly better prose; if the beat sheet carries an AI flavour, the prose never quite scrubs clean.

A human beat sheet builds the stage better

The high-concept issue's mean web-fiction feel of 6.53 is the highest of the three. The genre helps — but the human beat sheet had also laid out the pacing, conflict and hooks in advance. The models only had to write them down.

Scrubbing the AI flavour starts at the beat sheet

The beat sheet fixes the prose's structure and verbal habits. Editing the prose alone rarely gets it clean — treat the beat sheet and the prose together.

Beat-sheet quality is mixed into the score gaps between issues, so means alone cannot rank the models. In an AI writing pipeline, manage beat-sheet quality alongside prose quality.

03

The four models that ran all three issues

These four took part in the xianxia, high-concept and period issues back to back.

The second-round list kept the xianxia winner Gemini 3.1 Pro, the top two Chinese models Qwen3.8 Max and Kimi K3, and the then freshly released DeepSeek V4 Flash, as a counterpart for the later-added DeepSeek V4 Pro.

Gemini 3.1 Pro is the only model that stayed in the top tier all three times (ranks 1 / 1 / 3), with a range of just 0.66. Kimi K3 climbed the whole way, from 5th in xianxia to 1st in period. DeepSeek V4 Flash spans 1.25 points across the three issues, the most genre-bound of the four. Qwen3.8 Max caved in once, on high-concept (4th → 9th → 6th).

04

Each model against itself

GLM and Gemini Flash track version changes; GPT-5.6, three models from one family.

Version upgrades: GLM and Gemini Flash

GLM-5.2 → GLM-5.3 is the most successful upgrade here: +0.63 on high-concept and +1.17 on period, with cost per chapter up from $0.018 to $0.0233 — thirty percent more money for a whole tier of improvement. Gemini 3.6 → 3.7 Flash costs the same and gains +0.20 / +0.42 across the two genres.

Old and new versions were tested on different genres and beat sheets, so treat the gains as indicative. Both new versions gain on both high-concept and period, which at least points one way. The just-released Gemini 3.8 Flash will go into the next issue as soon as we can manage it.

GPT-5.6 Sol / Luna / Terra: three models from one family

Sol, Luna and Terra are three different models in the GPT-5.6 family. Judged by price per chapter, Sol is the flagship, Terra the mid-range and Luna the lightweight, with up to a fivefold gap between them.

GPT-5.6 Sol
$0.145 / chapter
Xianxia 6.40 (Issue 02 · No. 3)
GPT-5.6 Luna Better at writing
$0.029 / chapter
High-concept 6.66 · Period 6.62
GPT-5.6 Terra
$0.0725 / chapter
High-concept 5.33 · Period 5.43

Luna and Terra share genres and reviewers, so they can be compared head on: a 1.33-point gap on high-concept and 1.19 on period — while Terra costs 1.5× more per chapter. On these two web-fiction issues, the cheapest, lightweight Luna scores higher and costs less (both boards are in Issue 03 · High-concept and Issue 04 · Period).

Sol only took part in the xianxia issue, with a different genre and different reviewers, so it cannot be set directly against the other two: it placed 3rd, but at $0.145 a chapter it is the priciest of the three.

05

How far these comparisons hold

Cross-issue comparison is shaped by genre, beat sheets, reviewers and sample size.

The genre comparison is the most reliable
The high-concept and period issues entered the pool on the same day and were scored by the same reviewers over the same stretch. The main variables are the genre and the beat sheet that came with it.
Trajectories show adaptability
All three issues scored to one rubric, with means at most 0.37 apart. But the genres differ and reviewer headcounts moved, so the trajectories speak to cross-genre performance — not to a model's ability changing over time.
Version gains are indicative only
Old and new versions faced different genres and beat sheets, so other factors are mixed into the gap. The same-issue contrast between GPT-5.6 Luna and Terra is the more direct one.
The beat sheets have different origins
Xianxia used a plain AI beat sheet, period a hand-refined AI one, and high-concept a human beat sheet from a published work. The three issues' means cannot simply be ranked against one another.
Sample size
Three chapters per model per issue. A volatile model's mean carries little weight, and rank changes pick up randomness too.

Each issue's full leaderboard, axis breakdown and complete texts: Issue 02 · Xianxia · Issue 03 · High-concept · Issue 04 · Period