Skip to main content

Xu Ben Project: An Open Research Log — June 2026

June 17, 2026

This page is a project hub. It reports interim observations and the reasoning that produced them; it is not a finished paper, and the difference is specific rather than rhetorical. Concretely: no figure on this page carries an interval, no coding result carries an inter-coder reliability statistic, and no sampling frame is fully specified. Claims made here are therefore provisional in a way that the language of "findings" can obscure, and the section "What This Log Is Not Entitled To Claim" below is the operative constraint on how these numbers may be used — including by the author. It documents how a vague intuition — "Xu Ben seems to talk about civilisation more than institutions these days" — was tested, corrected, and rebuilt into a different kind of question over two days of intensive AI-assisted research.

Future substantial observations will normally be published as separate research-log entries and linked back here. This page will stay as the stable entry point: current model, timeline, public files, and guardrails.

One interim observation matters more for method than for Xu Ben specifically. Counting keywords and tracking the trend — the obvious approach — gave a different answer from counting them within genres, and neither answer is automatically the better one. Why that happens, and what it costs to choose between them, is set out under The Current Model below.

The Original Question

Why did Chinese-language public intellectuals in the 2020s appear to shift from specific institutional, historical, and social analysis toward diagnostic discourse centred on "subjectivity," "the complete human," "civilisational crisis," and "meaning"?

Xu Ben (徐贲), a Chinese-American professor who has published extensively in both mainland Chinese media and overseas, was chosen as the entry point — not as a representative case, but as one author whose output is large enough and spans enough platforms to test the question.

How The Hypothesis Changed

The project began with a simple expectation: aggregate keyword frequency would show a decline in institutional vocabulary (制度, 公民, 公共) and a rise in diagnostic vocabulary (人类, 文明, 意义, 主体性) across Xu Ben's writing over time.

This expectation survived first contact with data — barely. Then it was modified five times, each time in the same direction: a naive measure was found to diverge from a stratified one. No test left a measure intact; the project's only flat series (Caixin, 71 → 69) is discussed in Correction 2 below.

Correction 1: Genre Mix Effect

The aggregate diagnostic-share trend (25% → 31% → 35%) across three periods diverged from the genre-stratified trend. Within academic essays (论文) alone, the trend flattened at 29% in 2021-2026. The aggregate and the stratified figures disagree; which is the better estimate of the author depends on whether genre choice is exogenous noise or part of the trajectory being measured (see the methodological note below). Genre-stratified reporting was adopted for subsequent work in this project. (Note that it was not applied to the platform comparison in Correction 2 below, which reports unstratified aggregates; that comparison therefore remains open to the same composition objection.)

Correction 2: Platform Selection Effect

The institutional vocabulary decline seen on Aisixiang (爱思想, a mainland academic repost platform) — from 99 to 52 per 10,000 characters — did not appear on the Caixin blog (财新博客) for the same period (71 → 69, essentially flat). Paired same-article comparison (14 articles appearing on both platforms) showed nearly identical keyword density, consistent with the divergence being driven by which articles each platform collected rather than by within-article editing. (This rules out one mechanism — same-article political editing — but not platform-level content selection, which in a mainland repost platform may itself be a form of political filtering. See entry 3 below.)

Correction 3: Publication Pipeline Effect

Matching book tables of contents against platform article titles gave 46–89% apparent chapter-article overlap across four collections (57 title matches; matching was done on titles, not on article text, and no text-level comparison was performed). Two patterns are consistent with this data and are offered as hypotheses, not findings: that the Caixin blog held drafts later collected in 《颓废与沉默》(2015), and that Aisixiang holds an archived layer of material from 《人以什么理由来记忆》(2008). Neither direction can be established from this evidence: title matching does not order draft and reprint, and the Aisixiang timestamps that would order them are the same field that Correction 5 below shows to be contaminated by batch imports — including a 2008 cluster.

Correction 4: The Same Word Does Different Work

Close reading of 6 articles (assisted by three AI systems — GPT, Gemini, and Grok — across multiple rounds, with errors exposed and corrected along the way; see the AI statement below for slot assignments) suggested a provisional coding scheme in which the word 制度 performs at least four distinct text functions. This scheme has not been tested for inter-coder reliability and should be treated as a hypothesis, not a finding:

Table 1
Type What 制度 does
Institution-building Constructive framework: civic education, democratic governance
Institution-failure analysis Domination mechanism: propaganda, totalitarian education
Public-intellectual-failure diagnosis Background: focus on intellectual silence, cynicism
Humanistic-meaning diagnosis Exits the frame: AI, human subjectivity, civilisation

The same author could produce high-institutional and high-diagnostic texts in the same year (2016). On Aisixiang, the aggregate decline is not evidence of abandonment. Whether any author-level change occurred is not settled by these data: the Caixin series is flat over the same period, and the log has no rule for deciding which platform estimates the author.

Correction 5: Batch-Import Dates Distort Chronology

Aisixiang's "update time" (更新时间) includes batch-import events. One day in 2008 accounts for 12 articles uploaded within a single hour — clearly not a natural publication pattern. These are archive-loading dates, not writing dates. Any time-series analysis must detect and flag such clusters. (This correction changed no figure reported on this page: the metrics above are densities per 10,000 characters, and the batch cluster affects counts. It belongs here as a data-handling note rather than as a correction that revised a result — see entry 7 below.)

The Current Model

Within the sampled materials, the two platforms disagree. Aisixiang shows a decline in institutional vocabulary; Caixin does not. This log models the disagreement as arising from text-function composition and pipeline position. It does not estimate an author-level trajectory, and the data collected here cannot do so. The same author produces different text types — provisionally coded as institution-building, institution-failure analysis, public-intellectual-failure diagnosis, and humanistic-meaning diagnosis — across different platforms, genres, and publication cycles. The four-way scheme is a hypothesis with no reliability testing behind it; entry 4 below states what it cannot be used to claim.

The methodological lesson is conditional, not general: aggregate keyword trends and genre-stratified trends can diverge, and where they diverge, neither is automatically the better estimate. Stratifying by an exogenous variable (which platform collected an article, which crawl batch it arrived in) removes a source of contamination. Stratifying by an endogenous variable (which genre the author chose to write) can remove the effect being measured, because genre choice may be part of the shift rather than noise around it. This log stratified by both without distinguishing them, and does not know which of the two divergences it observed is which.

Research Timeline

Table 2
Date Stage Key Event
2026-06-16 Seed note Published observation notes on Xu Ben's 2025-2026 AI-era writings
2026-06-16 Aisixiang pilot 89 articles across 5 genres; keyword groups v0.2; genre-stratified analysis
2026-06-16 Methodology Literature review (9 clusters); peer review by AI reviewers (GPT, Gemini, Grok)
2026-06-17 Caixin comparison 42 Caixin articles; cross-platform paired comparison (14 pairs)
2026-06-17 Book overlap 4 book TOCs; 57 chapter-article matches; title-level overlap mapping
2026-06-17 Close reading 6 articles across 3 AI reviewers (9 slot-assignments: GPT 2, Gemini 1, Grok 6); provisional four-type text-function model (untested — see entry 4)
2026-06-17 Model consolidation v0.2 methodology summary; codebook; manual coding

What This Log Is Not Entitled To Claim

Each entry names a sentence type that does not appear on this page, and what would earn the right to write it. Entries are constraints on the text, not statements about the world: nothing below asserts a finding.

1. Not entitled to claim that stratified estimates are more accurate than aggregate ones — only that they differ. Forbids: calling the aggregate trend an artefact, a distortion, or misleading; calling the genre-stratified figure the corrected one. Earns it back: an independent criterion against which the two estimates can be scored — held-out full-text classification, the author's own account of his output, or a second corpus where the true direction is known by other means.

2. Not entitled to claim that any vocabulary shift in Xu Ben's writing has been established. Forbids: "the decline," "the diagnostic turn," "the visible vocabulary shift is real" as referring expressions; any author-level explanation of a platform-level pattern. Earns it back: the same direction on both platforms within matched genres, or a stated and defended rule for which platform estimates the author.

3. Not entitled to claim that political filtering has been ruled out. Forbids: presenting "content selection" as an alternative to censorship rather than a possible form of it; any sentence implying the censorship hypothesis was tested. The paired comparison (14 articles) tests within-article editing only, and by construction samples articles that passed both platforms' filters. Earns it back: comparing the political content of the Caixin articles Aisixiang did not collect against those it did — data already in hand, test not run.

4. Not entitled to claim any result derived from the four-type text-function codebook. Forbids: the Current Model as currently worded; the count "four"; any percentage produced by manual coding. No inter-coder reliability was computed, the coder and the hypothesis-former are the same party, and Type 4 is defined by the coded term's absence, which makes the scheme unfalsifiable as written. Earns it back: double-coding of a held-out subsample by an independent coder with a reported agreement statistic, plus a redefinition of Type 4 that does not rest on absence.

5. Not entitled to claim a direction of transfer between platform and book. Forbids: "functioned as a draft workspace," "functioned as an archive layer," "lifecycle direction." Matching was on titles, not text; ordering requires the Aisixiang timestamps that Correction 5 flags as contaminated by batch imports, including in 2008 — the same year as the archive-layer claim. Earns it back: dated print sources, or text-level comparison of article and chapter.

6. Not entitled to claim that the AI division of labour reduced error in this project. Forbids: "the most valuable AI contribution was disagreement"; "AI systems did not make final interpretive judgments" (the word "final" makes this unfalsifiable — it is satisfied by whatever the human did last). The reviewer outputs are not published, so no reader can check any attribution, and the one documented episode (the platform-editing question) eliminated only the weak form of the hypothesis it is credited with eliminating. Earns it back: publication of the reviewer transcripts for the six close-read articles, or per-finding provenance tags naming which system proposed and which challenged each claim.

7. Not entitled to narrate the five corrections as evidence of rigour. Forbids: "corrected five times" as a robustness claim. All five corrections run in one direction — a naive measure found contaminated. No test is reported as having left a measure intact. Correction 5 changed no figure reported here, since the metrics are densities and the batch cluster affects counts. The project's only flat series (Caixin, 71 → 69) is reported as a phenomenon rather than as a null result. Earns it back: a stopping rule — how many confounders were examined, and how many produced no effect — plus at least one reported null per stage.

8. Not entitled to the protections of a process note while using the language of results. Forbids: "results," "findings," "mandatory," "public package" — unless accompanied by the reporting standards those words imply (n per cell, interval or spread, sampling frame, coding reliability). Where those standards are not met, this page must say so at the point of the number, not once at the top. Earns it back: meeting the standards, or restating the numbers as descriptive counts without trend language.

Public Research Files

The following cleaned research files are available for review. They are the current public file set for this project; later stable versions may replace or supplement them.

  • Methodology Summary v0.2 — 2-page overview: hypothesis → corrections → current model
  • Interim Observations v0.1 — 8 observations with evidence and interpretation. None carries an interval, a sampling frame, or a reliability statistic; read them under entry 8 above
  • Source Map — Platform survey across Aisixiang, Caixin, CDT, and others
  • Publication Map — Xu Ben's book bibliography with publication clusters
  • Book-Article Overlap Table — Which chapters match which platform articles
  • Text-Function Codebook v0.1 — Coding rules for the four provisional text-function types (v0.1; no inter-coder reliability computed, and Type 4 is defined by the coded term's absence — see entry 4)
  • Keyword Groups v0.2 — The vocabulary classification used for scoring

Full article texts and raw crawl data are retained locally but not published (copyright and verification status). AI reviewer outputs and internal working notes are also unpublished, but for a different reason: they are unedited and unverified, not copyright-restricted. Entry 6 above treats their publication as the condition for claiming any error-reduction benefit from the AI division of labour, and that condition is not met here.

Update Policy

This page is not meant to absorb every new observation. Minor corrections and data-handling notes go into the project changelog. Stable files in /research/xuben/ are updated when the underlying evidence or model changes. A new public research-log post should be created when the project reaches a new analytical stage: for example, an overseas-media comparison, a CDT sampling result, a media-agenda pilot, or a revised text-function model.

AI-Assisted Research Statement

This project was conducted by the site author with assistance from multiple AI systems:

  • Claude Code (cc): data collection scripts, keyword scoring, cross-platform analysis, statistical comparison, close-reading file preparation, version control
  • Codex (coco): methodology design, consolidation of interim observations, media-agenda framework, proposal of the provisional four-type text-function model, editorial oversight
  • GPT: close reading (2 slots), theoretical framing, critical review
  • Gemini: close reading (1 slot + predictions), corrected twice on platform-editing claims
  • Grok: close reading (all 6 slots), self-corrected after re-reading, operational suggestions

Two things must be said about this arrangement, and only one of them is favourable.

The favourable one: one assumption — that platform divergence reflected within-article editing — was surfaced and dropped in the course of multi-round review rather than carried through. The mechanism was not reviewer-versus-reviewer disagreement: of the two correction episodes noted in the list above, one was a reviewer corrected by the author and one was a reviewer correcting itself on re-reading; and with 9 slot-assignments across 6 articles (GPT 2, Gemini 1, Grok 6), at least half the close-read articles had only one reviewer. (Note: this eliminated the weak form of the censorship hypothesis — same-article political editing — not the strong form, which is platform-level content selection. See Correction 2 and entry 3 above.)

The unfavourable one, stated as a hypothesis in the seed note of 2026-06-16 and reaffirmed here: AI assistance lowers the marginal cost of producing an additional supporting case to near zero while leaving the cost of finding a disconfirming case unchanged. The expected effect is a drift in the ratio. This log is consistent with that drift: five corrections are reported and all five run in the same direction (a naive measure found to be contaminated); no test is reported as having left a measure intact; the one flat series in the data (Caixin, 71 → 69) is presented as a phenomenon to be modelled rather than as the project's cleanest null result. The reviewer outputs that would let a reader audit this are not published. Readers should weight the two accordingly — the favourable one is an anecdote, the unfavourable one is a structural prediction that the contents of this page do not contradict.

All research decisions, hypothesis revisions, and publication choices are the responsibility of the human author.