Extended corpus: the persistence gap does not replicate¶
The pilot corpus result was a selection artefact
The persistence gap between emotional registers — ×9.2 versus ×2.9, \(p = 0.08\) across twenty-four hand-picked subjects — becomes ×3.04 versus ×2.90, \(p = 0.53\) across 440 category-derived subjects. The factor of three has vanished.
And the switching-rate gap is an audience effect
Accusation subjects switch regime three times more often (8.6 % versus 2.7 %, \(p = 0.014\)) — but they are also three and a half times more visited (\(p = 5\times10^{-14}\)). At comparable traffic, the odds ratio falls from 3.4 to 1.38 (\(p = 0.63\)).
The new protocol has a flaw of its own
Category membership is a noisy proxy for the register: "Lil Tay" sits in a hoax category, but its audience is that of a celebrity. Label noise attracts any gap towards zero — the null result is therefore consistent with the absence of an effect as much as with a diluted one.
Since settled — by blind annotation
The ambiguity this page leaves open was resolved by annotating the register by hand across all 440 subjects. Label noise is indeed massive — 40 % of subjects belong to neither register — and it carried the whole switching-rate gap: 8.6 % versus 2.7 % becomes 4.8 % versus 5.1 % (\(p = 1.00\)). The gap was not diluted, it does not exist.
Why change the selection protocol, not just the size¶
The pilot corpus comprised twenty-four hand-picked subjects. At that size the choice can be read and contested. At several hundred it can no longer be read — and that is precisely where selection bias becomes invisible: nothing in a list of three hundred titles distinguishes those that would have been retained for what they show.
The extended corpus therefore replaces the choice of subjects with the choice of
categories. Seventeen Wikipedia categories are declared in ide.catalogue.REGISTERS, and
everything they contain enters the pool. Whether an article belongs to a category is decided
by Wikipedia's contributors, not by the author of the analysis.
| Register | Categories | Available | Substantial | Retained |
|---|---|---|---|---|
| accusation | conspiracy theories, disinformation, fake news, hoaxes, political scandals, corruption, propaganda, moral panic | 3,126 | 1,935 | 220 |
| discovery | space probes and missions, physics experiments, astronomical surveys, exoplanets, Nobel laureates | 6,322 | 2,705 | 220 |
Four precautions against foreseeable confounds:
| Precaution | Rationale |
|---|---|
| events on both sides | pitting affairs against abstract concepts would compare subject types, not registers |
| disjoint classes | an article in both registers is discarded, not arbitrated |
| sampling by title fingerprint | alphabetical truncation would systematically over-represent certain subjects |
| substance filter (≥ 10,000 bytes) | "Astronomical surveys" explored to depth 1 returns nearly four thousand zero-traffic catalogue entries, which would decimate one register and not the other |
The resulting manifest is written to data/catalogue.json and versioned: it constitutes the
pre-registration, and a published result refers to a consultable corpus.
Results¶

The audience imbalance between registers; the disappearance of the rate gap once traffic is controlled; the non-replication of the persistence gap; and the switching rate following audience rather than register. Figure regenerated by notebook 11.
1. A gap appears, then vanishes¶
| Comparison | accusation | discovery | odds ratio | p |
|---|---|---|---|---|
| raw (440 subjects) | 8.6 % | 2.7 % | 3.4 | 0.014 |
| stratum ≥ 47 views/day | 16.5 % | 12.5 % | 1.38 | 0.63 |
| traffic-matched (173 pairs) | 6.9 % | 3.5 % | — | 0.18 (McNemar) |
Detection requires a prior regime of at least 50 views per day. Median traffic is 39 views/day for accusation against 11 for discovery, and the share above the threshold is 47 % versus 22 %. The register does not explain what audience already explains.
A direction remains — accusation subjects switch slightly more often at equal traffic — but nothing that warrants a conclusion.
2. Persistence does not replicate¶
| Corpus | accusation | discovery | p |
|---|---|---|---|
| pilot, 14 hand-picked subjects | ×9.20 (n = 8) | ×2.90 (n = 6) | 0.081 |
| extended, 440 category-derived subjects | ×3.04 (n = 21) | ×2.90 (n = 7) | 0.533 |
The pilot corpus contained the best-known conspiracy theories — QAnon at ×44, Pizzagate at ×18 — and that is exactly what hand-selection produces: the cases that come to mind are the extreme ones. The category-derived corpus also contains dozens of obscure subjects from the same register, and its median becomes indistinguishable from the discovery register's.
The selection protocol did not change the precision of the result: it changed its sense.
3. The register is a noisy proxy¶
The largest lifts retained on the accusation side are instructive:
| Subject | Date | Lift |
|---|---|---|
| Lil Tay | 27 June 2020 | ×7.9 |
| Watch Dogs (video game) | 6 December 2020 | ×6.3 |
| Million Dollar Extreme | 6 July 2016 | ×5.7 |
| The Capture (TV series) | 30 July 2022 | ×5.5 |
| Mossack Fonseca | 14 October 2019 | ×4.8 |
"Lil Tay" appears in a hoax category because of a hoax about her death, but the article's attention dynamics are those of a celebrity. The same holds for a video game or a television series filed under thematic categories without their audience being driven by a scandal.
Label noise does not bias the result in one direction: it attracts it towards zero. The null result is therefore consistent with two readings these data cannot separate — no effect, or a diluted effect.
4. Identification remains out of reach¶
Two fits out of twenty-eight were initially declared usable, with ratios of 697 and 5,431 — forgetting times of several years for a four-month fitting window. A near-step transition does not expose the knee that carries the information about \(\lambda\): the fit hugs the curve, the residual scatter is excellent, and the ratio runs away unconstrained.
An observability check was therefore added to the module — the forgetting time must fit inside the fitted window. After correction, as on the pilot corpus, no change yields usable parameters.
What this changes for the project¶
The only difference between emotional registers the project had measured does not survive verification. The emotional-charge mechanism \(\alpha\) remains without empirical support, neither through the amplification rate (calibration) nor through persistence.
For the memorandum the consequence is direct: the persistence indicator proposed in place of a cap on \(\gamma\alpha/\lambda\) does measure something — switch dates and amplitudes are robust — but it does not discriminate between emotional registers on this corpus. It remains usable as an instrument of observation, not as proof of a mechanism.
Neither protocol settles it — and what settled it¶
The pilot was biased by selection; the extended corpus is noisy in its labelling. This is not a dead end but a specification: a third protocol should combine a category-derived pool — for the absence of selection bias — with a subject-by-subject validation of the register, carried out blind, without seeing the series. That is annotation work, not computation.
Done — see Blind annotation
That third protocol was carried out. All 440 subjects were coded by hand from title and lede alone, under a rubric published before any annotation existed. Agreement between category and annotation is 59.5 %; the switching-rate gap falls to an odds ratio of 0.93 (\(p = 1.00\)); and the discarded subjects — those belonging to neither register — turn out to switch more often than either. They were what carried the gap.
Open leads¶
- ~~Annotate the register by hand on the category-derived pool, blind.~~ → done, across all 440 subjects rather than two hundred. It settled the question: no effect, rather than a diluted one.
- Match on traffic at construction time rather than afterwards: drawing subjects in pairs of comparable traffic would immunise the comparison against the principal confound. The annotation adds a requirement: match on subject kind too, since the two registers turn out to barely cover the same kinds of subject.
- Lower the detection threshold by aggregating weekly. A switching rate of 2 to 8 % leaves more than nine subjects in ten without a measurement.
- Establish a baseline switching rate on a control register — subjects with no particular emotional charge, traffic-matched — to measure the detector's specificity.
Implementation: ide.catalogue · Notebook:
11 — Extended corpus ·
pilot corpus · roadmap