Exposed Diversity Index¶
Measuring the diversity a recommender feed actually exposes — and testing what we think we know about it.
What this repository contains¶
One instrument, and nothing else: a measure of the diversity a news feed actually exposes to its reader, computable without access to the platform's code — with everything needed to establish it, attack it, and measure it on real logs.
Twelve executable notebooks, 256 tests, twenty-seven recorded corrections, and a state of affairs resembling neither what the project announced nor what it believed it had established.

Served diversity, measured over 232,887 feeds of the Danish daily Ekstra Bladet — and what the unknown order leaves undetermined. Figure regenerated by notebook 26.
The verdict, in one page¶
| Object | State |
|---|---|
| The index as first proposed | untenable: saturable at zero cost, then evadable by burial → adversarial test · adversarial rank |
| The retained form | defined and attacked: entropy of served items over a declared catalogue, weighted by the attention of each rank → Index |
| Exposure, on which that weighting depends | measured: \(0.88 \pm 0.05\) rather than \(1\) by convention, no cascade, portable from page to page to within 6 % → measured exposure · the page effect |
| The rank-blind index, on real feeds | measured for the first time: 0.50 per user-day, 5.1 effective sections out of 26 — but this is the form the adversarial test disqualifies, hence an upper bound → the index measured |
| The exposed index, on those same feeds | bounded, not measured — the order column of the only dataset offering it does not contain the rank → the index measured |
| The viewpoint catalogue | it decides the level — 0.50 over 26 sections, 0.92 over 3 — but the ordering between readers holds down to six → the catalogue |
| The measuring instruments | valid after correction: three had to be restricted, and three conclusions withdrawn → counter-expertise |
In one sentence: the index holds as a measurement, not as a standard. What computing it requires can be measured — attention, its shape, its portability. What imposing it would require — a level, a catalogue, an aggregate quantity — follows from no measurement, and the repository has stopped pretending otherwise.
What holds¶
Five objects, and they are the only thing this work asks to be taken from it. A statistical test applies to someone else's data and answers on its own; it does not ask anyone to trust whoever wrote it.
Three checks to run on a log, in this order. Exchangeability — does the recorded order say anything? — detects nothing in MIND (\(z = +0.12\)) and rejects at \(z = -206\) on Baidu-ULTR; identifiability — is there enough to estimate? — necessary and not sufficient; form — does examination depend on what was clicked above? → MIND · served rank · form test
An exposure measured rather than assumed. \(\eta = 0.88 \pm 0.05\) over 143 documents, by display rather than by click. Cascade is refuted on Baidu-ULTR, by two independent routes. And the \(R^{-\eta}\) law used throughout is the worst of three fits on the measured curve. → measured exposure
And that measurement carries from one page to another. This was the last objection, and it came from this repository: if the attention discount depended on the page's composition, no measurement could be written without describing every page served. With the item held fixed the dependence exists but amounts to 6 %, moving the index by \(0.0025\) — fourteen times less than the \(1/R\) convention it was meant to disqualify. → the page effect
Exact bounds when the order is missing. Composition being known and order not, the exposed index is not determined but constrained: median width \(0.104\) over 232,887 real feeds. → the index measured
And a first real figure, where there were only simulations. Served diversity is 0.50 per user-day — 5.1 effective sections out of 26 — and the catalogue moves that level far more than the feed itself does. With a caveat that must travel with it: the figure uses the rank-blind, label-based form, the one the adversarial test saturates. It is an upper bound, not exposed diversity. → the index measured · the catalogue
What fell¶
About the world, four negative results:
- no public dataset permits the announced measurement — but not for the published reason: EB-NeRD carries both columns, and it is its order column that does not contain the order → the index measured;
- examination is not a cascade on Baidu-ULTR, and it does not follow a power law → measured exposure;
- attention severity does not transport from one surface to another — 1.1 on a results page, 0.04 to 0.11 on a three-tile banner → counter-expertise;
- no counterfactual estimator replaces exploration — doubly robust does worse than the plain estimator → served rank.
About the repository's own proposals, five more:
- an index floor saturates at zero cost — 1.000 for zero content diversity;
- the first fix prescribed polarisation — the optimum of Rao entropy is bimodal → adversarial test;
- "proximity to target resists best" was a scale artefact;
- the "screen fold" explanation is false — rank predicts better than pixels → format and return;
- two thirds of the published format effect were composition — 18 % at fixed rank, 6 % with the item held fixed → the page effect.
What is not settled¶
- The exposed index has never been measured on a real feed — only bounded, to within 0.104. The gap is one of data, not of method: it would take a verifiable rank.
- A section is not a viewpoint. The only real figure the repository holds bears on topical diversity exposed. Two labelling axes of the same corpus agree only at \(\rho = 0.16\). → the catalogue
- Nothing has been validated from the outside. The 256 tests check that the code does what is claimed, not that what is claimed is true, and no reviewer has been through it. This is the one lock the repository cannot open by itself. → call for review
Where this came from, and what was removed¶
The project began from an analogy between quantum decoherence and the collapse of consensus, drew from it a formalism borrowed from statistical mechanics, a recommender algorithm and a regulatory framework. All three have been removed.
The analogy because the audit refuted each of its own claims. The formalism because its only testable prediction — the effect of emotional charge — was measured four times with no effect. The regulatory apparatus because its own measurements hollowed it out: a floor whose level depends on the catalogue, which one served item in 177 suffices to satisfy, and whose aggregate quantity invites treating the margin rather than the most enclosed readers.
What remains is what survived its own checks. The audit keeps the entire record of twenty-seven corrections, including those bearing on the removed halves: it is the record of what this work got wrong, and that is the part one does not delete.
Reproducing¶
Everything runs in Docker; nothing is installed locally.
docker compose run --rm test # 256 tests
docker compose run --rm notebooks # runs the notebooks, regenerates the figures
docker compose up site # http://localhost:8000
The twelve notebooks are executable and produce every figure. Raw logs are not versioned — their own licences, several gigabytes — but their digests are, and everything recomputes identically from them.
Notebook numbers are not contiguous: the gaps are those of the removed chapters, and they are left as they are rather than renumbered, so that the audit keeps designating what it is talking about.