Exposed Diversity Index¶
Measuring the diversity a recommender feed actually exposes — and testing what we think we know about it.
What this repository contains¶
An instrument — a measure of the diversity a news feed actually exposes to its reader, computable without access to the platform's code — and the adversarial method that put it to the test: every proposition here is attacked, and whatever falls is published as fallen, including when what falls came from the repository itself.
What remains: twenty-six executable notebooks, 623 tests, twenty-five recorded corrections, and a state of affairs resembling neither what the project announced nor what it believed it had established six chapters earlier.

The sharpest result in the repository: for six chapters this repository estimated exposure through clicks; a column of the file allowed it to be measured. The estimate overstated decay by 23 %. Figure regenerated by notebook 23.
The verdict, in one page¶
| Object | State |
|---|---|
| The quantum decoherence ↔ consensus collapse analogy | refuted, transfer by transfer → audit |
| The classical formalism it led us to borrow | coherent, one parameter calibrated, and its one proper prediction — the effect of emotional charge — tested four times with no effect |
| The persistence gap between emotional registers | does not exist — a selection artefact → extended corpus |
| The index as it was proposed to the regulator | untenable: saturable at zero cost, then circumvented by burial → adversarial test · adversarial rank |
| The retained form of the index | defined and costed, and its rank weighting is portable: the page served shifts it by only 6 % → the page effect |
| The rank-blind index, on real feeds | measured for the first time: 0.47 per user-day, 4.6 effective sections out of 26 → the index measured |
| The exposed index, on those same feeds | bounded, not measured — the order column of the only dataset offering it does not contain the rank; 6 feeds in 10 stay undecidable → the index measured |
| The algorithm (EDA) | on the exact frontier, and redundant: a 1998 heuristic does as well, and the gap fades further under noisy relevance → baselines · blind spots |
| The measuring instruments | valid after correction — three had to be restricted, and three conclusions withdrawn → counter-expertise · the page effect |
| Exposure, the quantity the whole edifice rests on | measurable, and measured: \(0.88 \pm 0.05\) — not \(1.09\) as estimated → measured exposure |
| Its dependence on the page served, which this repository called disqualifying | measured with the item held fixed, and three times smaller than published: 6 % of examination, \(0.0025\) of index → the page effect |
In one sentence: the theory did not hold. What holds is a quantity — exposure, the one that decides everything else, and which can be measured rather than assumed: \(0.88\) and not \(1\), with no cascade, and independent of the page served to within 6 %. That is enough to write a standard; what is missing is no longer a method but a verifiable rank — and the one public log claiming to supply it does not.
What holds¶
Six objects, and they are the only thing this work asks to be taken from it. A statistical test applies to someone else's data and answers on its own; it does not ask anyone to trust whoever wrote it. That is the criterion separating them from the method precepts this repository briefly promoted to the rank of result — wrongly. This page said so, and the audit says why it was an overclaim.
Three checks to run on a log, in this order. Exchangeability — does the recorded order say anything? — detects nothing in MIND (\(z = +0.12\)) and rejects at \(z = -206\) on Baidu-ULTR; identifiability — is there enough to estimate? — necessary and not sufficient; form — does examination depend on what was clicked above? → MIND · served rank · form test
An exposure measured rather than assumed. \(\eta = 0.88 \pm 0.05\) over 143 documents, by display rather than by click. Cascade is refuted on Baidu-ULTR, by two independent routes. And the \(R^{-\eta}\) law used throughout the repository is the worst of three fits on the measured curve. → measured exposure
And that measurement carries from one page to another. This was the last objection, and it came from this repository: if the attention discount depended on the page's composition, no standard could be written without describing every page served. With the item held fixed the dependence exists but amounts to 6 %, moving the index by \(0.0025\) — fourteen times less than the \(1/R\) convention it was meant to disqualify. → the page effect
Counterfactual estimators confronted with a ground truth. +2.5 % error against +32 % for the naive estimate — with the diagnostic that forbids celebrating it, an effective sample size of 1,513 out of 4 million impressions. No estimator replaces exploration: doubly robust does worse. → served rank
An exact frontier against which to judge a re-ranker. The repository's filter holds it — 0.0 to 1.0 % of engagement left on the table — but so does MMR, and the price of the standard depends on the reader: 3.8 % when their interests cut across viewpoints, 17.1 % when their preference is a viewpoint. → baselines
And a first real figure, where there were only simulations. Over 232,887 Danish feeds, served diversity is 0.47 per user-day — 4.6 effective sections out of 26. The exposed index is only bounded: the order column of the only dataset offering it does not contain the rank, and six feeds in ten stay undecidable at a floor of 0.40. → the index measured
A data access request that can be verified instead of argued. Four aggregate tables, with no
personal data, proven sufficient — for 95 times fewer rows than the log. Two columns have been
added since, displays and format, each because a measurement showed what was lost
without it. → Article 40 request
What fell¶
About the world, five negative results:
- the \(\gamma\alpha > \lambda\) criterion means nothing — satisfied by construction for any content that broke through → calibration;
- the persistence gap between registers does not exist — ×3.04 against ×2.90 (\(p = 0.53\)) across 440 subjects → extended corpus;
- it was not diluted by labelling — 40 % noise measured, the gap disappears anyway → blind annotation;
- no public dataset permits the announced evaluation — but not for the published reason: EB-NeRD carries both columns, and it is its order column that does not contain the order → the index measured;
- examination is not a cascade on Baidu-ULTR, and it does not follow a power law → measured exposure.
About the repository's own proposals, five more:
- an index floor saturates at zero cost — 1.000 for zero content diversity;
- the first fix prescribed polarisation — Rao's entropy has a bimodal optimum → adversarial test;
- "target proximity resists best" was a scale artefact — at equal exposed diversity, the choice of measure barely matters;
- the filter brings nothing a 1998 heuristic does not already bring, except at high floors → baselines;
- the "screen fold" explanation is false — rank predicts better than pixels → format and return;
- and two thirds of the published format effect were composition — 18 % at fixed rank, 6 % with the item held fixed, and a simulation carrying no effect reproduces the published figure → the page effect.
What is not settled¶
- The exposed index has never been measured on a real feed — only bounded, to within 0.104, over 232,887 Danish feeds. The gap is one of data, not of method, and the request states exactly what would be needed — a verifiable rank.
- A section is not a viewpoint. The only real figure the repository holds bears on topical diversity exposed. → the index measured
- The level of the floor is a political decision, as is the viewpoint catalogue. The measurement describes; it does not prescribe.
- The three annotation coders are instances of the same language model.
- Nothing has been validated from the outside. The 623 tests check that the code does what is claimed, not that what is claimed is true, and no reviewer has been through it. This is the one lock the repository cannot open by itself. → call for review
- Nothing demonstrates that human opinion obeys statistical mechanics.
The two instruments¶
| Object | State | |
|---|---|---|
| EDI / IDE | Exposed Diversity Index — entropy of served items over the declared reference catalogue, weighted by the attention of each rank, in \([0, 1]\) | form defined, never measured on a real feed; its weighting depends on the page served |
| EDA / ADE | Exposed Diversity Algorithm — a recommender filter optimising this index rather than raw engagement | on the exact frontier, and redundant: a 1998 heuristic does as well |
French acronyms IDE and ADE are used throughout the code and the French documentation; EDI and EDA are their English equivalents.
The regulatory memorandum turns the index into recommendations for national and European regulators under the Digital Services Act — each recommendation carrying the measurement that corrected it.
Where this came from¶
The project started from an analogy: the larger a quantum system, the faster it decoheres; the larger a population, the harder agreement becomes. That analogy produced measurable quantities, and none of its own claims survived verification. That is the subject of the critical audit, which records seventeen corrections — five of which invalidated a formula, and two of which were discovered by trying to measure.
What survives of the formalism is classical: free-energy landscape, phase transition, hysteresis, barrier crossing.

The three regimes of public opinion, obtained by changing two parameters of the same free-energy landscape. Figure regenerated by notebook 04.
Explore¶
The twenty-six notebooks are executable and produce every figure in the paper. Each stands on its own.
| Notebook | What it shows |
|---|---|
| 01 — Entropy and purity | a coherent superposition has zero entropy; the index and its normalisation |
| 02 — Ising | Onsager's critical temperature, recovered numerically |
| 03 — Voter model | consensus scaling laws, and why connectivity is not the culprit |
| 04 — Fokker-Planck | three regimes of public opinion within one free-energy landscape |
| 05 — Hysteresis | the memory of a false belief, and the two ways to erase it |
| 06 — Resonance | the \(\gamma\alpha > \lambda\) threshold and the attention limit cycle |
| 07 — Algorithm | a frozen feed reopening under annealing |
| 08 — Agent model | the filter bubble seen from the individual |
| 09 — Calibration | \(\gamma\alpha/\lambda\) measured across 19 public attention episodes |
| 10 — Regime change | 14 dated switches, and why the ratio is unidentifiable there |
| 11 — Extended corpus | 440 category-derived subjects: the persistence gap does not replicate |
| 12 — Blind annotation | 40 % label noise measured, and the double recoding at \(\kappa = 0.92\) |
| 13 — Adversarial test | an index floor saturated at zero cost, and the fix that prescribed polarisation |
| 14 — Rank and counterfactual | burying diversity, and offline evaluation wrong by 201 % |
| 15 — Adversarial rank | all four measures circumvented by order, and position-bias severity estimated |
| 16 — MIND's exploration | an order indistinguishable from a shuffle, and five severities from one dataset |
| 17 — Served rank | two logs that record the rank, and an estimator judged against the truth |
| 18 — Article 40 request | four aggregate tables that suffice, and the proof that they do |
| 19 — Baselines | the filter judged against four competitors and against the exact frontier |
| 20 — Counter-expertise | five counter-tests, one of which withdraws a published conclusion |
| 21 — Blind spots | the test holds under cascade, the power law does not |
| 22 — Form test | the third check, and the collider it nearly published |
| 23 — Measured exposure | one unread column, and six chapters of estimation made redundant |
| 24 — Format and return | cascade refuted twice, and one of my explanations withdrawn |
Notebook prose is in French; code, variable names and API are in English throughout.
Reproduce¶
Everything runs in containers. Nothing is installed on the host.
git clone git@github.com:s-geffroy/Indice-Diversite-Exposee.git
cd Indice-Diversite-Exposee
docker compose run --rm test # 623 tests
docker compose run --rm notebooks # regenerate every figure
docker compose up lab # JupyterLab on :8888
docker compose up site # this documentation on :8000
docker compose run --rm latex # compile the paper to PDF
The English paper is at paper/synthesis_note.tex.
Peer review¶
This work is open to critical review. Feedback on the formalism, on the index's viability as an instrument, or on the limitations listed in the audit is the most valuable. → Call for review
Code under MIT · Written content under CC BY 4.0 · GitHub repository