Skip to content

Exposed Diversity Index

Measuring the diversity a recommender feed actually exposes — and testing what we think we know about it.


What this repository contains

An instrument — a measure of the diversity a news feed actually exposes to its reader, computable without access to the platform's code — and the adversarial method that put it to the test: every proposition here is attacked, and whatever falls is published as fallen, including when what falls came from the repository itself.

What remains: twenty-six executable notebooks, 623 tests, twenty-five recorded corrections, and a state of affairs resembling neither what the project announced nor what it believed it had established six chapters earlier.

Exposure measured directly, against exposure estimated through
clicks

The sharpest result in the repository: for six chapters this repository estimated exposure through clicks; a column of the file allowed it to be measured. The estimate overstated decay by 23 %. Figure regenerated by notebook 23.

The verdict, in one page

Object State
The quantum decoherence ↔ consensus collapse analogy refuted, transfer by transfer → audit
The classical formalism it led us to borrow coherent, one parameter calibrated, and its one proper prediction — the effect of emotional charge — tested four times with no effect
The persistence gap between emotional registers does not exist — a selection artefact → extended corpus
The index as it was proposed to the regulator untenable: saturable at zero cost, then circumvented by burial → adversarial test · adversarial rank
The retained form of the index defined and costed, and its rank weighting is portable: the page served shifts it by only 6 % → the page effect
The rank-blind index, on real feeds measured for the first time: 0.47 per user-day, 4.6 effective sections out of 26 → the index measured
The exposed index, on those same feeds bounded, not measured — the order column of the only dataset offering it does not contain the rank; 6 feeds in 10 stay undecidable → the index measured
The algorithm (EDA) on the exact frontier, and redundant: a 1998 heuristic does as well, and the gap fades further under noisy relevance → baselines · blind spots
The measuring instruments valid after correction — three had to be restricted, and three conclusions withdrawn → counter-expertise · the page effect
Exposure, the quantity the whole edifice rests on measurable, and measured: \(0.88 \pm 0.05\) — not \(1.09\) as estimated → measured exposure
Its dependence on the page served, which this repository called disqualifying measured with the item held fixed, and three times smaller than published: 6 % of examination, \(0.0025\) of index → the page effect

In one sentence: the theory did not hold. What holds is a quantity — exposure, the one that decides everything else, and which can be measured rather than assumed: \(0.88\) and not \(1\), with no cascade, and independent of the page served to within 6 %. That is enough to write a standard; what is missing is no longer a method but a verifiable rank — and the one public log claiming to supply it does not.

What holds

Six objects, and they are the only thing this work asks to be taken from it. A statistical test applies to someone else's data and answers on its own; it does not ask anyone to trust whoever wrote it. That is the criterion separating them from the method precepts this repository briefly promoted to the rank of result — wrongly. This page said so, and the audit says why it was an overclaim.

Three checks to run on a log, in this order. Exchangeability — does the recorded order say anything? — detects nothing in MIND (\(z = +0.12\)) and rejects at \(z = -206\) on Baidu-ULTR; identifiability — is there enough to estimate? — necessary and not sufficient; form — does examination depend on what was clicked above? → MIND · served rank · form test

An exposure measured rather than assumed. \(\eta = 0.88 \pm 0.05\) over 143 documents, by display rather than by click. Cascade is refuted on Baidu-ULTR, by two independent routes. And the \(R^{-\eta}\) law used throughout the repository is the worst of three fits on the measured curve. → measured exposure

And that measurement carries from one page to another. This was the last objection, and it came from this repository: if the attention discount depended on the page's composition, no standard could be written without describing every page served. With the item held fixed the dependence exists but amounts to 6 %, moving the index by \(0.0025\)fourteen times less than the \(1/R\) convention it was meant to disqualify. → the page effect

Counterfactual estimators confronted with a ground truth. +2.5 % error against +32 % for the naive estimate — with the diagnostic that forbids celebrating it, an effective sample size of 1,513 out of 4 million impressions. No estimator replaces exploration: doubly robust does worse. → served rank

An exact frontier against which to judge a re-ranker. The repository's filter holds it — 0.0 to 1.0 % of engagement left on the table — but so does MMR, and the price of the standard depends on the reader: 3.8 % when their interests cut across viewpoints, 17.1 % when their preference is a viewpoint. → baselines

And a first real figure, where there were only simulations. Over 232,887 Danish feeds, served diversity is 0.47 per user-day — 4.6 effective sections out of 26. The exposed index is only bounded: the order column of the only dataset offering it does not contain the rank, and six feeds in ten stay undecidable at a floor of 0.40. → the index measured

A data access request that can be verified instead of argued. Four aggregate tables, with no personal data, proven sufficient — for 95 times fewer rows than the log. Two columns have been added since, displays and format, each because a measurement showed what was lost without it. → Article 40 request

What fell

About the world, five negative results:

  • the \(\gamma\alpha > \lambda\) criterion means nothing — satisfied by construction for any content that broke through → calibration;
  • the persistence gap between registers does not exist — ×3.04 against ×2.90 (\(p = 0.53\)) across 440 subjects → extended corpus;
  • it was not diluted by labelling — 40 % noise measured, the gap disappears anyway → blind annotation;
  • no public dataset permits the announced evaluation — but not for the published reason: EB-NeRD carries both columns, and it is its order column that does not contain the orderthe index measured;
  • examination is not a cascade on Baidu-ULTR, and it does not follow a power law → measured exposure.

About the repository's own proposals, five more:

  • an index floor saturates at zero cost — 1.000 for zero content diversity;
  • the first fix prescribed polarisation — Rao's entropy has a bimodal optimum → adversarial test;
  • "target proximity resists best" was a scale artefact — at equal exposed diversity, the choice of measure barely matters;
  • the filter brings nothing a 1998 heuristic does not already bring, except at high floors → baselines;
  • the "screen fold" explanation is false — rank predicts better than pixels → format and return;
  • and two thirds of the published format effect were composition — 18 % at fixed rank, 6 % with the item held fixed, and a simulation carrying no effect reproduces the published figure → the page effect.

What is not settled

  • The exposed index has never been measured on a real feed — only bounded, to within 0.104, over 232,887 Danish feeds. The gap is one of data, not of method, and the request states exactly what would be needed — a verifiable rank.
  • A section is not a viewpoint. The only real figure the repository holds bears on topical diversity exposed. → the index measured
  • The level of the floor is a political decision, as is the viewpoint catalogue. The measurement describes; it does not prescribe.
  • The three annotation coders are instances of the same language model.
  • Nothing has been validated from the outside. The 623 tests check that the code does what is claimed, not that what is claimed is true, and no reviewer has been through it. This is the one lock the repository cannot open by itself. → call for review
  • Nothing demonstrates that human opinion obeys statistical mechanics.

The two instruments

Object State
EDI / IDE Exposed Diversity Index — entropy of served items over the declared reference catalogue, weighted by the attention of each rank, in \([0, 1]\) form defined, never measured on a real feed; its weighting depends on the page served
EDA / ADE Exposed Diversity Algorithm — a recommender filter optimising this index rather than raw engagement on the exact frontier, and redundant: a 1998 heuristic does as well

French acronyms IDE and ADE are used throughout the code and the French documentation; EDI and EDA are their English equivalents.

The regulatory memorandum turns the index into recommendations for national and European regulators under the Digital Services Act — each recommendation carrying the measurement that corrected it.

Where this came from

The project started from an analogy: the larger a quantum system, the faster it decoheres; the larger a population, the harder agreement becomes. That analogy produced measurable quantities, and none of its own claims survived verification. That is the subject of the critical audit, which records seventeen corrections — five of which invalidated a formula, and two of which were discovered by trying to measure.

What survives of the formalism is classical: free-energy landscape, phase transition, hysteresis, barrier crossing.

The three regimes of public opinion: free-energy landscape, stationary distributions,
and the splitting of an initially moderate society

The three regimes of public opinion, obtained by changing two parameters of the same free-energy landscape. Figure regenerated by notebook 04.

Explore

The twenty-six notebooks are executable and produce every figure in the paper. Each stands on its own.

Notebook What it shows
01 — Entropy and purity a coherent superposition has zero entropy; the index and its normalisation
02 — Ising Onsager's critical temperature, recovered numerically
03 — Voter model consensus scaling laws, and why connectivity is not the culprit
04 — Fokker-Planck three regimes of public opinion within one free-energy landscape
05 — Hysteresis the memory of a false belief, and the two ways to erase it
06 — Resonance the \(\gamma\alpha > \lambda\) threshold and the attention limit cycle
07 — Algorithm a frozen feed reopening under annealing
08 — Agent model the filter bubble seen from the individual
09 — Calibration \(\gamma\alpha/\lambda\) measured across 19 public attention episodes
10 — Regime change 14 dated switches, and why the ratio is unidentifiable there
11 — Extended corpus 440 category-derived subjects: the persistence gap does not replicate
12 — Blind annotation 40 % label noise measured, and the double recoding at \(\kappa = 0.92\)
13 — Adversarial test an index floor saturated at zero cost, and the fix that prescribed polarisation
14 — Rank and counterfactual burying diversity, and offline evaluation wrong by 201 %
15 — Adversarial rank all four measures circumvented by order, and position-bias severity estimated
16 — MIND's exploration an order indistinguishable from a shuffle, and five severities from one dataset
17 — Served rank two logs that record the rank, and an estimator judged against the truth
18 — Article 40 request four aggregate tables that suffice, and the proof that they do
19 — Baselines the filter judged against four competitors and against the exact frontier
20 — Counter-expertise five counter-tests, one of which withdraws a published conclusion
21 — Blind spots the test holds under cascade, the power law does not
22 — Form test the third check, and the collider it nearly published
23 — Measured exposure one unread column, and six chapters of estimation made redundant
24 — Format and return cascade refuted twice, and one of my explanations withdrawn

Notebook prose is in French; code, variable names and API are in English throughout.

Reproduce

Everything runs in containers. Nothing is installed on the host.

git clone git@github.com:s-geffroy/Indice-Diversite-Exposee.git
cd Indice-Diversite-Exposee

docker compose run --rm test          # 623 tests
docker compose run --rm notebooks     # regenerate every figure
docker compose up lab                 # JupyterLab on :8888
docker compose up site                # this documentation on :8000
docker compose run --rm latex         # compile the paper to PDF

The English paper is at paper/synthesis_note.tex.

Peer review

This work is open to critical review. Feedback on the formalism, on the index's viability as an instrument, or on the limitations listed in the audit is the most valuable. → Call for review


Code under MIT · Written content under CC BY 4.0 · GitHub repository