Skip to content

The index, measured on a real feed

The figure missing since day one

Over 232,887 real feeds of the Danish daily Ekstra Bladet, served diversity is 0.47 per user-day — 4.6 sections equally served out of 26. The repository had never published this number outside a simulation.

But the rank column does not contain the rank

EB-NeRD finally carries the two columns missing everywhere else: the served list and a label. Yet the repository's first check says the recorded order carries no position information — \(z = +1.05\), \(p = 0.29\) — and this is not a power failure: this log would have detected a severity 134 times smaller than the one measured on Baidu-ULTR.

What remains possible without rank: convicting, rarely acquitting

Composition being known, the exposed index is bounded: median width 0.104. At a floor of 0.40, 32.4 % of feeds breach it whatever their ordering — an enforceable finding without the missing column — but 59.7 % remain undecidable. That is the exact price of what the platform holds and does not publish.


A gap wrongly stated

Six chapters of this repository rested on one sentence:

No public dataset carries both the served rank and an interpretable viewpoint label. → MIND · served rank

It is false. EB-NeRD — Ekstra Bladet's log published for the RecSys Challenge 2024 — gives article_ids_inview, the served list, and attaches to each article a declared section, category_str. Both columns are there.

What is missing lies elsewhere, and is more interesting: the order column does not contain the order.

The index measured

Served diversity measured on real feeds, what the unknown order leaves undetermined, the three verdicts available without the rank column, and their dependence on the assumed severity. Figure regenerated by notebook 26.

1. The first check removes the order

Log Exchangeability Reading
Baidu-ULTR \(z = -206\) the recorded order is the served rank
MIND \(z = +0.12\) the order says nothing
EB-NeRD \(z = +1.05\) (\(p = 0.29\)) the order says nothing

A test that fails to reject says nothing until one knows what it could have rejected. This log would have detected a severity of 0.0066, where Baidu-ULTR measures 0.88: the silence is not lack of power, it is an answer.

article_ids_inview is a served set. The exposed index — the one this repository proposes to a regulator — is therefore not measurable here. The rank-blind index is, and it never had been on real data.

2. Served diversity, finally quantified

Window Median Quartiles Effective viewpoints Below 0.40 Below 0.50
served feed 0.417 [0.372, 0.489] 3.89 of 26 35.0 % 78.8 %
user-day 0.466 [0.411, 0.507] 4.56 of 26 18.8 % 71.2 %

The second row is the one that counts: it is the window the protocol prescribes, and the regulated quantity there is the share of the population below the floor, not the mean.

This figure does not say whether that is little or much — the level of a floor is a political decision, and the repository does not settle it. It gives the order of magnitude missing from any discussion of a floor: a reader receives, in one day, the equivalent of four to five sections equally served out of twenty-six available.

3. What the unknown order leaves undetermined

The order is missing, the composition is known: the exposed index is not determined, it is constrained. ide.entropy.exposed_index_bounds enumerates every distinct ordering and returns the exact bounds — not a heuristic.

Floor Certainly below Certainly above Undecidable
0.30 8.2 % 50.6 % 41.3 %
0.40 32.4 % 7.9 % 59.7 %
0.50 91.5 % 0.7 % 7.7 %

Without the rank one can convict; one almost never acquits. A third of feeds breach a floor of 0.40 whatever their ordering: an enforceable finding obtained without the missing column. But six feeds in ten remain undecidable, and that is what its absence costs.

Sensitivity. The width of the bounds grows with the assumed severity — 0.062 at \(\eta = 0.5\), 0.104 at 0.88, 0.127 at 1.1. The undecidable share, however, peaks near 0.9 and falls back: at higher severity the whole interval slides below the floor and the feed becomes decidable again — by conviction.

4. What this changes for the access request

The Article 40 request asked for "the served rank". This chapter shows that is not enough: a platform can supply an order column that is not the rank, without lying and without anyone seeing it. What must be asked for is the verifiable rank, and the exchangeability test is what verifies it — on the delivered data, before any other measurement.

That also makes the repository's first check more than a methodological preliminary: it is an admissibility clause.

Reservations

category_str is a section — "nyheder", "sport", "krimi" — not a viewpoint in the index's sense. The measurement bears on topical diversity exposed. It is the substitute already used on MIND, and must be read as such: two articles in the same section may argue opposite theses, and two different sections may say the same thing.

The severity \(\eta = 0.88\) is transported from Baidu-ULTR, a search engine. Chapter 25 established that it transports from page to page — not from platform to platform. Hence the sensitivity test, which does not reverse the conclusion.

The 26-section catalogue is the corpus's own, not a reference catalogue imposed by a regulator. Changing \(k\) changes the index level, never its ordering.

The exact bounds cover only feeds of ten items or fewer, 64 % of the log: beyond that, enumeration ceases to be practicable.

Finally, EB-NeRD is distributed for research use only. The raw log is not versioned; only the digest of composition signatures is — counts per section, sorted and anonymous — which four tests verify returns the same figures.


Notebook: 26 — The index measured · MIND · served rank · the page effect · Article 40 request