Skip to content

Rank and counterfactual: two corrections before any evaluation

A fourth adversary: complying by burying

At rigorously identical composition, moving divergent items to the bottom of the feed yields 10 % more engagement and takes the divergence from 0.525 to 0.630. None of the measures retained so far saw it: they look only at composition, never at order.

Naive evaluation of a re-ranking is off by 201 % at the median

Across sixty item sets, the naive estimate of a diversity filter's cost departs from the true value by 201 % at the median and by up to 851 %. It overstates the cost in 56 cases out of 60 — but understates it in the other 4. A naive figure is not an upper bound.

Both corrections exist, and they are cheap

RADio's rank discount closes the burial loophole. Counterfactual estimators recover the true value to within a point. These are not refinements to add after the evaluation on real data: they decide whether that evaluation will measure anything at all.


Why these two corrections come now

The adversarial test corrected the index's definition twice: measure the items rather than the labels, then do not measure by Rao's entropy, which prescribed polarisation. The next move announced by the roadmap is the evaluation of the algorithm on a public recommendation dataset.

Two assumptions still carried that programme, both implicit:

  1. that an item's position in the feed does not matter;
  2. that a logged feed can serve to evaluate a filter that did not produce it.

Both are false, and the second would invalidate the principal result of the next stage.

1. Burial

A reader consults the first item of a feed far more often than the eighth. A platform held to a diversity floor can therefore comply by placing divergent items at the bottom — changing nothing about its feed's composition.

Eight-position feed Composition Position entropy Rank-aware divergence Engagement
diversity on top [5, 1, 1, 1] 0.774 0.525 1.934
diversity buried [5, 1, 1, 1] 0.774 0.630 2.118

Same multiset, same position entropy — and 10 % more engagement. A standard blind to rank gives that gain away for free.

Rank and counterfactual

Burial at constant composition; the effect of the rank discount by its form, the undiscounted curve being flat by construction; estimators against the true value; and the distribution of replay errors across sixty item sets. Figure regenerated by notebook 14.

The rank discount

RADio weights each position by the attention it receives:

\[Q^*(x) = \frac{\sum_i w_{R_i}\,\mathbb{1}[i \in x]}{\sum_i w_{R_i}} \qquad w_{R_i} = \frac{1}{R_i}\]

The figure's second panel shows it most briefly: at fixed composition, sliding the divergent block from the top to the bottom of the feed leaves the undiscounted measure perfectly flat, while the discounted measures climb from 0.29 to 0.68. Rank-awareness is exactly what closes this loophole.

The choice of discount — reciprocal \(1/R\) or logarithmic \(1/\log_2(R+1)\) — shifts the result, and must therefore be published with it.

2. The declared reference

RADio's second contribution matters as much as the first, and it answers the defect of principle raised by the adversarial test: entropy assumes the uniform is ideal, Rao's entropy assumes separation is, and neither says so.

RADio's five measures are the same divergence applied to different pairs of distributions. What distinguishes them is the choice of reference, and that is what carries the normative value:

Measure Distribution served Reference
calibration feed categories the reader's reading history
fragmentation one reader's feed another reader's feed
activation affective intensity of the feed intensity in the supply
representation viewpoints in the feed viewpoints in the supply
alternative voices minority voices in the feed minority voices in the supply

Which direction of divergence is desirable is not given by the mathematics. A calibration of 0.013 describes a feed perfectly matching what the reader already read: that is a liberal recommender's objective and a deliberative recommender's definition of a bubble. The divergence measures; it does not decide — and that is the best thing it has to offer a regulator, who must then declare what is aimed at instead of hiding it in the choice of a formula.

3. The offline-evaluation trap

Re-rank logged feeds and measure the loss of relevance: that is what the programme announced, and that measurement is wrong.

Logged clicks were not produced by the filter under evaluation. A click depends on the item's relevance and on the exposure it was given. An item the platform buried has few clicks — not because nobody was interested, but because nobody saw it. And a diversity filter promotes exactly those items.

What each estimator measures

Estimator Estimated cost Departure from truth
true cost (known here, never on real data) 6.6 %
naive replay 5.0 % −24 %
IPS 6.6 % −0.0 %
SNIPS 6.6 % −0.1 %
IPS clipped at 5 6.6 % −0.0 %

Replay is the estimator actually used: re-rank the candidates, then sum the observed clicks weighted by the exposure the new ranking would give them. Its bias is structural — the observed click rate already carries the exposure the platform granted, so the estimator applies it twice.

The magnitude, and the sign

Across sixty item sets drawn at identical configuration:

median relative error of replay 201 %
worst case +851 %
overstates the cost 56/60
understates it 4/60

One would like to say the bias is conservative — that it always overstates the cost, so a favourable naive result would remain defensible. It is not. The sign depends on the item set, that is, on data one does not choose.

A naive figure is not an upper bound. It is a figure wrong by a considerable amount and in a direction nothing guarantees.

4. What must be published alongside the figure

An unbiased estimator does not suffice. Three quantities must accompany any counterfactual result, failing which it is neither interpretable nor reproducible.

The propensity model. Everything rests on knowing the policy that produced the data. On a public dataset it is not supplied: it must be modelled, typically by a position bias \(e(R) = R^{-\eta}\). That is an assumption, not a measurement.

The effective sample size. The further the evaluated policy sits from the one that produced the data, the fewer observations remain to estimate it:

Aggressiveness of the re-ranking Effective size (out of 60,000)
none 60,000
moderate 47,660
strong 10,026

An unbiased estimate resting on a few hundred effective observations is not a measurement, it is a number.

The clipping cap, if any. Its choice alone shifts the estimate — from −6.5 % at a cap of 1.2 to −0.0 % with none — and without it the result is not reproducible.

What this changes for the programme

Evaluating the algorithm on real data is not abandoned: it is conditioned.

  1. diversity is measured there by a rank-aware divergence to a declared reference, not by a point value on the feed's composition;
  2. the cost in relevance is estimated by IPS or SNIPS, never by replay;
  3. the propensity model, the effective size and the cap are published with the figure;
  4. the position-bias severity \(\eta\) is estimated, not posited, and its uncertainty is propagated to the result. Positing \(\eta\) by eye costs up to 179 % of error — the order of magnitude of the bias one claimed to correct. → Adversarial rank and severity

Without those conditions, the announced trade-off frontier would chiefly measure the position bias of the platform that produced the data.

The assumptions that remain

The position-bias model — exposure depends only on rank — is an assumption. Everything above therefore moves the problem one step along, from "clicks are labels" to "exposure is modelled by rank". The second statement is far better than the first, and it remains a statement.

Three of RADio's five references further require attributes this repository does not have — affect scores, viewpoint annotations, minority coding. The extended corpus showed what it costs to take an available label for the attribute one would like to measure.

Open leads

  1. ~~Estimate the severity of position bias rather than posit it.~~ → done: \(\hat\eta = 1.013 \pm 0.019\) by a content fixed-effects regression — and a refusal to estimate when the platform did not vary its rankings.
  2. ~~Redo the adversarial test under rank-aware measurement.~~ → done: all four measures are circumvented by burial, and the rank-aware floor doubles the engagement cost.
  3. Compare against tuned baselines — MMR, random re-ranking, popularity — not against the pure engagement filter alone, which is a straw man.
  4. Instantiate the five references on real data, which presupposes first solving the labelling problem the extended corpus documented.

Implementation: ide.radio, ide.offpolicy · Notebook: 14 — Rank and counterfactual · adversarial test · algorithm · roadmap