Skip to content

Adversarial rank and severity: what a rank-blind standard lets through

A platform certified at 0.70 exposes only 0.36

Judged on ordered feeds rather than compositions, all four measures of the adversarial test certify a diversity the reader does not receive. Rao's entropy certifies 0.750 for an exposed diversity of 0.355; position entropy, 0.774 for 0.443.

A rank-aware floor closes it — and it costs

Exposed diversity then reaches the floor, the divergent items moving up the feed. The engagement price doubles: from 8.2 % to 18.9 % for Rao's entropy, from 10.7 % to 20.9 % for position entropy. A standard that cost no more would close nothing.

Position-bias severity can be estimated, and it had to be

\(\hat\eta = 1.013 \pm 0.019\) over 40,000 impressions, by a content fixed-effects regression. Positing \(\eta\) by eye costs up to 179 % of error — the very order of magnitude of the bias one claimed to correct.


Two debts, settled together

Notebook 14 left two explicit debts, and they share a root: rank.

The four measures compared in the adversarial test had been tested on compositions — what share of attention goes to which viewpoint — never on ordered feeds. And the counterfactual estimators rested on a position-bias model whose severity was posited, not measured.

Adversarial rank and severity

Displayed diversity against what is actually exposed; the gap by the standard's stringency; severity recovered, and the threshold beyond which it ceases to be; and what a posited \(\eta\) costs. Figure regenerated by notebook 15.

1. The adversarial test, on ordered feeds

The feed has \(n\) positions to fill from a catalogue of \(k\) viewpoints, hence \(k^n\) possible feeds — 65,536 here. They are all enumerated: the optimum is exact, not a heuristic's output. These are still negative results, and an optimum missed by a solver would look exactly like a standard that holds.

Under a rank-blind floor

Measure Cost Displayed Exposed Feed served
Rao (ILD) 8.2 % 0.750 0.355 00000033
position entropy 10.7 % 0.774 0.443 00000123
Gaussian ILD unattainable
target proximity 5.9 % 0.750 0.628 00000012

The optimal feed always has the same shape: six items of the preferred viewpoint, then the divergent ones relegated to the last positions, where attention no longer reaches.

Under a rank-aware floor

Measure Cost Displayed Exposed Feed served
Rao (ILD) 18.9 % 0.938 0.702 00033300
position entropy 20.9 % 0.953 0.702 00013122
target proximity 10.6 % 0.750 0.701 00100200

The loophole is closed, and the cost doubles. That is the standard's price, and it had to be quantified before being proposed.

The gap grows with the stringency

Blind floor 0.40 0.50 0.60 0.70 0.80
Rao (ILD) 0.262 0.306 0.350 0.395 0.397
position entropy 0.173 0.249 0.263 0.331 0.319
target proximity 0.000 0.066 0.066 0.122 0.166

The more diversity a blind standard demands, the more profitable burying it becomes. This is not the artefact of one particular floor.

Note that target proximity resists best — a null gap at floor 0.40, 0.122 at 0.70 — where Rao's entropy reaches 0.395. That is the third respect in which it stands out, after being the only one to make the intended shape of exposure explicit.

A gap that is, this time, thresholdable

The adversarial test had to give up thresholding the gap between two different indices: the index and Rao's entropy are not on the same scale, and an honest feed already showed 0.36.

The gap measured here is of another nature: it is the same measure applied twice to the same feed, once blind to rank and once accounting for it. It is zero for a feed that does not relegate, and hence directly interpretable.

2. Estimating position-bias severity

The model posits \(P(\text{click}) = R^{-\eta}\,g(i)\), hence

\[\log \mathrm{CTR}(i, R) = \log g(i) - \eta \log R\]

Relevance \(g(i)\) is a content fixed effect: we do not try to estimate it, we eliminate it by centring within each item. What remains is the only variation identifying \(\eta\) — that of the same item seen at different ranks. This is the simplest form of intervention harvesting, and it requires no experiment.

True \(\eta\) 0.40 0.70 1.00 1.30 1.60
\(\hat\eta\) 0.409 0.693 1.005 1.315 1.649
standard error 0.0045 0.0064 0.0085 0.0121 0.0192

What conditions the estimate

Ranking exploration \(\hat\eta\) Standard error Identifiable
0.00 no
0.02 1.595 0.247 yes, but worthless
0.05 0.800 0.092 yes
0.15 1.031 0.036 yes
0.50 1.004 0.011 yes

Under a deterministic policy, \(\eta\) is not in the data. No item changes rank, so there is no variation to exploit, and the estimator refuses to return a figure rather than invent one. That is precisely the case where positing the value instead of estimating it would be undetectable in the results.

Under thin exploration it does return one — but the standard error says it is worth nothing.

Since corrected: this check is necessary and not sufficient

It counts rank variation without saying where it comes from. A log whose order was shuffled before publication shows more of it than any real platform, and the estimator declares itself identifiable there before returning five incompatible severities. The missing check is an exchangeability test, and it must precede estimation. → MIND's real exploration

3. Why it had to be estimated

Posited \(\eta\) 0.5 0.8 1.0 1.2 1.5 2.0
estimated cost 2.7 % 4.9 % 6.6 % 8.6 % 11.9 % 18.4 %
error −59 % −26 % +0.6 % +30 % +80 % +179 %

The true cost is 6.6 %. Positing \(\eta\) wrongly costs up to 179 % of error — that is, the order of magnitude of the 201 % bias the counterfactual correction claimed to eliminate.

With \(\eta\) estimated from the data — \(1.013 \pm 0.019\) — the cost sits between 6.6 % and 6.9 % against a true value of 6.6 %.

Correcting does not suffice: the correction's parameter must be estimated, and its uncertainty published. Otherwise one replaces a known bias with a bias of the same order that merely looks like a correction.

What this changes for the programme

Memorandum recommendation 1 already required a rank-aware measure. This work gives its price — the engagement cost doubles — and a control quantity: the gap between the blind and the rank-aware measure of the same feed, which is zero for a platform that does not relegate.

For counterfactual evaluation, a fourth requirement joins the three already set: \(\eta\) must be estimated and its uncertainty propagated, and the dataset's exploration must be checked first — it is what decides whether estimation is possible at all.

Withdrawn conclusion: \"target proximity resists best\"

This page compared the four measures at the same nominal floor of 0.70. They do not live on the same scale: a 0.70 floor on target proximity does not demand what a 0.70 floor on entropy demands.

Compared at equal actually exposed diversity, both measures cost the same, within 0.6 % of the exact bound. What matters is not the choice of measure but the level required and rank-awareness. → counter-expertise

The assumptions that remain

The form \(e(R) = R^{-\eta}\) is posited, and only its severity estimated. An examination model depending on the content, or on what the reader has already seen, would give different propensities.

Exhaustive enumeration further bounds the feed sizes studied: eight positions over four viewpoints. Nothing guarantees the observed behaviour carries to a fifty-item feed, and claiming so would require an optimisation whose exactness would no longer be guaranteed.

Open leads

  1. Check that burial carries to scale. Beyond a few tens of thousands of feeds an approximate optimisation would be needed — and it would then have to be established that it does not invent the result.
  2. Enrich the examination model: confidence or cascade models make attention depend on what the reader has already consulted, not on rank alone.
  3. ~~Measure a public dataset's actual exploration before drawing anything from it: it decides whether \(\eta\) is identifiable there.~~ → done, and negative: the order recorded in MIND is indistinguishable from a shuffle.
  4. Compare against tuned baselines — MMR, random re-ranking, popularity — which remains the evaluation programme's outstanding debt.

Implementation: ide.ranking, ide.offpolicy · Notebook: 15 — Adversarial rank and severity · adversarial test · rank and counterfactual · memorandum