Skip to content

Technical and ethical regulatory memorandum

For the attention of — digital regulators (national authorities, European Commission)

Subject — thermodynamic stabilisation of the informational space and countering the algorithmic resonance of false information

Status — working document, open to critical review. Numerical values quoted are illustrative values from simulations, not numerical recommendations.


Preamble: what this memorandum can and cannot claim

The underlying reasoning is formalised and numerically verified (models, tests). Exactly one of its quantities has been measured on real data — the ratio \(\gamma\alpha/\lambda\), estimated on public attention series; the other parameters are not calibrated. It therefore proposes a metrological framework and quantities to measure, not thresholds ready to be written into law.

That attempt at measurement in fact forced recommendation 2 to be rewritten twice, and then to conclude that the quantity it targeted is not one a regulator can establish. An argument for measuring before legislating, not for taking the framework as settled.

A second reservation must be stated at the outset. The original thread concluded that "regulation ceases to be arbitrary censorship and becomes an engineering of stability". The phrase is appealing and should be treated with suspicion: an engineering of stability is an intervention in public debate. It may be legitimate, but it must be justified as such, with corresponding democratic safeguards — not naturalised by vocabulary borrowed from thermodynamics.


I. Technical recommendations

The premise common to all three: content verification acts after the kinetics have played out. The models show the phenomenon is governed by structural parameters of the algorithm, not by content taken item by item. Those parameters are what must be made observable.

1. Impose a floor on the Exposed Diversity Index

What is recommended today — after three corrections

The standard. A floor on the entropy of the items served, projected onto the bins of a reference catalogue declared by the regulator, and weighted by the attention each rank commands. Not on announced labels, and not on feed composition alone.

The two quantities to publish with it:

  • the gap between the rank-blind and the rank-aware measure of the same feed — zero for a platform that does not relegate, and the only quantity in this repository that thresholds directly;
  • the signature excess — the gap between two indices computed on the same feed, relative to what an honest catalogue would show, which grows with label/content decoupling.

The price, quantified: a rank-aware floor costs 10.6 % to 20.9 % of engagement depending on the measure, against 5.9 % to 10.7 % for a rank-blind floor that protects nothing.

What the regulator sets: the reference catalogue, the aggregate quantity (the share of the population below the threshold, not the mean), and the level of the threshold — which no computation here determines.

What the regulator must have measured rather than fix: the attention discount, which is not even a property of the surface but of the page served — on real data, a rich format above removes 8.4 points of examination from what follows, at equal rank and comparable height (format and return). Two feeds of identical composition therefore do not expose the same thing depending on the formats used.

The surface's attention discount: It is about 1.1 on a results page and a tenth of that on a three-thumbnail banner, and on its value depends the very existence of the burial loophole. A conventional discount imposed on every surface would be wrong for most. → counter-expertise

What matters less than was written: the choice of diversity measure. At equal actually exposed diversity, entropy and divergence-to-target cost the same. It is the level and rank-awareness that make the standard.

How to publish the figure: as an effective number of viewpoints (\(k^{\mathrm{EDI}}\)) rather than a normalised entropy. "2.6 effective viewpoints out of 4" is intelligible; "0.70" is not.

What remains unverified: this form has never been measured on a real feed, for want of a dataset carrying both the served rank and a viewpoint label. → Article 40 request

Two anticipated objections, and what measurement says about them:

  • "we would have to redesign our engines"no. A diversification heuristic published in 1998 holds a floor of 0.80 with no measurable loss of engagement, and relevance ranking remains the limiting case at zero coefficient. What the standard requires is measuring, not rebuilding. → baselines;
  • "the cost will be prohibitive"it depends on the reader, and that must be conceded: from 3.8 % of engagement for a reader whose interests cut across viewpoints to 17.1 % for a reader whose preference is a viewpoint, at the same 0.90 floor. The standard costs most where it serves most, which is precisely where the objection will be raised.

The three corrections that led here are kept below, in the order they occurred.

Measure. Require very large online platforms to keep the index of individual feeds above a threshold \(H_{\text{critical}}\). Below it, the platform is required to reinject a "cooling flow" of semantically diverse content.

Rationale. A collapsed index is the signature of a zero local social temperature, i.e. a frozen state in the Ising sense. Below that temperature, the memory of false beliefs becomes persistent (notebook 05).

What the regulator must set itself:

  • the reference catalogue \(k\) — without an imposed denominator, the index flatters the most closed feeds;
  • the aggregate quantity — the share of the population below the threshold, not the mean: a satisfactory mean can conceal a wholly enclosed minority;
  • the threshold itself, which remains to be calibrated empirically.

What the regulator should not set: the implementation. The algorithm is one way to meet the objective, not the only one.

Recommendation revised after the adversarial test — the floor must be on Rao

The reservation below, left open, was put to the test by simulation (adversarial test), and it holds beyond what was supposed: a platform able to decouple label from content obtains an EDI of 1.000 — full marks — for strictly zero content diversity, without giving up a point of engagement. It need not even go that far: at half decoupling, the constraint retains only 36 % of its force.

A floor on label entropy is therefore not a tenable standard. The measurement must be made on the items served, not on the labels announcing them.

Corrected a second time — not on Rao's entropy

This recommendation first named Rao's quadratic entropy. That was an error, and of the worst kind: Rao's entropy is the intra-list distance, whose constrained optimum is bimodal. It awards 1.000 to a feed serving the two edges and nothing between, against 0.750 to a spread feed — a Rao floor would prescribe polarisation.

The retained floor is on position entropy: the Shannon entropy of the items served, projected onto the bins of the reference catalogue. It is the index with one substitution — items instead of labels — hence the same reading, the same scale, and a lower compliance cost than Rao's entropy.

The largest gap is published beside the floor. Entropy is nominal: it counts occupied viewpoints without seeing their spacing. The diagnostic covers what it misses, and it is what makes bimodality observable.

And the measure must be rank-aware

A measure bearing on a feed's composition alone is satisfied by burying the divergent items at the bottom of the ranking: at identical composition that yields 10 % more engagement without moving the measure by a point. The floor must therefore bear on a distribution weighted by the attention each rank receives. → Rank and counterfactual

Its price is quantified: the engagement cost doubles — from 8.2 % to 18.9 % for Rao's entropy, from 10.7 % to 20.9 % for position entropy. A platform certified at 0.70 by a blind measure in fact exposes only 0.36. → Adversarial rank and severity

Associated control quantity: the gap between the blind and the rank-aware measure of the same feed. Unlike the excess signature, it compares a measure with itself and is therefore directly thresholdable — it is zero for a platform that does not relegate.

What the regulator must then fix in addition: the reference catalogue, which serves as both grid and unit. It is the same political question as the choice of \(k\), moved one step along.

Complementary provision. Publish both indices on the same feed and monitor the excess signature — the EDI − Rao gap relative to what an honest catalogue would show at the same index. It is zero for an honest platform and grows with gaming. → Adversarial test of the index

Successor to be worked out. A target proximity — divergence between the exposure served and a distribution the regulator publishes — is the only measure tested that makes the intended shape of exposure explicit rather than assumed. It is also the entry point to the field's normative-diversity metrics.

Original reservation, kept on record. The index is gameable: label diversity can satisfy a threshold without diversifying the argument. A credible standard must pair automated measurement with qualitative sampling.

2. Cap the amplification-to-damping ratio

Recommendation revised after measurement

This recommendation was originally written as a prohibition on configurations where \(\gamma\alpha > \lambda\). The empirical calibration shows that formulation to be inapplicable: the ratio exceeds 1 in all 19 measured episodes, under every estimator. This follows logically — an observable attention episode necessarily went through a growth phase. Checking the sign tells you nothing.

Measure. Impose a ceiling on the ratio between a piece of content's amplification rate and its natural damping rate:

\[\frac{\gamma\alpha}{\lambda} \leq \rho_{\max}\]

Rationale. Beyond \(\gamma\alpha = \lambda\), effective damping of the feedback loop becomes negative: the system accumulates energy instead of dissipating it (notebook 06). That regime is the norm, not the exception — measurement places the information ecosystem between 1.5 and 12, median 2.5 to 4.2 (notebook 09). The relevant regulatory quantity is therefore the margin, not the crossing.

Why this remains the most solid recommendation. It presupposes no malicious intent to be demonstrated. At uniform gain, a more emotional item should cross the threshold where a factual one does not: the bias would be mechanical, and auditing \(\gamma\) more enforceable than auditing editorial intent.

A caveat for the record. That last point is not supported by the data: the measurement detects no difference in ratio between accusation content and scientific announcements (\(p \geq 0.13\)). The mechanical argument remains theoretically sound, but a regulator must not present it as demonstrated.

Operational difficulties, now quantified.

  • \(\rho_{\max}\) is not determined by theory. The measurement provides a descriptive reference, not a normative threshold.
  • The estimated value is method-dependent: the median varies by a factor of 1.7 with the fitting window. A threshold anchored to a single value would be contestable; the estimation protocol must be standardised alongside the threshold.
  • The available measurement concerns an ecosystem gain, not a platform's internal \(\gamma\). Reaching the latter requires DSA Article 40 access.
  • The peak method is blind to installed regimes. A second method detects them well but does not identify the ratio on those cases — and a theoretical limitation compounds this: under logistic saturation, \(\gamma\alpha/\lambda\) is unidentifiable regardless of data quality.

A replacement indicator does measure — but proves nothing

Regime-change detection yields two robust quantities the amplification ratio does not: the date of the switch and the lift of the plateau. They are measurable on public data, without a form assumption, and they address what a regulator actually seeks to establish: not the speed of a flare-up, but how long a false belief stays installed.

They do not discriminate between emotional registers, however. A first measurement on twenty-four subjects suggested a ×9.2 versus ×2.9 gap; verification on 440 category-derived subjects reduced it to ×3.04 versus ×2.90 (\(p = 0.53\)), and showed the switching-rate gap to be an audience effect. Blind annotation of the register closed the question: with the label corrected, the switching rate goes from 8.6 % versus 2.7 % to 4.8 % versus 5.1 % (\(p = 1.00\)). The gap was not diluted by approximate labelling, it does not exist.

A regulator can therefore use it to observe a durable switch, not to establish that a category of content produces more of them. → Regime changes · Extended corpus · Blind annotation

3. Throttle super-spreader reach on kinetic anomaly

Measure. Impose dynamic limits on cascading share reach as soon as a propagation anomaly is detected.

Rationale — and an important correction. The original reasoning held that the small-world structure of social networks makes consensus impossible. This is measurably false: consensus time grows as \(N^2\) on a local network and only as \(N\) in mean field — global connectivity accelerates convergence (notebook 03, audit, point 12).

What fragments is not link density but the directional bias of algorithmic micro-fields, together with the homophily that compartments the graph.

This recommendation therefore stands as an emergency measure — slowing a cascade buys time for verification — but should not be presented as the structural remedy. Recommendations 1 and 2 are better supported.


II. Ethical and behavioural recommendations

1. Neutralise the "engagement tax"

Treat the maximisation of retention time through the exploitation of negative emotions as a societal nuisance, on the model of environmental externalities, and create fiscal or legal incentives to decouple the business model from permanent friction.

2. Transparency of social-potential assessment

Guarantee every citizen's right to know the shape of the social potential they are subject to: a legible gauge showing the diversity level of their own feed, and the extent to which their decision space has been curved by algorithmic micro-fields.

Reservation. This right requires measuring individual feeds. The protocol must be aggregative and differentially private, or transparency is paid for in surveillance.

3. A right to thermal noise and algorithmic forgetting

Establish a principle of bias disconnection: the ability to enable, in one click, a "fluid exploration" mode that artificially raises social temperature and disables collaborative filtering, breaking the hysteresis sustained by one's history.

A measured nuance, and it matters. Noise is not monotonically beneficial: beyond a certain level, exposure diversity degrades again (notebook 08). The original thread anticipated this — "injecting thermal noise permanently makes society chaotic and illegible". It is therefore not the quantity of noise that matters but its dosage, which argues for an on-demand mode and cyclical annealing rather than permanent noise.


III. Control framework: a metrology of the informational space

[Platform data feeds]
[Regulator's Fokker-Planck simulator]
          ├──▶ unimodal distribution, centred  ─────▶ compliant
          └──▶ bimodal distribution with no
               central moderation zone         ─────▶ alert, then DSA sanction

"Phase scanners"

Rather than counting reports of false information — a lagging and gameable indicator — the regulator simulates the state of opinion from distributions supplied via platform APIs, and detects phase transitions.

The relevant quantity is not the number of problematic items but the shape of the distribution: sharp bimodality with no central moderation zone characterises a degraded informational space, independently of the content of any single message.

What a log must contain to be auditable

The DSA opens platform data to researchers and regulators (Article 40). This section states what that access must require if it is not to be empty.

The served rank, or failing that the exposure propensity. A click depends on two things: the item's relevance and the exposure it was given. Without the rank the two stay conflated, and no counterfactual evaluation of a re-ranking is possible — not by the regulator, nor by the platform itself. Publishing that column reveals nothing of the ranking code: it says where an item was shown, not why.

What anonymisation must not destroy. Shuffling the display order before publication does not make a log unbiased: the clicks were produced under the real order and carry its bias in full. The shuffle removes only the variable that would allow correction — it makes the log uncorrectable. That is the case of MIND, the reference dataset for news recommendation. → MIND's real exploration

Acceptance check. Before any audit resting on a supplied log, run the within-feed exchangeability test: given the feed, are clicks distributed independently of the recorded position? The test is exact, and its power calibration states what it would have detected. A log that "passes" it — in the sense that its order is indistinguishable from a shuffle — is not auditable counterfactually, and it is better to establish that before the audit than after.

Why this check is not optional. Without it, exposure estimation does not halt: it returns a figure. On MIND the severity estimator accepts the dataset, declares itself identifiable and produces five incompatible severities depending on a mere nuisance parameter, from \(-0.13\) to \(+0.25\), each with a standard error below 0.007. An audit resting on any one of them would be indistinguishable from a correct one.

The requirement is not utopian: a public dataset already meets it. The Open Bandit Dataset publishes the served position and the true propensity of every display, and contains a bucket served by a uniformly random policy. On it, the counterfactual estimate of a never-deployed policy's value lands within 2.5 % of its directly measured value, where the naive estimate is off by 32 %. What is asked here therefore has an industrial precedent, published under a free licence by the platform itself.

Further provision: require a fraction of exploration. The same dataset shows why publishing propensities is not enough. Over 4,077,727 impressions served by an optimised policy, the estimate's effective sample size falls to 1,513, i.e. 0.04 %: a policy very sure of itself renders its own logs nearly unusable for evaluating anything but itself. A serious audit therefore presupposes that a fraction of traffic be served deliberately at random — that is the price, low and quantifiable, of auditability.

Associated control quantity: effective sample size, published with any counterfactual estimate. It states how many observations actually carry the figure, and it alone tells a measurement from a coincidence.

Further provision: standardise the form of the request. What an audit requires fits in four aggregate tables — feed profiles, clicks by rank, (item, rank) cells, exposure by declared viewpoint — which are verified to recompute identically the exchangeability test, position-bias severity and both diversity measures. None contains personal data, and together they weigh about a hundred times fewer rows than the corresponding log.

Standardising this format would serve both sides: it removes the platform's Article 40(5) ground for refusal — counts by rank reveal neither the ranking nor its parameters — and it gives the regulator a deliverable whose compliance can be checked, instead of an access whose scope is negotiated. The low-count suppression threshold must be fixed and published there: on real data, moving from 5 to 20 impressions per cell shifts the estimated severity by 27 %.

What is missing to make this operational

Gap Nature
calibration of \(\gamma\alpha/\lambda\) donemeasured between 1.5 and 12 across 19 public episodes, with its caveats
calibration of \(J\) and \(T\) on real data no procedure proposed
normative value of \(\rho_{\max}\) the measurement describes, it does not prescribe
estimating a platform's internal \(\gamma\) requires DSA Article 40 access
detection of installed disinformation regimes done14 dated changes, QAnon and health disinformation included
identification of \(\rho\) on installed regimes fails: real scatter four times too high, and unidentifiable under logistic saturation
calibration of the persistence indicator done, and negative — the pilot corpus's ×9.2 versus ×2.9 gap failed to replicate across 440 subjects, and blind annotation eliminated it for good. The indicator observes; it does not discriminate
existence of an emotional-charge effect \(\alpha\) four measurements, no effect: amplification, persistence on the pilot then the extended corpus, switching rate
privacy-preserving audit protocol not designed
normative definition of the viewpoint catalogue \(k\) political choice unresolved
resistance of the index to gaming done, and negative — an EDI floor saturates at zero cost (adversarial test); the measurement must be made on Rao's entropy
calibration of the position-entropy floor no procedure — the adversarial test establishes the form of the standard, not its level
choice of the exposure target an unsettled political question, which a divergence measure makes explicit instead of burying
cost in perceived relevance of an index floor quantified in simulation — 0 to 17 % depending on the floor and on the alignment between relevance and viewpoint (baselines); never measured on a real feed
counterfactual evaluation on a public dataset done, and negativeMIND does not retain display rank: exposure is not identifiable there, and estimating it anyway returns five incompatible severities
requiring the served rank in supplied logs proposed here, not instrumented on the regulator's side — but a public dataset already meets it
value of \(\eta\) to adopt for a news feed measured elsewhere, not transportable: 1.10 on a results page, 0.04 to 0.11 on a three-thumbnail banner
fraction of random exploration to require of a platform proposed here; without it an audit's effective sample size falls to 0.04 %
standardised format for a data access request donefour aggregate tables, verified sufficient, with no personal data
actually filing an Article 40 request beyond this repository: Art. 40(8)(a) requires affiliation to a research organisation

This memorandum should therefore be read as a framework to harden, not a ready-to-use mechanism. Its contribution is to name measurable quantities where regulatory debate still reasons in volumes of removed content. The roadmap sets out how each gap could be closed.


See also: critical audit and limitations · roadmap · call for review