I. Introduction
In June 2017, the European Commission fined Google €2.42 billion for self-preferencing its own comparison shopping service in search results. The preferencing was visible, stable, and reproducible, and that visibility was not incidental to the legal finding. It was load-bearing.
The question this article addresses is what happens when that visibility disappears.
Generative AI systems, large language models integrated into the products of vertically integrated platforms, do not rank results. They synthesise them. For a user asking an AI assistant which productivity software to use or which insurance provider to consider, no ranking appears. A recommendation does.
The argument developed here is that this shift does not simply make self-preferencing harder to prove under Article 102 TFEU, the standard framing in the existing literature, which emphasises opacity and information asymmetry. It makes the primary legal methodology structurally inapplicable. Effects-based abuse analysis, as developed through Google Shopping, Post Danmark, and Intel, depends on the construction of a counterfactual: what would a neutral, non-dominant actor have done? Generative AI outputs are stochastic, non-reproducible, and baseline-free in a way that makes this counterfactual impossible to construct with the reliability that law requires. The result is not merely an evidentiary gap. It is a doctrinal one.
II. Effects-Based Analysis and the Counterfactual Requirement
Article 102 TFEU prohibits the abuse of a dominant position. Since the mid-2000s, the Commission and the Court of Justice of the European Union have progressively moved toward an effects-based approach to establishing abuse: one concerned not with the form of conduct but with its likely or actual competitive effects. The Intel judgment crystallised this shift: even for conduct that is facially exclusionary, a dominant firm is entitled to demonstrate that its behaviour lacked the capacity to restrict competition.[1]
This effects-based framework relies on the construction of a comparator: a baseline against which allegedly anticompetitive conduct can be assessed. That comparator need not be perfectly observable — Intel and Post Danmark II both permit probabilistic and inferential assessments of foreclosure, and Google Shopping itself relied on an analytically constructed notion of ranking neutrality rather than a naturally occurring benchmark.[2] What the doctrine requires is a principled basis for constructing the comparison. In margin squeeze cases like TeliaSonera, that basis is the margin an as-efficient competitor would need to remain viable.[3] The Commission’s own guidance acknowledges that identifying a realistic counterfactual is central to distinguishing aggressive competition from actionable exclusion.[4]
In Google Shopping, the counterfactual was explicit and empirically constructed. The Commission compared the positions of Google’s own comparison shopping service against those of rivals across millions of search queries. The counterfactual, neutral, relevance-based ranking was both conceptually definable and empirically testable. The conduct was reproducible, the comparator was stable, and the deviation was measurable.[5]
III. Generative AI: Why the Comparator Dissolves
Large language models generate text through probabilistic processes. Given a prompt, the model assigns probability distributions across possible next tokens and samples from those distributions. Identical prompts do not reliably produce identical outputs across sessions, users, and deployment contexts.[6]
For competition law analysis, three consequences follow that go beyond the familiar critique of algorithmic opacity.
The first is the absence of a visible treatment. A ranked list exposes the ordering decision. A synthesised answer absorbs those decisions into a conclusion. When an AI assistant recommends a particular cloud storage service, a particular legal database, or a particular financial product, the user receives an output, not the chain of retrieval, weighting, and generation choices that produced it.
The second is the absence of an agreed neutral baseline. In search, neutrality has an operational, if contested, definition: ranked by relevance signals, not by ownership or commercial relationship. What would a “neutral” large language model say about which accounting software best suits a small business, or which platform’s terms of service are most favourable? The answer depends on training data composition, retrieval augmentation configurations, fine-tuning choices, and post-training alignment processes. None of these has an agreed neutral state. Without a baseline, no deviation can be measured; without a deviation, no preferencing can be established.
The third is non-reproducibility at legally relevant timescales. Competition investigations take years. Large language models are retrained, fine-tuned, updated, and redeployed on rolling cycles that bear no relationship to investigatory timelines. The model generating outputs in the year an investigation concludes is not the model that generated allegedly preferential outputs when a complaint was filed. Evidence gathered through adversarial testing today describes a system that may no longer exist.[7]
IV. Why This Is a Doctrinal Problem, Not Just an Evidentiary One
The existing literature on algorithmic opacity in competition law tends to frame these difficulties as problems of evidence: self-preferencing is hard to prove because the system is opaque and the underlying data is hard to obtain.[8] This framing is accurate but insufficient. It implies that the solution lies in enhanced access rights: deeper discovery, mandatory audits, regulatory access to model weights and training data. But they do not resolve the problem, because the problem is not primarily one of access.
Even with tools that force output reproducibility, the stochasticity problem persists at the level of legal inference. Any individual output is a sample from a probability distribution. A single preferential recommendation may reflect nothing more than sampling variance; a pattern of preferential recommendations may be statistically detectable, but detecting a probabilistic tendency is not the same as establishing an identified act of abuse. Article 102 TFEU is not written in the language of probabilities. It is written in the language of conduct.
This is the structural misalignment. Competition law has historically treated system design and architectural choices as conduct in themselves — training parameters, retrieval configurations, or fine-tuning decisions that favour an undertaking’s own ecosystem could, in principle, constitute the relevant act. But identifying the conduct in this way shifts rather than resolves the problem. Effects-based abuse still requires a counterfactual: what would neutral training or neutral fine-tuning have produced? The baseline problem examined in Section III reasserts itself at the level of conduct, and with equal force.
V. Does the Digital Markets Act Close the Gap?
One might argue that Article 102 TFEU no longer needs to carry this load, because the Digital Markets Act provides a lex specialis for the largest platforms. Article 6(5) DMA prohibits designated gatekeepers from favouring their own services or products in ranking over those of third parties.[9]
But it does not resolve the counterfactual problem for two reasons.
First, Article 6(5) was drafted with ranking systems in mind. Whether Article 6(5) reaches generative AI outputs that produce no visible ranking, only a synthesised answer, remains legally unsettled.
Second, the DMA applies only to platforms formally designated as gatekeepers under its quantitative and qualitative thresholds. A generative AI system operated by a dominant but non-designated firm, or by an AI-native entrant not yet meeting those thresholds, falls outside the DMA’s scope entirely and back within Article 102’s reach, where the counterfactual problem remains unresolved.
VI. Conclusion
The counterfactual is not a procedural nicety in effects-based abuse analysis. The Google Shopping case could be built on a stable ranking, a definable neutral baseline, and reproducible evidence gathered over a fixed period. Each of those conditions was produced by the underlying technology.
Generative AI self-preferencing, if it occurs, operates as a distributional property of a probabilistic system, a tendency embedded at training, not a ranking decision made at retrieval. The counterfactual, in this context, cannot be constructed. And without a counterfactual, effects-based analysis cannot function.
It is an argument that the current framework, designed for a world of deterministic, observable, comparable conduct, is not equipped to reach it without structural reform. What it cannot take is the form of applying existing doctrine and hoping that the technology eventually becomes observable enough to fit it.
The law needs a new comparator. It does not yet have one.
[1]Case C-413/14 P Intel Corporation Inc v Commission [2017] ECLI:EU:C:2017:632, paras 133–140; Case C-23/14 Post Danmark A/S v Konkurrencerådet [2015] ECLI:EU:C:2015:651, para 65.
[2]Case C-413/14 P Intel Corporation Inc v Commission [2017] ECLI:EU:C:2017:632, paras 133–140; Case C-23/14 Post Danmark A/S v Konkurrencerådet [2015] ECLI:EU:C:2015:651, para 65.
[3]Case C-52/09 Konkurrensverket v TeliaSonera Sverige AB [2011] ECLI:EU:C:2011:83, para 44.
[4]European Commission, ‘Guidance on the Commission’s enforcement priorities in applying Article 82 of the EC Treaty to abusive exclusionary conduct by dominant undertakings’ [2009] OJ C45/7, para 21.
[5]European Commission, ‘Google Search (Shopping)’ (Case AT.39740) Commission Decision of 27 June 2017, paras 344–395.
[6]Tom B Brown and others, ‘Language Models are Few-Shot Learners’ (2020) 33 Advances in Neural Information Processing Systems 1877, 1877–1878. For a legal-oriented discussion of stochasticity in AI systems, see also Crémer, de Montjoye and Schweitzer (n 6) 112–113.
[7]Jacques Crémer, Yves-Alexandre de Montjoye and Heike Schweitzer, ‘Competition Policy for the Digital Era’ (European Commission 2019) 117–119.
[8]See, paradigmatically, Ariel Ezrachi and Maurice E Stucke, Virtual Competition: The Promise and Perils of the Algorithm-Driven Economy (Harvard University Press 2016) chs 2–3; Nicolas Petit, Big Tech and the Digital Economy: The Moligopoly Scenario (Oxford University Press 2020) 198–202.
[9]Regulation (EU) 2022/1925 of the European Parliament and of the Council of 14 September 2022 on contestable and fair markets in the digital sector [2022] OJ L265/1, Art 6(5).
