CAGE


Written by Devanshi Sharma holds an LL.M. in Corporate and Business Laws from Gujarat National Law University (GNLU).

Large language models are built on a premise that sits uneasily against the core architecture of data protection law. Privacy regimes from the GDPR to India’s DPDP Act 2023[1] are organised around limiting principles which provides to collect only what is necessary, use it only for the stated purpose and give the individual a mechanism to withdraw or erase. Model training runs on the opposite logic. More data scraped from more sources almost always produces a better model and the individual whose comment, biography or photo caption gets swept into a training corpus is rarely asked and rarely told. The result is not a hypothetical clash of principles. Regulators, courts and computer scientists are now working through the same underlying question from three different directions namely, the data minimization, consent and the right to be forgotten. none of the three has produced a clean answer yet.

Data minimization meets a scale-hungry technology

Article 5(1)(c) of the GDPR[2] requires that personal data be “adequate, relevant and limited to what is necessary” for the purpose of processing. That instruction assumes that someone can specify in advance what data a given purpose requires. When foundation model are trained by the developers , they can not predict that which peice of information will help the AI model to train in future to perform tasks. Therefore they tend to train the model using a very large amount of data  Purpose limitation is similarly strained: a dataset scraped for one model is routinely reused, filtered and recombined for the next one and the “purpose” specified at collection time rarely anticipated that a general-purpose conversational agent would be built from it years later.

This is not merely an academic mismatch. It is the substance of the finding that Italy’s data protection authority, the Garante, made against OpenAI. Its analysis is worth walking through because it is the most developed regulatory record to date on how minimization and lawful-basis requirements apply to LLM training specifically.

Consent under strain: the Garante’s case against OpenAI

On 30 March 2023 the Garante issued an emergency order halting ChatGPT’s processing of Italian users’ personal data and it did so on grounds that went directly to the training pipeline.[3] The core finding was that OpenAI had processed personal data to train its models without an adequate legal basis under Article 6(1) GDPR[4], alongside separate transparency failures and the absence of any age-verification mechanism for a service capable of generating content unsuited to minors. Following measures implemented by OpenAI concerning transparency, data-subject rights, the legal basis for algorithmic training and age-related safeguards, the Garante allowed ChatGPT to become accessible again in Italy.

The story did not end there. In December 2024 the Garante fined OpenAI €15 million, again centred on the finding that personal data had been used to train ChatGPT before any adequate legal basis existed. the authority located the violation as early as the model’s public release. The regulator also rejected OpenAI’s attempt to rely on a legitimate-interest assessment to justify the training and separately faulted the company for failing to report a March 2023 data breach. As a remedy, the Garante exercised a power it had not previously used by ordering a six-month public information campaign explaining to Italians how to object to their data being used for generative AI training.[5]

However, the enforcement decision did not survive judicial challenge. On 18th march 2026, the Rome tribunal upheld OpenAI’s contentions and annulled the grante’s decision.[6]The grante removed temporarily the 2024 decision from its website. Importantly, the judgment does not amount to a finding that the underlying processing complied with the GDPR rather, the case was resolved on the question of the Garante’s competence to adopt the contested decision.”

Two things stand out from that record. First, consent in the ordinary sense was never really on the table to which OpenAI’s defence rested on legitimate interest and not on having asked users for permission because asking billions of internet-scraped data subjects for opt-in consent before scraping is not operationally feasible. Second, even the fallback basis of legitimate interest failed scrutiny once a regulator actually tested it against the GDPR’s balancing test. That is a meaningful signal for any jurisdiction weighing how to authorise AI training on personal data: we didn’t have another practical option is not on its own a lawful basis.

The right to be forgotten was built for a search index, not a set of weights

The right most directly implicated by AI training is erasure and its origin story matters. India’s right to be forgotten appreared in DPDP act 2023, where under section 12[7] right to erasure is provided. But despite this two gaps stands out. First, nothing in the DPDP Act or the 2025 Rules[8] distinguishes personal data absorbed during AI pretraining from personal data held in an ordinary database even though the two present entirely different erasure mechanics, leaving practitioners to reason by analogy from provisions drafted for conventional processing. Second, the right to be forgotten as Indian courts have built it runs against intermediaries operating a queryable index, search engines and legal databases and has not yet been tested against a foundation-model developer whose “index” is a trained parameter set rather than a database.

Where this leaves the law

None of the frameworks currently on the books were written with parameter-entangled models in mind and it shows. A durable answer will probably need to combine three things that presently sit in different silos.  A lawful-basis standard for training that does not quietly default to “it was public” as a substitute for legitimate-interest balancing, erasure obligations that are calibrated to what unlearning can actually deliver today rather than to a search-index-era assumption of clean deletion and cross-border convergence, since a model trained once is deployed everywhere making jurisdiction-by-jurisdiction consent rules only partially effective at the point they matter most before training begins. Until that convergence happens, enforcement actions will keep doing the doctrinal work that legislatures have not yet finished.

Footnotes

  1. Digital Personal Data Protection Act 2023 <https://www.meity.gov.in/static/uploads/2024/06/2bf1f0e9f04e6fb4f8fef35e82c42aa5.pdf>
  • Digital Personal Data Protection Act 2023, s 12(3).
  • Digital Personal Data Protection Rules 2025

[1]Digital Personal Data Protection Act 2023 <https://www.meity.gov.in/static/uploads/2024/06/2bf1f0e9f04e6fb4f8fef35e82c42aa5.pdf>

[2]General Data Protection Regulation (EU) 2016/679, art 5(1)(c) https://gdpr-info.eu/art-5-gdpr/

[3]Garante per la protezione dei dati personali, Provvedimento n 112 del 30 marzo 2023 [doc web n 9870832] <https://www.garanteprivacy.it/home/docweb/-/docweb-display/docweb/9870832>

[4]General Data Protection Regulation (EU) 2016/679, art 6(1) <https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32016R0679>

[5]Garante per la protezione dei dati personali, COMUNICATO STAMPA – ChatGPT, il Garante privacy chiude l’istruttoria. OpenAI dovrà realizzare una campagna informativa di sei mesi e pagare una sanzione di 15 milioni di euro.https://www.garanteprivacy.it/home/docweb/-/docweb-display/docweb/10085432

[6]Tribunale di Roma, sentenza n 4153/2026 (18 March 2026),Wilson Sonsini, ‘OpenAI Prevails in Landmark Italian AI and GDPR Enforcement Case’ (31 March 2026) https://www.wsgr.com/en/insights/openai-prevails-in-landmark-italian-ai-and-gdpr-enforcement-case.html, Cross-Border Data Forum, ‘Generative AI and GDPR Enforcement in Europe: A Lot of Noise, One Fine, Zero Survivors’ (19 March 2026) https://www.crossborderdataforum.org/generative-ai-and-gdpr-enforcement-in-europe-a-lot-of-noise-one-fine-zero-survivors/

[7]Digital Personal Data Protection Act 2023, s 12(3).

[8]Digital Personal Data Protection Rules 2025

Leave a Reply

Discover more from CAGE

Subscribe now to keep reading and get access to the full archive.

Continue reading