MeitY's 50 AI Curation Units: Data Extraction Without Privacy Laws
By The Squirrels·
The Cart Before the Algorithm
India is currently executing one of the most aggressive state-led artificial intelligence integrations in the global south. As of March 2026, the Ministry of Electronics and Information Technology (MeitY) is fast-tracking the deployment of 50 "AI Curation Units" across various central ministries. According to official government output documents, the mandate is clear: parse fragmented government data—ranging from nationwide health statistics to agricultural crop surveys—and feed it into the IndiaAI Datasets Platform (AIKosh) to train domestic AI models.
However, a systemic decode of this rollout reveals a glaring institutional vulnerability. The state is constructing a massive, centralized algorithmic data pipeline before finalizing the statutory guardrails meant to protect the 1.4 billion citizens generating that data.
By prioritizing data extraction over legal compliance, MeitY is engineering a technological reality that outpaces its own regulatory framework. The result is a high-stakes experiment in algorithmic governance, operating in a void of algorithmic accountability.
The Architecture of Extraction: By the Numbers
The deployment of these AI Curation Units has been characterized by rapid financial allocation juxtaposed against delayed bureaucratic execution, culminating in a sudden, aggressive rush in early 2026.
The financial blueprint underscores the state's ambition. In March 2024, the Union Cabinet officially approved the overarching IndiaAI Mission with a massive five-year outlay of ₹10,371.92 crore. The initial momentum was slow; the Union Budget for 2024-25 allocated a modest ₹551.75 crore specifically to kickstart the mission's infrastructure.
However, the fiscal year 2025-26 marked a turning point. The FY26 Budget drastically increased the IndiaAI Mission allocation to ₹2,000 crore—a nearly four-fold increase from the previous year's revised estimates. Government output documents explicitly targeted the establishment of the first 20 AI curation units in central ministries during this fiscal year.
"Finding and filtering these often fragmented data and making them available for AI modelling that can provide unique applications has been a key aim." —MeitY Official Statement
Now, in March 2026, after nearly two years of slow progress—reportedly due to pushback from individual ministries highly protective of their existing data silos—MeitY has announced it is "fast-tracking" the initiative. Credible reports indicate the final 20 of the 50 total units are slated to be established over the next few months. This extraction architecture is further supported by a network of over 174 "Data and AI Labs" established nationwide to support data cleaning and curation.
The DPDP Contradiction: A Regulatory Void
The most critical systemic flaw in the rollout of the 50 AI Curation Units is its direct contradiction with the implementation status of the Digital Personal Data Protection (DPDP) Act.
While MeitY is actively extracting and curating citizen data for AI models, the regulatory body meant to police this exact activity does not fully exist. According to the government's own Outcome Budget for 2025-26, the target for setting up the Data Protection Board's (DPB) digital office—which includes basic recruitment and platform development—is slated for only 75% completion for the fiscal year.
Furthermore, official sources confirm that the government is still in the process of notifying the 25 core rules under the DPDP Act.
Consequently, the demographic, health, and logistical data of 1.4 billion citizens is being funneled into the AIKosh platform without a fully operational, independent regulatory authority in place. There is currently no institutional mechanism equipped to investigate data breaches within these curation units, audit the resulting algorithmic biases, or impose statutory penalties.
The "Non-Personal" Data Fallacy
The narrative surrounding the AI Curation Units is deeply polarized between the state's imperatives for efficiency and civil society's concerns over privacy.
The official government stance frames this initiative as a necessary leap toward data-driven governance and sovereign AI development. IT Minister Ashwini Vaishnaw has publicly noted that these units will be responsible for curating datasets containing "non-personal data such as passport data."
However, privacy advocates and legal experts warn that the infrastructure for data extraction is vastly outpacing the infrastructure for legal compliance. Industry analysts point out that while the Digital India Act and DPDP regulations are meant to establish a new privacy regime, they require urgent and substantial budget allocations to "ensure these frameworks have teeth," including dedicated funding for enforcement bodies.
More critically, technology policy researchers argue that the line between "non-personal" and "personal" data is highly porous.
Without documented, legally mandated "human-in-the-loop" oversight, supposedly anonymized datasets—such as aggregated health statistics or demographic surveys—can easily be re-identified by advanced AI models. When an AI model cross-references "non-personal" geospatial data with "non-personal" agricultural and logistics data, the resulting synthesis can easily pinpoint individual behaviors, vulnerabilities, and identities.
Centralization and the Surveillance Risk
In the absence of a finalized statutory framework, the curation of government data operates in a dangerous legal gray area. Globally, the standard for algorithmic accountability requires strict transparency regarding how models are trained, exactly what data is utilized, and how inherent biases are mitigated. India currently lacks a dedicated algorithmic accountability law.
By centralizing geospatial, health, and demographic data across 50 distinct ministries into a single pipeline (AIKosh), the state is inadvertently building a foundational architecture that could easily be repurposed for mass surveillance.
Experts warn that without strict, legally enforceable data-siloing and purpose-limitation mandates, the curation units risk creating a "data swamp." In this scenario, the sheer volume and velocity of data merging into AIKosh compromise citizen anonymity by default, regardless of the original intent of the collection.
The Aadhaar Echo: Implement First, Regulate Later
For technology policy researchers and institutional critics, the rapid rollout of the AI Curation Units feels distinctly, and uncomfortably, familiar. It perfectly mirrors the early, chaotic days of the Aadhaar implementation.
The Unique Identification Authority of India (UIDAI) began collecting biometric data for the Aadhaar project in 2009. This massive data collection apparatus operated for years before the Aadhaar Act was formally passed by Parliament in 2016.
That "implement first, regulate later" approach led to severe institutional consequences: unchecked mission creep, massive data leaks, and a decade of retroactive, exhausting legal battles in the Supreme Court over the fundamental right to privacy.
By launching 50 AI Curation Units to process the data of over a billion citizens while the Data Protection Board remains literally under construction, MeitY is repeating historical precedent. The state is once again building the sprawling technological architecture of tomorrow on the regulatory void of today.
Conclusion: The Cost of Algorithmic Ambition
The establishment of 50 AI Curation Units across central ministries represents a monumental shift in how the Indian state views and utilizes citizen data. The ₹10,371.92 crore IndiaAI Mission has the potential to drive significant domestic innovation, optimizing everything from agricultural yields to public health logistics.
However, innovation cannot serve as a proxy for accountability.
Extracting, cleaning, and centralizing the data of 1.4 billion people into the AIKosh platform without a fully operational Data Protection Board or notified DPDP rules is a systemic failure of sequencing. It places the burden of risk entirely on the citizen while granting the state unprecedented, unregulated access to behavioral and demographic insights.
Until the statutory frameworks for algorithmic accountability and data protection are fully funded, staffed, and legally enforceable, MeitY's AI Curation Units will remain a high-speed train operating without brakes. The data is already flowing; the question is whether the law will ever catch up
