New batches starting this week Β· Limited seats

India's DPDP Act for AI Teams: What Engineers Building AI Applications Should Know

An engineering guide to India's Digital Personal Data Protection Act and the DPDP Rules, 2025 for teams building LLM, RAG and agent applications, translating each obligation into controls you can build and test.

DPDP Act engineering controls for AI: map personal data, notice and consent, minimise and redact, erase everywhere, breach response
Last updated Β· 14 min read Β· 3,162 words

India's Digital Personal Data Protection Act, 2023 (DPDP Act) treats AI applications like any other software that processes personal data: the prompt, the retrieved chunk, the log line and the training record all count. For AI teams, DPDP compliance is mostly an engineering problem: know where personal data flows through prompts, logs, RAG indexes, training sets and agent memory, tag it with its consent and purpose, and make erasure, retention and breach response reach every one of those stores. The DPDP Rules, 2025 were notified in November 2025, and most business obligations apply from May 2027, so design them in now.

Not legal advice. This is a simplified engineering guide based on the Act and Rules as published at the time of writing. It does not cover sectoral rules (RBI, IRDAI, SEBI, health) or your specific facts. Work with your legal and privacy team, and check the official text from MeitY, as rules and timelines can be amended.

The DPDP Act and Rules in one page

Parliament passed the Act in August 2023, but it became operational only when the Ministry of Electronics and Information Technology (MeitY) notified the Digital Personal Data Protection Rules, 2025, in the Gazette dated 13 November 2025. The Rules commence in three phases:

PhaseDate (as notified)What switches on
Immediate13 November 2025Definitions and the provisions that constitute and run the Data Protection Board of India (Rules 1, 2, 17 to 21)
12 months13 November 2026Registration and obligations of Consent Managers (Rule 4)
18 months13 May 2027Notice, security safeguards, breach intimation, retention and erasure, children's data, Significant Data Fiduciary duties, rights, cross-border transfer (Rules 3, 5 to 16, 22, 23)

In early 2026 MeitY consulted industry on shortening the timeline for Significant Data Fiduciaries. Law-firm trackers reported it as a proposal, not notified law, so confirm the current position with counsel before planning against either date.

The vocabulary engineers need

  • Data Principal: the individual the data is about; for a child (under 18), this includes the parent or lawful guardian.
  • Data Fiduciary: whoever decides the purpose and means of processing, usually the company running the AI product.
  • Data Processor: anyone processing on the fiduciary's behalf, such as a cloud host, model API provider or vector database SaaS. The fiduciary stays responsible and may engage processors only under a valid contract.
  • Significant Data Fiduciary (SDF): a fiduciary the government notifies based on factors such as volume and sensitivity of data. SDFs need a Data Protection Officer based in India, an independent data auditor, and periodic impact assessments and audits. The Rules add that an SDF must verify its technical measures, including algorithmic software, do not pose a risk to Data Principals' rights, a clause aimed squarely at AI systems.
  • Consent Manager: a Board-registered entity offering individuals one interoperable platform to give, manage, review and withdraw consent.
  • Data Protection Board of India: inquires into breaches and complaints and imposes penalties.

What the law requires, in plain terms

  • Lawful ground: consent, or a listed "legitimate use" such as data voluntarily provided for a specified purpose, legal obligations or medical emergencies.
  • Notice and consent: consent must be free, specific, informed, unconditional and unambiguous, given by clear affirmative action, and limited to necessary data. Withdrawing must be as easy as giving.
  • Purpose limitation: "customer support" consent does not quietly become "train our model" consent.
  • Rights: access, correction, completion, updating and erasure, grievance redressal (answered within 90 days under the Rules) and nomination.
  • Security and breaches: reasonable safeguards; on a breach, inform affected individuals and the Board without delay, with a detailed report to the Board within 72 hours of becoming aware.
  • Children: verifiable parental consent; no tracking, behavioural monitoring or targeted advertising directed at children, subject to notified exemptions.
  • Cross-border: transfers are allowed unless the government restricts a country by notification; stricter sectoral laws still apply.

The Schedule to the Act sets maximum penalties per breach: up to β‚Ή250 crore for failing to take reasonable security safeguards, and up to β‚Ή200 crore each for failures on breach notification or children's obligations (source: the Schedule to the Digital Personal Data Protection Act, 2023). The Board decides actual amounts.

Why AI systems make DPDP harder

A CRUD app keeps personal data in known tables; an LLM application copies it everywhere. A pasted PAN lands in the prompt, request log, traces and possibly the provider's abuse-monitoring store. An email becomes chunks, vectors and cache entries. An agent's "memory" that a customer recently lost a job is personal data too. A fine-tuning job bakes examples into weights, where row-level erasure is impractical. This is data engineering for RAG and agents, done with privacy as a first-class requirement.

Step 1: map personal data across the AI pipeline

You cannot erase data you do not know you have. Map every AI feature's stores:

User --> App API --> [PII redactor] --> Model API
            |              |                |
            v              v                v
      prompt/resp log   trace store   provider logs
            |
            v
   RAG ingest --> chunks + vectors --> cache
            |
            v
   agent memory    eval sets    fine-tune data
            |
            v
         backups (all of the above)

For each store record what personal data can land there, the lawful ground and purpose, access, retention, region and provider, and how deletion works. Evaluation sets and red-team transcripts count too; teams routinely copy production conversations into them. Keep the map in your AI governance use-case inventory so it is reviewed whenever the feature changes.

Purpose limitation is enforceable only when every record carries its purpose with it:

  • Give every document, chunk, memory item and log record a principal_id (or pseudonymous key), a purpose code, a consent_id or legal_basis, and an expires_at.
  • Filter on that metadata at retrieval time. A support assistant retrieves only support-purpose chunks; a marketing agent never sees chunks tagged for credit decisions.
  • Consume consent-withdrawal events (from your consent service or, later, a Consent Manager) and mark affected records ineligible immediately, deleting them asynchronously.
  • Store the consent record itself: notice version, timestamp, channel.

In pgvector or any vector store this is ordinary metadata filtering; the discipline is making it mandatory in the retrieval layer, not optional per prompt.

Step 3: make erasure propagate everywhere

When consent is withdrawn or the purpose is served, the Act expects erasure, including by your processors. Design one erasure job keyed on principal_id that fans out to:

  1. Source systems: CRM, ticketing, document stores.
  2. Chunks and vectors: delete by metadata; rebuild indexes if your store does not truly delete.
  3. Caches: semantic and response caches holding the person's data.
  4. Agent memory: see memory in AI agents for structuring memory so it can be deleted per user.
  5. Logs and traces: delete or irreversibly redact prompt content.
  6. Processors: vendor deletion APIs or the contracted process, with evidence.
  7. Backups: typically short expiry plus a "re-delete after restore" runbook step, agreed with legal.

Fine-tuned weights are the hard case: avoid training on identifiable personal data, prefer de-identified or synthetic examples, and keep a training manifest.

One tension: the Rules also require personal data, associated traffic data and processing logs to be retained for at least one year for specified purposes such as breach investigation. Reconcile erasure and minimum retention store by store with your privacy team.

Step 4: set retention limits on prompt logs

Prompt logs are often the fastest-growing and least governed personal data store. Split them into two streams:

  • Processing metadata: request and user IDs, timestamps, model, tool calls, retrieved document IDs, policy decisions. Keep for audit and minimum-retention needs.
  • Content: raw prompt and response text. Redact before writing, keep only as long as debugging and evaluation need, restrict access, expire automatically.

Configure AI observability tools to mask fields by default, and check what your model provider retains on its side.

Step 5: minimise and redact before the model sees it

Consent is limited to data necessary for the purpose, which pushes you to send models less:

  • Detect PII, including Indian identifiers such as Aadhaar, PAN, mobile and account numbers, in inputs and retrieved context before the model call.
  • Swap values for reversible tokens (<ACCOUNT_1>) held in a vault; re-insert only in the final response, and only if the user is entitled to see them.
  • Give tools narrow outputs: a balance tool returns the balance, not the customer profile.
  • Redact before indexing when the identifier adds no retrieval value.

These controls overlap heavily with AI security for enterprise LLM apps: what a manipulated model cannot see, it cannot leak.

Want to build these patterns hands-on? Cloudsoft's AI, GenAI and Agentic AI course covers RAG pipelines, agents and guardrails, in Ameerpet or live online.

Step 6: treat model providers as processors

Every model API, embedding service, vector database SaaS and observability vendor that touches personal data is a processor and needs a valid contract. Give legal the facts:

  • Does the provider train on your inputs or outputs, and is the answer in the contract?
  • What does it log, for how long and why (for example, abuse monitoring)? Is reduced retention available?
  • Which regions and sub-processors handle the data?
  • Can it notify you of breaches and support deletion fast enough for your own 72-hour report to the Board?
  • Which safeguards does it commit to? The Rules expect safeguard obligations in processor contracts.

Step 7: plan for data residency

The Act takes a "negative list" approach: transfers abroad are permitted except to countries the government restricts by notification, and Rule 15 adds that fiduciaries must meet any government requirements on making such data available to foreign states. Stricter sectoral rules still apply; RBI's directions on storing payment system data in India are the obvious fintech example. For SDFs, the government can also specify personal data that must not leave India.

So: prefer Indian cloud regions for model endpoints and vector stores where available, record region in the data map, avoid "global" endpoints by default, keep a self-hosted model option for the most sensitive workloads, and make region a configuration switch rather than a rewrite.

Step 8: breach response for AI systems

AI adds new breach paths: a prompt injection that reveals another customer's data, a broken retrieval filter, an agent emailing the wrong person, a leaked trace export. Prepare:

  • Log retrieval and tool calls with the requesting user and the principal_id of returned data, so you can scope exposure in hours.
  • Add AI scenarios to the incident runbook, with kill switches per feature, tool and index.
  • Pre-draft notices to individuals and the Board.
  • Run a tabletop on "the chatbot showed customer A's loan details to customer B"; missing logs surface quickly.

For deeper detection and response skills, the Cyber Security course is the complementary path.

Step 9: children's data

If an edtech tutor or family app may be used by anyone under 18, you need verifiable parental consent, with due diligence that the parent is an identifiable adult (the Rules allow reliable identity and age details or a virtual token from an authorised entity). No tracking, behavioural monitoring or targeted advertising directed at children, subject to exemptions listed in the Rules. In AI products this usually means age-gating, turning off long-term memory and profiling for child accounts unless clearly permitted, and excluding child conversations from training and evaluation data.

DPDP concept to engineering control: an illustrative map

Illustrative only: your legal team decides what your system actually requires.

DPDP conceptEngineering control in an AI system
Notice and consentVersioned notice service; consent record per principal; consent check at ingestion and retrieval
Purpose limitationMandatory purpose metadata; retrieval filters enforced in the data layer
Legitimate useslegal_basis per data flow, reviewed by legal
Data minimisationPII detection and tokenisation before model calls and indexing; narrow tool outputs
Access, correction, erasureQuery and delete by principal_id across sources, vectors, caches, memory, logs and processors
Grievance redressalTicket workflow tracked well inside the 90-day limit
Security safeguardsEncryption, masking, access control, monitoring, backups, prompt-injection defences
Breach intimationRetrieval and tool-call audit logs; AI incident runbook; 72-hour reporting drill
Retentionexpires_at on every record; split metadata and content logs; purge jobs
Children's dataAge gating; parental verification; no profiling or long-term memory by default
Processor contractsVendor register with training, retention, region and deletion terms
Cross-border transferRegion attribute per store; India-region endpoints; config-driven routing
SDF dutiesAudit-ready system documentation, evaluation results and risk tests

An illustrative Indian fintech example

Consider a Bengaluru lending and payments app that wants an AI assistant to answer "why was my EMI higher this month?" and help agents summarise disputes. A careful design:

  • Data map: chat text, loan records via tools, dispute documents in a RAG index; stores include chat logs, traces, vectors, a semantic cache and case notes.
  • Purpose: servicing the customer's loan is documented with legal; reusing chats to train a cross-sell model is not assumed and needs its own analysis, likely separate consent.
  • Minimisation: a get_emi_breakdown tool returns only needed fields; account and PAN numbers are tokenised before any model call.
  • Residency: endpoints pinned to an Indian region; payment data stays under the RBI storage directions compliance already follows.
  • Erasure: account closure plus expiry of the retention period triggers the fan-out job; documents under regulatory hold are excluded from retrieval.
  • Breach readiness: every retrieval logs whose documents went to which user.

The banking AI assistant project walks through a similar build end to end. Delivering systems like this alongside a client's compliance and security teams is typical Forward Deployed Engineer work, which is why Cloudsoft's FDE PRO program includes a Secure Banking AI Assistant project.

Frequently asked questions

Does the DPDP Act apply to AI and LLM applications?

Yes. It covers digital personal data processed in India, and processing abroad connected with offering goods or services to people in India. There is no AI exemption, so prompts, logs, embeddings, agent memory and training data containing personal data are covered.

When do the DPDP Rules come into force?

The Rules were notified in November 2025. Board provisions applied immediately, Consent Manager registration follows after 12 months, and most business obligations apply after 18 months, in May 2027. MeitY has consulted on shortening some timelines, so check the current official position.

Can I send personal data to a foreign LLM API under the DPDP Act?

The Act permits transfers abroad unless the government restricts a country by notification, but sectoral rules such as RBI directions can be stricter and Significant Data Fiduciaries may face extra restrictions. You also need a processor contract and should minimise what you send. Confirm with your legal team.

Can I use customer chat logs to fine-tune a model?

Only with a lawful basis covering that purpose; support consent does not automatically extend to training. Because erasing data from weights is impractical, many teams fine-tune on de-identified or synthetic data instead.

How do I handle an erasure request in a RAG system?

Tag every chunk and vector with the Data Principal identifier, then run one job that deletes source records, chunks, vectors, cached responses, agent memory and log content, and asks processors to delete their copies. Handle backups through short expiry and a re-delete step on restore.

Is publicly available personal data outside the DPDP Act?

The Act excludes personal data made public by the Data Principal or by someone legally obliged to publish it. The exclusion is narrow and fact-specific, so do not assume scraped web data is exempt; take legal advice before training on it.

What are the penalties under the DPDP Act?

The Schedule to the Act sets maximum penalties per breach, for example up to 250 crore rupees for failing to take reasonable security safeguards and up to 200 crore rupees for breach-notification or children's-data failures. The Data Protection Board decides actual penalties.

Is this article legal advice?

No. It is an engineering guide that simplifies the law. Work with your legal and privacy team and check the official Act and Rules from MeitY before making compliance decisions.

Privacy-aware AI engineering is now part of shipping AI in India, not an afterthought. To learn to build RAG, agents and guardrails with these controls designed in, explore Cloudsoft's GenAI and Agentic AI training in Hyderabad, in Ameerpet or live online. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us