Beyond Compliance: Why RBI's Data Governance Guidelines Require a New AI Architecture
Executive Summary
For twenty years, enterprise security protected where data lived. Encryption protected databases. Firewalls protected networks. Access policies protected applications. AI changed that model with agents breaking data security.
The RBI's draft data governance guidance, open for comment until August 17, sets expectations that most banks' current data lakes and AI pipelines cannot meet today. For every data element, an institution has to show where it came from, who owns it, how it may be used, and that no uncontrolled copy exists, and it has to hold those guarantees even when an AI agent or an outside model touches the data in real time.
The cost of getting it wrong is climbing: the average data breach in India now runs to a record ₹25.5 crore (IBM). Static, design-time controls like masking in test environments, tagging conventions, and standing database credentials were built for a world of fixed reports, and they cannot answer the question the RBI is really asking, which is to prove control at the moment of use. The institutions that treat this as an architecture decision, ensuring control and protection travel with the data itself, will satisfy the regulator and unblock their AI programs with the same move. The sections below map the draft's core requirements to that architecture, and to where Skyflow fits.
The Draft, in Plain Terms
On July 15, 2026, the Reserve Bank of India released its draft Guidance on Regulatory Expectations for Data Governance. Comments are open until August 17. On first read, it looks like standard governance fare: board oversight, committees, roles, a policy or two. Read it closely, and it's something more unusual. A regulator has written down, in plain language, why most data lake and AI architectures at Indian banks and NBFCs won't hold up.
The guidance draws on the Basel Committee's BCBS 239 principles for risk data aggregation. It applies to nearly every RBI-regulated entity, from commercial and cooperative banks to NBFCs, the All-India Financial Institutions, Asset Reconstruction Companies, and Credit Information Companies. Unlike guidance that stays at the policy level, this draft reaches into the data stack: single source of truth, metadata and lineage, classification, tokenization, third-party sharing, and cross-border processing.
Four Requirements a Typical Data Lake Does Not Meet
Most banks built their current data platforms, the Cloudera, Snowflake or Databricks instance feeding the analytics team, the lake behind the fraud model, the RAG pipeline behind the new AI assistant, for one purpose: get all the data into one place so anyone can use it. The RBI draft asks a harder question: Can you prove, element by element, where this data came from, who owns it, what it's allowed to be used for, and that no copy of it is drifting outside your control?
Four provisions bring that tension into focus (see below). Few banks have quietly cleared this bar, more than a decade after BCBS 239 introduced these principles, the Basel Committee's own reviews still find not one of them fully implemented across the banks it assessed, and the RBI is now extending that bar to nearly the entire Indian financial sector.
Here are the four provisions:
- Single Source of Truth (SSOT): The draft requires an SSOT for each data element, with "no parallel or competing SSOT sources"; downstream systems, models, and processes must derive from it. Most warehouse migrations, log exports, analytics sandboxes, and model-training extracts create the parallel copies it is written to eliminate.
- Metadata and lineage that travel with the data: Foundational metadata (owner, classification, purpose, permitted use) must be captured at origination and flow downstream "without loss of basic attributes created at source," updated on each transformation. Most ETL and AI pipelines flatten that context as data moves.
- Named security controls: The draft names the pattern directly: "data encryption, tokenization, and anonymisation," applied wherever data is processed, shared, or transformed. Gartner's Hype Cycle for Data Security Technologies, 2026 points the same way: digital tokenization preserves referential integrity, keeps data safe for cross-border transfer, and shrinks regulatory scope to token-handling systems.
- Provable need-to-know sharing: Data shared externally, which now includes foundation-model API calls, must stay "traceable to the designated SSOT," limited to "need to know," with no "unauthorised reuse, sharing, or duplication." A RAG pipeline that sends customer PII into an LLM prompt fails that test, more so when the model is hosted outside India, where the draft requires cross-border processing not to "impair the RE's ability to access, retrieve, and manage its data."
The underlying philosophy is old, BCBS 239 and the DPDP Act, 2023, now operationalized by the notified DPDP Rules, 2025, translated into requirements about how data actually moves through the modern AI data stack.
Why AI Turns Data Governance Into a Runtime Problem
The timing raises the stakes. Gartner had earlier predicted that 30% of AI projects would be abandoned after proof of concept. Poor data quality and weak risk controls were among the leading reasons stated. For a bank, those risk controls are now the RBI's to define, and this draft is where it starts.
AI systems do not access data the way a fixed report does. An agent decides at query time which records to pull, which fields to reason over, and what to place in a prompt. This loops thousand times per day continuously, unsupervised. A governance model built on static rules has nothing to say about a decision an agent made 3 seconds ago in response to a prompt no compliance officer will read.
That is the case for runtime data control for agentic AI. Design-time governance answers whether a pipeline was built correctly. Runtime data control answers what happened the last time a system touched customer data, the question supervision actually tests. A second tension: a data lake exists for democratized, self-serve access, while the draft's classification and traceability rules read as restriction. Reconciling them requires the data element to carry its own controls as it moves, not a separate layer that can be bypassed.
What “Compliant by Design” Actually Looks Like
This is the specific gap Skyflow was built to close. Skyflow's Runtime AI Data Control layer intercepts data flowing into datastores, models and agents. Skyflow becomes the authoritative source for regulated data protection the draft asks for. Read against the draft, the mapping is close to one to one.
End users of data like analysts, data scientists, and agents work on tokens that are safe to hold. Skyflow's runtime data controls sanitize (tokenize) sensitive data before it reaches an LLM and rehydrates responses only for permitted requesters (decided at request time). One policy governs data across the stack. Skyflow delivers a system-agnostic, model-agnostic solution that does not block or excludes data. it detects sensitive entities, sanitizes them, and unblocks major AI initiatives.
The Real Choice Ahead
The comment period closes August 17, but drafts like this rarely soften on their way to becoming binding directions. RBI has form here: its 2018 payment data localization directive moved from guidance to hard enforcement, with the RBI in 2021 barring American Express, and Diners Club from onboarding new customers for non-compliance. The exposure is concrete: under the DPDP Act, penalties run up to ₹250 crore per instance. Institutions that treat this as a documentation exercise, patching over the non-prod-only masking, log-and-hope access, and catalog-as-control habits that got them this far, will spend 2027 finding out, one supervisory inspection at a time, that those habits don't satisfy a requirement for enforceable SSOT and provable, runtime need-to-know access.
Institutions that treat the draft as a documentation exercise will meet its SSOT and runtime need-to-know requirements the hard way, one inspection at a time. Institutions that treat it as an architecture mandate will find that the same design that satisfies supervision also unblocks AI. The data finally becomes safe to use across the stack, including by agents.
That second path is the one that pays off. If you are mapping your data lake, warehouse, or AI roadmap against this draft, talk to Skyflow.
(This article reflects Skyflow's reading of the RBI's draft guidance and is not legal advice. Institutions should review the guidance in full and consult counsel.)
