Why does your data estate look the way it does?
Three clocks ran at different speeds for thirty years. Platform technology evolved on a vendor cycle, compliance on a legislative cycle that never tracked it, and business systems on a quarterly one. Nobody was asked to reconcile them.
The platform technology changed first. Thirty years took organisations from relational databases and reporting warehouses, through the data lake, into lakehouse and streaming architectures. Each step solved a real limitation of the one before it, and each step also left the previous layer running, because the regulatory reports built on it still had to be produced.
The compliance regime changed on its own schedule and never tracked the technology. The Privacy Act 1988 predates the data lake by more than two decades. APRA's CPS 234 arrived in 2019. The Security of Critical Infrastructure Act 2018 was amended to cover data storage systems in December 2024. Each obligation was written against the technology of its moment, and each one is still in force.
Business systems were bought to solve specific problems for specific business leaders on quarterly timeframes. A finance system, a case management system, a student management system, a claims platform. Interoperability was not the criterion the purchase was judged against, and a common data fabric was not on anyone's evaluation matrix. Until recently there was no consumer that required the three clocks to be reconciled. Reporting could live with fragmentation. An AI agent cannot.
What actually breaks when you point AI at an ungoverned estate?
Not that AI stops working. It keeps working, assembling confident answers from whatever it can reach, and nobody downstream can tell a grounded answer from an invented one.
In a fragmented estate the model reaches the file share as readily as the system of record. When the real answer is not findable in the chaos, it fills the gap. The output arrives with the same tone and formatting as a correct answer. Cursor's support bot is the clearest public example: in April 2025 users were logged out when they moved between devices, and the bot told them a new policy limited each subscription to one device. No such policy existed. The cause was a session management bug, so there was no record to retrieve, and users cancelled subscriptions over a rule the company had never made.
Recent Australian research found 42 per cent of organisations are deploying AI agents faster than they can standardise and govern them, and 43 per cent have no human-in-the-loop process for agentic workflows at all (Insight Enterprises, Assistance to Autonomous, June 2026, 318 Australian decision-makers). An ungoverned estate produces unreliable answers, and an ungoverned process means nobody catches them.
The regulator has been explicit. In its October 2024 guidance on commercially available AI products, the OAIC noted that generative AI is probabilistic and does not understand the data it handles, and that the APP 10 obligation to take reasonable steps to ensure personal information is accurate must be weighed against the increased risk in an AI context. Accuracy is not a quality metric in that framing. It is a privacy obligation.
What is a modern data platform, actually?
Six capabilities working together: ingestion, storage, transformation, governance, security, and serving and consumption. Naming them is what lets you tell an architecture apart from a licence.
Ingestion covers getting data in, broken into sub-capabilities for different data types, because a batch extract behaves nothing like an event stream and neither behaves like an unstructured document. Storage covers where data lands and how long it stays, which is where retention limits and immutability are decided. Transformation turns raw inputs into something usable, consistent and joined, and fixes the business meaning of a field.
Governance covers quality, labelling, classification, lineage and, most importantly, ownership: who is accountable for this data being correct. It is the capability most often written into a policy and least often assigned to a named person. Security covers monitoring, access control and permissions applied to the data rather than only to the systems holding it, because an agent reaches content, not applications.
Serving and consumption gets the right data into the right systems, tools and agents in a form they can act on. Most reference architectures draw these six as a stack of equals, left to right, with serving as the last box. The drawing is the problem. It implies a design and build order, and organisations follow it.
Why does centralising your data first fail?
Centralising is foundational but it is not the design. It tells you where your data is. It does not tell you whether an agent should be permitted to use it to decide something.
Two things get missed. The first is interoperability, which is not a state you reach but a constantly moving picture. Every new business initiative brings a new system, each arriving with its own model of entities you already hold, so a centralisation exercise scoped against the systems you had eighteen months ago will be incomplete on the day it lands.
The second decides whether the programme delivers value: the end goal. What systems does the centralised model serve, and how will those systems and agents consume the data. Consider a platform supporting grants or payments. The decision at the serving layer is whether an entity is eligible and should be paid. Work backwards and the requirements become concrete. The entity record has to be current and correct. The eligibility measures have to be defined as fields, not inferred from an attached document. Those records have to be immutable, because the determination has to be defensible after the fact.
Every one of those requirements reaches back through the stack. Immutability is a storage and governance decision. Defined eligibility measures are a transformation and modelling decision. Currency of the entity record is an ingestion decision. None can be retrofitted cheaply, and none would have been identified by an exercise that started at ingestion and worked forwards. In our experience this is where programmes stall, because nobody has written down who consumes the data and what they decide.
What do Australian regulations demand of the data serving layer?
Four regimes attach to four different things: the decision, the information asset, the information itself, and the system. None of them match the way an agent reads across an estate.
APP 1.7 to 1.9 of the Privacy Act, inserted by Part 15 of Schedule 1 to the Privacy and Other Legislation Amendment Act 2024, commence on 10 December 2026 and attach to the automated decision. An APP entity that has arranged for a computer program to make, or substantially and directly support, a decision significantly affecting a person's rights or interests must set out specified information in its privacy policy. It is a transparency obligation rather than a prohibition, and that is precisely why it is a platform problem: you cannot describe what personal information feeds a decision unless your serving layer can tell you.
APRA's CPS 234 paragraph 20 attaches to the information asset, requiring classification by criticality and sensitivity, including assets managed by related and third parties. PSPF Release 2025 Policy 8 attaches to the information itself. Since 20 December 2024, section 9(7) of the SOCI Act attaches to the system, deeming data storage systems holding business critical data to form part of the critical infrastructure asset where the responsible entity owns or operates them and compromise could affect the asset.
That last point has a counterintuitive consequence. Build a single centralised store holding business critical data and you may have brought a new system inside the Part 2A risk management program, a regime that previously only touched your operational environment. That is not a reason to avoid centralising. It is a reason to know what you are doing before you do it. Underneath the fragmentation, obligations remain outcome-only: APP 11.1 requires reasonable steps but does not say what a reasonable step looks like for an index built over records you hold for a specific purpose.
How to sequence the work before your next AI deployment
- Name the consumers and the decisionsBefore any platform work, write down which systems, tools and agents will consume data from the platform, and what decision each one makes or supports. Be specific. "Customer service" is not a consumer. "The agent that tells a member whether their insurance cover has lapsed" is. If a consumer cannot be named, it does not belong in the first release.
- Sort those decisions by consequenceDecisions that significantly affect a person's rights or interests carry the APP 1.7 obligation from 10 December 2026. Eligibility, payment, entitlement, assessment and access to a service almost always fall in that group, and they demand the most upstream guarantee.
- Trace each high-consequence decision backwardsFor the top two or three, work back through serving, security, governance, transformation, storage and ingestion, and write down what each layer must guarantee for the decision to be defensible. This produces an evidenced scope rather than an aspirational one.
- Assign ownership before you assign technologyEvery data set in the traced path needs a named accountable owner in the business, not in the technology function. This is the step that gets deferred, and it determines whether the governance layer is real or decorative.
- Treat interoperability as a standing capabilityNew systems will keep arriving. Decide now how a new source is onboarded, how its entities map to the ones you already hold, and who approves that mapping. Without this the platform is accurate on launch day and drifting by the end of the quarter.
- Build the serving contract, then build backwardsSpecify what the serving layer will hand each named consumer, including the classification, the provenance and the permitted purpose travelling with the data. Only then commission the ingestion, storage and transformation work required to satisfy it.

