Why doesn't a fixed AI budget work?

Because much of your AI spend is priced on consumption, not seats. A flat annual budget assumes cost tracks headcount, but token, API and agent costs move with what each team actually does.

You are a commercial or technology leader in a financial services or superannuation organisation. Over the past six to twelve months you have rolled out generative AI tools across underwriting, member services, and engineering, and adoption has gone well by every measure your business cared about at the time. Then the board asks why AI spend has tripled this quarter, and who authorised it. The honest answer is that no one did, not in the way the question assumes. Usage grew, not procurement, and you are the one standing in front of the Risk Committee without the data to explain how. This is not a failure of judgement. It is the predictable result of governing a usage-based cost with a seat-based mental model, and it is happening across financial services and superannuation right now.

Most organisations set an AI budget the way they would set any software licensing budget: a headcount, a per-seat rate, an annual total. That model works when cost tracks seats. It does not work when cost tracks usage, and increasingly it does not.

Some AI providers are still sold on a subscription basis, where a seat costs a fixed amount regardless of how hard it is used. Others, and a growing share of the market, are priced on consumption: tokens processed, API calls made, agent runs completed. Where consumption-based pricing applies, the monthly bill has no natural ceiling tied to a headcount. It moves with what people actually do, and what people actually do varies enormously by team and by moment. An engineering function running a heavy sprint has fundamentally different usage from the same team in a quiet BAU month. Finance has different usage again, and a different usage curve, from underwriting. A single flat budget, set once a year, was never going to fit all of that at once.

Comparison table of seat-based and usage-based AI pricing. Seat pricing is fixed per seat, has low sensitivity to usage spikes, and fits variable team demand poorly because heavy and light users cost the same. Usage pricing scales with tokens, calls and agent runs, is highly sensitive to spikes, and fits variable demand accurately only if usage is attributed by team. Most organisations run both models at once.
A budget built for predictable seats does not fit consumption that was never predictable to begin with.

One in three Australian businesses exceeded their AI budget in the last financial year, and 32 per cent subsequently paused, cancelled, or scaled back AI deployments because the spend could no longer be justified (Elastic AI cost research, 2026). That is not a story about profligate teams. It is a story about a budgeting model that assumed predictability the underlying tools were never built to provide.

Why can't we just cap AI seats or negotiate a better rate?

Because this is an observability and attribution problem, not a procurement one. Caps and better rates don't tell you where spend is coming from, and you can't negotiate your way out of not knowing.

The instinct, when a board asks about AI cost, is to reach for the levers that work on every other piece of enterprise software: renegotiate the contract, cap the number of licences, push back on the vendor. Those levers matter, but they are solving the wrong problem. You cannot negotiate your way out of not knowing where the spend is coming from.

Australian research bears this out. Ninety per cent of executives report confidence in their organisation's visibility into AI tool usage. At the same time, Australia ranks second globally for unauthorised AI tool use, with 60 per cent of workers using AI tools their organisation has not sanctioned, trailing only the United States. That gap between executive confidence and actual visibility is not evidence of anyone failing to do their job. It is evidence that most organisations have not yet built the layer that would let that confidence be tested.

A gateway helps close part of that gap, and it should be the primary source of usage data in any AI FinOps model, but it is not the whole answer. A gateway only sees traffic that is routed through it. Embedded AI features inside existing SaaS tools, and any use of AI tools outside sanctioned channels, sit outside that view. In practice, most organisations will land on a coverage figure, not a guarantee: an assessment along the lines of, for illustration, "we estimate most of our AI usage is visible through the gateway, with host-based or endpoint controls needed to close the remaining gap." The actual figure will differ for every organisation and has to be tested, not assumed. What matters is stating a position at all, and treating it as a starting point for closing the gap rather than as evidence the problem is solved.

What does an AI FinOps operating model look like?

Five parts working as one system: a metering layer, a tagging and attribution scheme, a stated coverage position, budget envelopes with tolerance thresholds, and alerting tied to those thresholds, reviewed every quarter.

An operating model that can survive a board question needs five things working together. None of them is a single product, and no single tool available today closes this gap in full. What matters is that they operate as one system, not five disconnected initiatives.

1. A metering layer, most practically an AI gateway, sitting between users and AI providers wherever the environment allows traffic to be routed through it. This is where the majority of usage data should originate, and it is the foundation the rest of the model is built on.

2. A tagging and attribution scheme that assigns every request, where technically possible, to a user, a team, a department, and where relevant a project. Without this layer, a metering tool tells you total spend, not where it came from, which is the exact gap that put a commercial or technology leader in front of the board with no answer.

3. A stated coverage position, treated as a risk assessment and reviewed on a cadence rather than assumed. Knowing, for example, that most visibility comes through the gateway, with a defined remainder addressed by other controls, is a defensible position, provided the figure is the organisation's own, tested number rather than a borrowed rule of thumb. Assuming full visibility without having tested it is not.

4. Budget envelopes with tolerance thresholds, replacing the fixed annual number. Finance sets an upper band, not a static figure, and usage is free to move within it. Practitioners working through this problem elsewhere have converged on the same shape: seat-based caps are being abandoned in favour of attribution by role and task, because a principal engineer running large workloads and a Finance analyst running occasional queries were never going to sit sensibly under the same limit.

5. Alerting tied to thresholds, not to the monthly invoice. By the time a bill arrives, the quarter that drove it is already over. An alert triggered when a team crosses a defined proportion of its envelope gives the organisation the chance to act inside the period, not after it.

None of this is static once built. Usage genuinely changes as projects start and finish, and the operating model has to change with it. That means a standing review cadence, typically quarterly, where actual usage by user, team, and project is checked against the envelope, the coverage position is reassessed, and budgets are re-aligned to where the organisation actually is, not where it was set to be twelve months earlier. The practitioner consensus treats AI budgeting as a three-way responsibility: finance sets the tolerance, engineering turns that tolerance into enforceable technical controls, and a function bridging the two, whether inside FinOps or as a standalone capability, translates the resulting data into decisions either side can act on. A token quota set by only one of those groups tends not to hold.

Flow diagram of the AI FinOps operating model as a closed loop. Sanctioned AI tools, embedded SaaS AI features and direct or unsanctioned API use feed a metering or gateway layer with a stated, tested coverage position. Usage is then tagged to user, team, department and project, which drives a budget envelope set as a tolerance band and threshold alerting triggered inside the period. Both feed a quarterly review that re-aligns envelopes and reassesses coverage, looping back to the gateway.
This is a cycle of measurement, review and re-alignment, not a control you switch on once.

What do CPS 230 and SPS 515 mean for AI spend?

An AI tool embedded in critical operations may fall within CPS 230's material service arrangement reporting, and for super trustees, SPS 515 requires AI expenditure to be justified and monitored against member outcomes.

For an APRA-regulated entity, an AI tool that has become part of daily operations across underwriting or member services may not sit outside CPS 230's scope simply because it was procured as commodity software. CPS 230 requires that an entity not rely on a service provider unless it can continue to meet its prudential obligations in full, and material arrangements, which explicitly include core technology, require senior management reporting commensurate with how heavily they are used. APRA has not, to date, issued specific guidance naming AI tools as material service arrangements, so this is Lumaris's reading of how the standard would likely apply, not settled regulatory guidance. Where an AI tool has become embedded in a critical operation, an entity would be prudent to at least test that reading against its own arrangement rather than assume it falls outside scope.

For a superannuation trustee, SPS 515 adds a more direct question. Expenditure decisions must be justified against the fund's strategic objectives and the outcomes sought for beneficiaries, and the standard requires the entity to define, in advance, how that expenditure will be monitored. It also requires an assessment of whether operating costs are adversely affecting members' financial interests. An AI rollout that triples in cost without an attribution model to explain it is not simply an internal budget overrun. It sits close to a question SPS 515 already requires the fund to be able to answer: whether the spend still serves members, was monitored the way it said it would be, and did not erode value on their behalf.

How does AI FinOps map to NIST AI RMF, Australia's Guidance for AI Adoption and ISO/IEC 42001?

Closely. Each framework expects accountability, ongoing measurement and a documented review cycle from AI governance. An AI FinOps model applies that same shape to spend, which checks the model rather than adding a separate compliance exercise.

The US National Institute of Standards and Technology's AI Risk Management Framework organises AI governance into four functions: govern, map, measure, and manage. The operating model above is, in effect, a FinOps-specific implementation of the measure and manage functions, applied to cost rather than to model risk narrowly defined. A metering layer and a stated coverage position are a measure function. Budget envelopes, threshold alerting, and the quarterly review cadence are a manage function, closing the loop the framework describes as continuous rather than one-off. The NIST framework does not speak in FinOps or budgeting terms directly; the mapping is Lumaris's reading of how a cost-attribution model fits its structure, not a claim the framework itself makes.

Closer to home, the National AI Centre's Guidance for AI Adoption, published in October 2025, sets out six essential practices for organisations deploying AI, building on but not replacing the earlier Voluntary AI Safety Standard. Two of those practices bear directly on this model. "Decide who is accountable" asks an organisation to name a person or function answerable for an AI system's outcomes, which is exactly the gap a board question about tripled spend exposes when no one owns the attribution data to answer it. "Measure and manage risks" asks for ongoing testing and monitoring rather than a point-in-time assessment, which is the same logic behind treating a coverage position as a tested, reviewed figure rather than a fixed assumption. The guidance also encourages organisations to maintain a register of the AI systems in use, a natural companion to the attribution layer: the register names what is in use, the attribution layer shows how it is being used and by whom.

ISO/IEC 42001, the international standard for AI management systems, follows the same plan-do-check-act structure as other management system standards such as ISO 27001. At that level, a documented, reviewed operating model of the kind described above is consistent with the discipline the standard expects of an AI management system, though certification against it would require a more detailed gap assessment than this briefing attempts.

Taken together, none of these frameworks tell an organisation to build an AI FinOps model as such. What they do is establish that accountability, ongoing measurement, and a documented review cycle are the expected shape of AI governance more broadly, and this cost model is that same shape applied to spend.

What does a defensible answer to the board look like?

A gateway with honestly stated partial coverage, tagged to team and project and reviewed quarterly against a tolerance band, is a materially stronger position than a flat annual budget and a seat count.

None of this requires waiting for a perfect system. The next time the board asks why spend moved, the answer won't require guessing. It will name what drove the increase, by team and by project, state what the organisation can currently see and what it cannot, and set out what changes next quarter as a result.

Action

What to do before your next board or Risk Committee meeting

  1. Route sanctioned AI traffic through a gatewayPut a metering layer between your users and AI providers wherever traffic can be routed, and make it the primary source of usage data.
  2. Tag every request to its ownerAttribute usage to user, team, department and, where relevant, project, so any movement in spend can be traced to who drove it.
  3. Write down your coverage positionEstimate how much of your AI usage the gateway actually sees, test that figure, and name the host-based or endpoint controls covering the rest.
  4. Swap the annual figure for tolerance bandsHave finance set an upper band for each team, and have engineering turn that band into enforceable technical controls.
  5. Alert on thresholds, not invoicesTrigger an alert when a team crosses a defined share of its envelope, so you can act inside the period rather than after the bill lands.
  6. Test your CPS 230 and SPS 515 positionCheck whether any AI tool embedded in a critical operation is a material service arrangement and, if you're a super trustee, define in advance how AI expenditure will be monitored against member outcomes.
  7. Book the quarterly reviewPut a standing quarterly review in the calendar with finance, engineering and whoever owns the attribution data, and re-align envelopes to actual usage each time.
Todd Noller
Principal Consultant · Lumaris Consulting

Todd Noller is Principal Consultant at Lumaris Consulting, an Australian-owned, vendor-neutral advisory firm specialising in AI, data, cyber security, cloud, and critical infrastructure. Todd is a delivery and operations leader who builds technology functions, governance frameworks, AI platforms, and the teams that run them. With over sixteen years across data, AI, and cyber delivery in regulated environments, Todd led the Complex Programs portfolio at a Defence Prime: a ~$40M book of cyber programmes spanning federal digital identity, enterprise integration, and security uplift. Prior to that role, he stood up an investment management firm's entire technology function from scratch across New York and Australia, launching multiple SaaS products to SEC, FINRA, APRA CPS 234, and GDPR compliance within two years. At the Department of Defence, he coordinated national-scale cyber incident responses and served as Business Product Owner for a real-time threat intelligence platform. Todd spent three years as a Palantir specialist, including as Country Manager for Palantir's Australian deployments.

View LinkedIn profile