Why do adopt everywhere and wait for perfect both feel like the only options?

Neither instinct is a failure of judgement. Three real pressures collide, industry noise that adoption is urgent, technical debt that AI will not resolve on its own, and compliance obligations you will not jeopardise, and the way through is choosing deliberately rather than at either extreme.

Trying to do everything ignores the debt and the compliance obligations. Waiting for a perfect plan ignores that no plan removes all risk; it only delays learning whether AI actually helps your team. Deliberate sequencing, based on where AI genuinely adds value against your risk appetite rather than on hype or fear, beats both extremes.

What if AI is already running ahead of you?

The most common position we see is sanctioned and unsanctioned AI running side by side with no consistent governance, and retrofitting assurance is a known problem you do not need to halt the work to solve.

AI often arrives through developers experimenting with coding assistants, teams adopting copilots without a formal decision, and pilots that quietly became production dependencies. As agents go live faster than the policies written to cover them, which AI agent touched which change, on whose instruction, reviewed by whom, and whether it shipped with a known flaw, becomes unanswerable once work has spread across ad hoc tools.

That may be tolerable for one developer experimenting. It is not tolerable inside a regulated supply chain, a public-trust mandate, or critical infrastructure. You need visibility into what exists, a consistent pattern applied going forward, and an evidence trail that starts from today rather than waiting for a clean-sheet rebuild.

What does the evidence actually say about AI-generated code right now?

Caution is justified by data: only about 55% of AI generation tasks produce secure code, a known flaw introduced in roughly 45% of cases, even as syntax correctness now exceeds 95%.

Veracode's most recent testing against current flagship models found that security performance has not kept pace with syntax correctness. The technical debt picture reinforces it: technical debt has risen 30 to 41% following AI tool adoption, and legacy code refactoring has fallen 74% since 2023, as teams lean toward building new rather than tending what already exists.

None of this means AI has no place in your SDLC. It means where you start, or where you intervene first, matters more than the decision to start. Code that compiles, passes tests, and ships on schedule is not the same as code that is secure, and the gap between the two has been widening, not closing.

What is the Governed Delivery Loop, a pattern for governing AI in the SDLC?

One governed control point between your AI agents and your codebase, with four stages applied to every change: isolate, govern, evidence, and approve.

Most tools already in your stack solve one piece of this. AI assistants make developers faster, CI/CD platforms move code, GRC platforms track obligations, and AI governance point tools flag risk after the fact. None of them governs AI agents inside the delivery pipeline itself, with human sign-off structurally required and continuous evidence generated as a by-product of the work rather than reconstructed for an audit. That gap is where the Governed Delivery Loop sits.

Isolate runs every AI-executed task in its own contained workspace, often a git worktree or container, with access scoped to that task alone, so agents never collide or leak into each other. Govern checks every applicable policy and obligation automatically, every time rather than sampled, tags each commit with the identity that produced it so AI-authored changes get a stricter rule set, and widens the checks to catch AI-specific risks such as hallucinated or non-existent package dependencies.

Evidence records every check, pass or fail, the moment it happens, written by the control point itself rather than the agent self-reporting, so a complete entry answers who, what, when, where, and why. Approve keeps a human as the final, accountable sign-off before anything reaches customers, patients, members, or students. Because the loop is structural rather than procedural, governance travels with the work, one control point can evaluate several frameworks concurrently, and running on your own infrastructure keeps the evidence inside your existing security boundary.

Does every efficiency gain actually need to come from AI?

No. Linting, static analysis, and automated security scanning are mature, deterministic, and already understood by auditors, and they carry none of the security or evidentiary uncertainty of AI-generated code.

The useful question is not whether AI could do a task, but whether it adds measurable efficiency and cost savings over the automation you already have or could reach for. If a static analysis rule already solves the problem reliably, adding an AI agent adds cost and risk without adding value.

AI earns its place where judgement, pattern recognition, or language-heavy tasks genuinely outperform deterministic tooling, such as summarising ambiguous legacy code for review, drafting test scenarios from natural-language requirements, or triaging unstructured incident reports, not as a default answer to every inefficiency.

What does this look like across different regulated sectors?

The underlying obligations differ by sector, but the shape of the problem does not: whatever regime applies, you must be able to show, for any AI-assisted change, what happened, who checked it, and who signed off.

Healthcare carries TGA Software as a Medical Device evidence obligations where AI touches diagnosis, monitoring or treatment; financial services extends APRA CPS 234 and CPS 230 to AI-assisted delivery, not just production; higher education applies research and student data obligations; critical infrastructure extends SOCI Act positive security obligations to AI agents inside delivery pipelines; defence sits inside the Defence Industry Security Program; energy extends AESCSF maturity assessments; and state and territory government applies VPDSS or its equivalent. ISO/IEC 27001, Essential Eight and the Privacy Act are the cross-sector baseline. This is a starting orientation, not a compliance opinion, and your own regulatory advice should confirm what applies to your organisation.

Whichever regime applies, the evidence obligation is the same shape: continuous, specific to the change, and ready before it is requested, precisely what a structural control point is built to produce.

How do you actually choose your first, or next, AI-SDLC use case?

Inventory where time goes, score each candidate on risk, its proximity to certified or regulated scope, and on value, the time, cost, or rework it would save, then choose where measurable value and manageable risk intersect, not the lowest-risk or most impressive option.

A structured approach removes the guesswork, whether you are choosing where to start or bringing an already-running AI workflow under control. Score every candidate workflow, including anything AI is already doing informally, against both risk and value in the same sitting, so the choice is comparative from the start rather than one workflow assessed in isolation.

The right first use case sits where genuine value and manageable risk intersect, not the lowest-risk corner by default and not the one that would look best in a board presentation. Define your evidence approach before you start, then run the pilot, or the retrofit, as a genuine test with a defined success metric before deciding what is next.

Action

Six steps to choosing your first, or next, AI-SDLC use case

  1. Inventory your SDLC workflows, including AI already in informal useList where time and effort actually go across your delivery lifecycle: code review, test generation, documentation, refactoring, incident triage. Be honest about where debt, manual effort, and any unofficial AI use are concentrated.
  2. Score each workflow on riskAssess how close it sits to code that is safety-critical, or within your certified or regulated scope.
  3. Score each workflow on valueEstimate what genuine adoption would save in time, cost, or reduced rework compared to your current approach, including any existing automation.
  4. Choose a balanced candidate, not the safest or the most impressive oneThe right first use case sits where measurable value and manageable risk intersect, not the lowest-risk option by default and not the one that would look best in a board presentation.
  5. Define your evidence approach before you startDecide upfront how you will record what an AI agent touched, what was checked, and who signed off before it counted as done, extending your existing compliance obligations rather than sitting outside them.
  6. Run the pilot, or the retrofit, with a defined success metric, then decide what is nextTreat it as a genuine test, not a foregone conclusion. If it proves out, you have a defensible, evidenced case for where to extend next.
Todd Noller
Principal Consultant · Lumaris Consulting

Todd Noller is Principal Consultant at Lumaris, an Australian-owned, vendor-neutral advisory firm specialising in AI, data, cyber security, cloud, and critical infrastructure. A delivery and operations leader, he has over sixteen years across data, AI, and cyber delivery in regulated environments, including standing up an entire technology function to SEC, FINRA, APRA CPS 234, and GDPR compliance within two years.

View LinkedIn profile