Your Change Board Meets Monthly. The Exploit Arrived Tuesday Afternoon.

A proof-of-concept drops on Friday. Mass exploitation begins Monday. And somewhere in between sits your change advisory board, meeting on the second Wednesday of the month and requiring three signatures for any production change.

A proof-of-concept lands on a Friday. By Monday there’s a mass-scanning campaign hitting every exposed instance on the internet. Somewhere in that timeline sits your change advisory board, which meets on the second Wednesday of the month and requires three sign-offs for anything touching production.

That’s the whole story, really. The Azure Blog’s recent argument — that the window between disclosure and active exploitation has collapsed from weeks to hours or days — isn’t news to anyone who’s watched a mass-exploitation event unfold in real time. What’s new is that Microsoft is finally framing it as an architectural problem rather than a nagging operational one. The implication is uncomfortable: a manual approval process that takes longer than the attacker’s dwell time isn’t caution. It’s exposure dressed up as diligence.

Let me bring in the voice I always hear in my head when I write something like this — the skeptical platform lead who has been burned by automation before and isn’t thrilled about being told to trust it more.

“So the pitch is: let the robots patch prod. We’ve met.”

That’s not the pitch. The pitch is narrower and more defensible: move the decision about whether a patch class gets applied out of a ticket queue and into policy, so that the default state of your estate is “current” and exceptions are the thing you file paperwork for — not updates.

Right now most shops have it backwards. Patching is opt-in per cycle, gated on human review, and the exception — “we skipped this ring because of the holiday freeze” — is invisible until an auditor or an incident finds it. Flip it. Make patched the enforced baseline through Azure Policy, and make not patching the deliberate, logged, time-boxed exception with a named owner.

“Zero-touch” is a marketing word. What you actually want is low-touch with fast rings: automation applies, telemetry watches, humans intervene by exception. The human is still there. They’re just not standing in the doorway blocking every single update.

“Fine. But ‘the estate’ isn’t one thing. You’ve got VMs, App Service, and whatever Arc is pretending to manage.”

Correct, and the blast radius differs for each — which is exactly why you can’t have one policy and call it done. The services below change their surface area often enough that you should treat the specifics as representative patterns and confirm the current behaviour, SKUs and API versions against the product docs for your regions before you wire anything up. The shape of the approach is stable even when the plumbing moves.

Azure VMs are the straightforward case. Azure Update Manager handles assessment and scheduled deployment for Windows and Linux: you enable periodic assessment so the platform actually knows what’s missing instead of guessing, set a maintenance configuration, and attach VMs by scope so the platform orchestrates the deployment. The exact agent and scheduling model has shifted over time — verify what applies to your fleet rather than trusting an old runbook.

run.shbash — zsh
# Set a VM to platform-managed patching + assessment.
# IMPORTANT: 'az rest' has NO dry-run. There is no --what-if here.
# So scope tightly: run this against ONE non-prod VM first, confirm the
# result, and look up the CURRENT api-version for Microsoft.Compute
# virtualMachines in the docs before you paste a version string in.
az rest --method patch \
  --url "/subscriptions/<sub-id>/resourceGroups/<rg>/providers/Microsoft.Compute/virtualMachines/<vm>?api-version=<current-api-version>" \
  --body '{
    "properties": {
      "osProfile": {
        "windowsConfiguration": {
          "patchSettings": {
            "patchMode": "AutomaticByPlatform",
            "assessmentMode": "AutomaticByPlatform"
          }
        }
      }
    }
  }'
# Then bind the machines to a maintenance configuration (the schedule ring)
# via the maintenance assignment commands — test on the non-prod ring first.

App Service is where people get complacent. The OS and runtime underneath are Microsoft’s problem — you don’t patch the host. But the stack version you pinned eighteen months ago is your problem, and it doesn’t move unless you move it. A deprecated runtime is an unpatched runtime wearing a nicer hat. Enforce supported stacks with policy and stop treating PaaS as maintenance-free.

Arc-connected machines are the real test of whether you mean it. On-prem boxes and other-cloud VMs projected into Azure can be brought under the same update-management and policy surface through the Connected Machine agent — but confirm which capabilities are generally available for Arc in your environment, because the hybrid story lags the native one and the gaps are where you get hurt. The catch is operational: your policy scopes, your identity model, and your network egress all have to actually reach those machines. Arc is where “we have a patching strategy” quietly becomes “we have a patching strategy for the things that were easy.” Scope your policy at the management group and include the Microsoft.HybridCompute/machines type explicitly, or your hybrid fleet silently falls out of the baseline.

“And when an auto-applied patch takes down a production workload at 2am, who’s holding that pager?”

You are. Same as you are now. The difference is you get to choose whether the failure mode is “a canary ring broke and self-contained” or “the whole estate took the bad update simultaneously because we have one big-bang schedule.”

This is the part the vendor blog soft-pedals, so I’ll say it plainly: faster patching does break things. The mitigation is not slower patching. It’s staged patching — rings — so breakage is detected on a population you can afford to lose before it reaches the population you can’t.

Model it as maintenance configurations with staggered schedules:

  • Ring 0 — canary: non-prod and a thin slice of stateless prod. Patches land first, within hours of availability.
  • Ring 1 — broad prod: 24–48 hours behind the canary, gated on health signals, not on a human clicking approve.
  • Ring 2 — crown jewels: the databases and stateful tier, behind the broad ring, with an actual rollback plan tested more than once.

The health gate between rings is where your automation earns trust. Wire Defender for Cloud and your monitoring so that if Ring 0 shows regressions, Ring 1 doesn’t fire. That’s an orchestration problem, not an approval-meeting problem — and orchestration runs at 2am without needing coffee.

“Where does Defender for Cloud actually fit, or is it just the thing that generates recommendations nobody reads?”

Fair shot, and partly deserved. Defender for Cloud’s recommendation list is where good intentions go to accumulate. But two of its functions matter here and are worth turning on deliberately rather than by accident — check the current plan structure, because Microsoft renames and re-packages these more often than anyone would like.

First, its integrated vulnerability assessment is meant to tell you what’s exposed right now against what’s missing — the correlation between “this CVE is being exploited” and “these machines of mine are vulnerable to it.” That prioritisation is the whole game when the window is hours. Blanket “patch everything” is noise; “patch these internet-facing machines carrying the actively-exploited CVE, first” is signal. Treat the exact scoring and exposure signals as a capability to validate in your tenant, not a promise to take on faith.

Second, the agent/extension provisioning settings are what get the assessment agents onto new resources without someone remembering to install them. The failure mode you’re designing out is the brand-new VM that nobody onboarded — the one that’s invisible to your baseline and therefore permanently unpatched. Configure provisioning at the plan/subscription level and then verify that Arc machines inherit what you expect rather than assuming they do.

Making “patched” the default, in policy

The control-plane shift is concrete, not a metaphor. You use Azure Policy to audit and enforce the patch posture. Before you hand-roll guest configuration, check for built-in definitions and initiatives in your tenant that already cover machine update assessment and pending-update auditing — assign those scoped at the management group so new subscriptions inherit the baseline on day one.

run.shbash — zsh
# Find candidate built-in initiatives first — don't assume the name or ID.
az policy set-definition list --query "[?contains(displayName, 'update')].{name:displayName, id:id}" -o table
​
# Assign with identity so any DeployIfNotExists effects can remediate.
# Default to audit, promote to enforcement after you've read the compliance report.
az policy assignment create \
  --name "enforce-update-assessment" \
  --scope "/providers/Microsoft.Management/managementGroups/<mg-id>" \
  --policy-set-definition "<initiative-definition-id>" \
  --mi-system-assigned --location <region> \
  --enforcement-mode DoNotEnforce

Start in DoNotEnforce (audit). Read the compliance report. It will be worse than you think — it always is — and that gap is your actual risk register, not the one in the spreadsheet. Then promote to enforcement ring by ring.

The risk-based exception is the governance you keep. An exemption with an expiry date and a justification is a managed decision. Grant it at the resource-group or resource scope, and make it expire so “temporary freeze” can’t quietly become permanent drift:

audit.ps1PowerShell
# Time-boxed exemption — governance that expires instead of rotting.
New-AzPolicyExemption `
  -Name "freeze-legacy-erp-q4" `
  -PolicyAssignment (Get-AzPolicyAssignment -Name "enforce-update-assessment") `
  -ExemptionCategory Waiver `
  -Scope "/subscriptions/<sub-id>/resourceGroups/rg-legacy-erp" `
  -ExpiresOn (Get-Date).AddDays(30) `
  -Description "ERP vendor validating patch set; owner: platform-team; review 2026-11-03"

What the blog won’t tell you

Two things. First, none of this removes the need for change control — it relocates it. You’re moving governance upstream, into policy definitions and ring design and exemption reviews, where it runs continuously instead of in a monthly meeting. If your organisation reads “automated patching” as “no governance,” it will build neither the rings nor the exemption discipline, and the first bad patch will be used as evidence that automation was the problem. It wasn’t. The missing canary ring was.

Second, the hard part was never the technology. Update orchestration, policy enforcement, Defender assessment — these have been shippable for a while. The hard part is telling a change board that its approval step is now slower than the threat and therefore has to become an exception process. That’s a political conversation, and the collapsing window is the only leverage you’ll get to have it. Use it before an incident has the conversation for you.

The attacker automated exploitation years ago. The question on the table is whether your defence is still waiting for a human to click approve.