AKS Now Defaults New Clusters to StandardV2 NAT Gateway—Check the Cost

Microsoft changed the default NAT Gateway SKU for new AKS clusters, and the new one bills per gateway-hour plus per gigabyte of data processed. If you're deploying clusters with managedNATGateway outbound type, the new…

Defaults are promises the platform makes on your behalf, and they change without asking. The latest one: new AKS clusters that use an AKS-managed NAT gateway for outbound now get a StandardV2 NAT gateway instead of the previous default. Microsoft shipped this to general availability (Azure update 574430), and on paper it’s an upgrade — more concurrent connections, zone resilience, more throughput headroom before SNAT ports run dry.

On the invoice, it’s a different conversation. StandardV2 bills per gateway-hour and per gigabyte of data processed. That second line is the one that bites a chatty cluster six months in, long after whoever ran az aks create has forgotten which SKU they got. Nobody gets paged for a NAT gateway SKU. They get paged for the intermittent connection timeouts when SNAT ports exhaust — and that’s the failure this change actually addresses.

Here’s what matters: this is narrow, it’s opt-in for anything that already exists, and it only touches one specific configuration. Work through these in order before your next cluster deploy.

1. Confirm the new default even reaches you

Four outbound types exist in AKS: loadBalancer (the historical default), managedNATGateway, userAssignedNATGateway, and userDefinedRouting. This change touches only clusters set to managedNATGateway. If you’re on loadBalancer — and most clusters are, because that’s what you get when you don’t specify — nothing here applies. If you’re on userAssignedNATGateway, you’re bringing your own gateway and you already own the SKU decision. Check before you worry:

run.shbash — zsh
az aks show -g <rg> -n <cluster> \
  --query "networkProfile.outboundType" -o tsv

If that returns anything other than managedNATGateway, skip the rest and go read something else. The blast radius is exactly the set of clusters that both use the AKS-managed VNet path and chose managed NAT gateway as the egress route. That’s a deliberate subset, not your whole fleet.

2. Find out what SKU you actually have right now

The NAT gateway doesn’t live in your cluster’s resource group. AKS provisions it in the node resource group — the one with the MC_ prefix it manages on your behalf. So you need two hops:

run.shbash — zsh
# 1. find the managed node resource group
NODE_RG=$(az aks show -g <rg> -n <cluster> \
  --query nodeResourceGroup -o tsv)
​
# 2. read the NAT gateway and its SKU in that RG
az network nat gateway list -g "$NODE_RG" \
  --query "[].{name:name, sku:sku.name, zones:zones}" -o table

Existing clusters are untouched. If yours was created before this rolled out, it keeps whatever SKU it had — the new default only decides for new clusters at creation time in supported regions. The update lists availability by region rather than enumerating every one; verify your target region there rather than assuming, because “supported regions” is doing real work in that sentence and the list grows over time.

3. Model the two-part bill before you celebrate the resilience

This is the step people skip, and it’s the one that matters. A NAT gateway’s data-processing charge is per gigabyte of traffic flowing through it — all of it, not just the expensive-looking stuff. A cluster that pulls large container images on every scale event, ships verbose logs to an external SIEM, or backs up to a non-private storage endpoint can push serious volume through that gateway. The per-hour charge is predictable and small. The per-GB charge is the one that compounds.

For a low-traffic internal cluster, the simpler older model can genuinely be cheaper. For a high-connection-count workload that was already flirting with SNAT exhaustion, StandardV2’s headroom is worth paying for — you’re trading a few dollars of data processing against 3am outbound failures. Decide which cluster you have before the finance team decides for you. Pull current egress volume from the gateway’s metrics rather than guessing.

4. Leave your existing clusters alone unless you have a reason

There is no forced migration. Existing clusters do not change SKU because a default moved. Resist the urge to “standardise the fleet” with a mass upgrade — a NAT gateway change is an egress-path change, and egress-path changes are how you discover that a downstream firewall allowlists your current public IPs. The only cluster worth touching today is one that’s actively hitting SNAT port limits or connection ceilings. Everything else can wait for its next planned rebuild, where the new default picks it up for free anyway.

5. Pin the choice in IaC so the default stops deciding for you

A changed default is a signal: you’re letting the platform make an architectural decision at deploy time. For anything production, make the choice explicit in code so a cluster recreate — or a copy-paste into a new region — doesn’t silently inherit whatever the default happens to be that quarter.

snippet.txtBICEP
resource aks 'Microsoft.ContainerService/managedClusters@2024-05-01' = {
  name: clusterName
  location: location
  properties: {
    networkProfile: {
      outboundType: 'managedNATGateway'
      natGatewayProfile: {
        managedOutboundIPProfile: {
          count: 2
        }
        idleTimeoutInMinutes: 4
      }
    }
  }
}

Pin outboundType, the managed outbound IP count, and the idle timeout explicitly. The IP count and idle timeout are what actually govern your available SNAT ports — more outbound IPs means more ports before exhaustion. Those knobs matter more to your outbound reliability than the SKU label does. Run your IaC with --what-if (Bicep) or terraform plan and read the diff before you apply; an in-place change to a NAT gateway profile can be more disruptive than it looks.

6. Survey the whole fleet in one query

Before any of the above means anything, you need to know your actual exposure across every subscription. Azure Resource Graph answers this in one read-only shot — no per-cluster looping:

run.shbash — zsh
az graph query -q "Resources
| where type =~ 'microsoft.containerservice/managedclusters'
| project name,
          rg = resourceGroup,
          sub = subscriptionId,
          outbound = properties.networkProfile.outboundType
| where outbound =~ 'managedNATGateway'"

That gives you the exact list of clusters this default change can reach. Everything not on that list is noise for this particular fire.

If you only do one thing: run the Resource Graph query. Knowing which handful of clusters use managedNATGateway turns a vague “did the default change affect us” into a short, named list you can actually reason about. A better NAT gateway you didn’t ask for is still a bill you didn’t budget for — and the first rule of changed defaults is that you find out what they cost before the platform tells you.