Skip to content

Azure IPAM v4.0.0 Release - #377

Open
DCMattyG wants to merge 233 commits into
mainfrom
ipam-3.7.0
Open

Azure IPAM v4.0.0 Release #377
DCMattyG wants to merge 233 commits into
mainfrom
ipam-3.7.0

Conversation

@DCMattyG

@DCMattyG DCMattyG commented Mar 11, 2026

Copy link
Copy Markdown
Contributor

Azure IPAM v4.0.0

This is a major release that delivers significant framework upgrades, a data grid migration, authentication modernization, comprehensive documentation overhaul, revamped examples, and numerous bug fixes.


Major Framework & Dependency Upgrades

  • React 18 → 19: Full migration including removal of forwardRef wrappers (ref-as-prop), removal of PropTypes, explicit null for useRef() calls, and migration of LoadingButton to Button
  • MSAL v4 → v5 (@azure/msal-browser 4.x → 5.x, @azure/msal-react 3.x → 5.x): Removed obsolete config, consolidated event types, fixed silent token timeout recovery (timed_out error code), and prevented iframe fallback timeout loops
  • Inovua React Data Grid → AG Grid (ag-grid-community / ag-grid-react 36.x): Complete migration to AG Grid including centralized DataGrid component, custom styling, column state persistence, unified data loading overlays, and custom cell renderers (drill-down, info, progress)
  • Vite 7 → 8 (vite 7.x → 8.x, @vitejs/plugin-react 5.x → 6.x): Migrated to Vite 8 which replaces Rollup with Rolldown and esbuild with Oxc for bundling, transforms, and minification. Removed vite-plugin-eslint2 (redundant with editor-based linting and incompatible with Vite 8)
  • ESLint 10: Upgraded from ESLint 9 to 10 with modern React linting plugins (@eslint-react/eslint-plugin v5, eslint-plugin-react-hooks v7 — which now consolidates the React Compiler lint rules per the React Compiler 1.0 release), replacing eslint-plugin-react. Added dist/ ignore, fixed no-useless-assignment violations, and removed unused eslint-plugin-jest. Resolved all remaining lint warnings as part of the React 19 modernization — useContextuse, <Context.Provider><Context>, ref naming conventions, stable list keys, hoisted static styled components, migration of SnackbarUtils to notistack's standalone enqueueSnackbar, and moved "ref assigned during render" patterns into useEffect
  • MUI v7 → v9 (@mui/material 7.3.x → 9.0.x, @mui/icons-material 7.3.x → 9.0.x): Major two-version jump. Migrated deprecated component props to the unified slots/slotProps API (largely via @mui/codemod), moved deprecated system props into sx, replaced the removed Unstable_Grid2 with the new default Grid (using size={{ xs: N }}), renamed removed Outline (no "d") icon exports to their Outlined counterparts, and removed @mui/lab (no longer needed — LoadingButton's loading prop is now native to Button)
  • React Router 7 → 8 (react-router 7.x → 8.x): Major version upgrade. The UI uses declarative-mode routing (BrowserRouter / Routes / Route) with all imports already sourced from react-router, so no application code changes were required. v8 raises the minimum runtime to Node.js 22.22.0 (and React 19.2.7+)
  • Updated NPM packages across the board (Vite 7.3.x, React Router 7.13.x, MUI 7.3.x, etc.)

Engine & Backend

  • Azure Function Blueprints: Implemented Blueprint-based function naming for improved clarity
  • Python dependency cleanup: Removed msal, azure-common, azure-keyvault-secrets, and six; added azure-mgmt-resource-subscriptions to address Azure SDK module separation
  • Python linting: Added pyproject.toml with Ruff linter configuration (pycodestyle, pyflakes, isort). Resolved all lint violations across 18 engine files — bare except clauses, wildcard imports replaced with explicit names, unused imports/variables, invalid escape sequences, import sorting and grouping
  • Reservation logic hardening: Fixed CIDR overlap detection during auto-fulfillment, added validation for all in-block vNet prefix overlap checks, and made auto-fulfillment idempotent by deduping existing block vNet associations
  • Endpoint fix: Standalone NICs are now included in the Endpoint list (fixes Standalone NICs not listed in Endpoint node #371)
  • Network associations fix: Resolved improper handling of missing vNETs and vHUBs (fixes Error fetching available IP Block networks #350)
  • Next available network fix: nextAvailableVNet aborted its entire Block list search when the first Block could not satisfy the requested size under smallest_cidr, returning a 500 instead of evaluating the remaining Blocks. max() was called on an empty candidate list, and the resulting ValueError escaped the loop. The same missing guard in nextAvailableSubnet and in single-Block Reservation creation replaced their intended error messages with an unhandled exception (fixes nextAvailableVNet fails when using multiple blocks and smallest_cidr option true if no IP range is available in first block #384)
  • HTTP status codes corrected for client and capacity errors: Conditions the API explicitly anticipates were reported as 500, making them indistinguishable from genuine server faults, prompting pointless client and proxy retries, and counting against server-error metrics. Requests that cannot be satisfied by current allocation state now return 409, as does a JSON patch that conflicts with the stored document; a malformed JSON patch now returns 400. Both already matched the codes used by neighboring checks in the same handlers. Errors originating from Azure, Cosmos DB, or the site configuration are unchanged and still return 500
  • Development reload flag removed from production startup: uvicorn was started with --reload in both production init scripts. Under a read-only run-from-package mount its watcher tracked thousands of files that can never change, and its supervisor process stayed alive when the application crashed — so the container kept running while serving nothing, and health checks failed against an apparently healthy process. A startup failure now exits, so the platform restarts the application and the fault is visible. engine/Dockerfile.dev retains the flag, where hot reload is intended
  • TLS validation moved to the OS trust store: The JWKS fetch backing token validation used requests, which verifies against the CA bundle shipped inside certifi rather than the operating system trust store. WEBSITES_INCLUDE_CLOUD_CERTS populates the OS store, so sovereign cloud roots were present but invisible to that code path — secret cloud (IL6) deployments rejected every token with CERTIFICATE_VERIFY_FAILED while passing in every commercial cloud, where certifi covers the endpoints. Every other outbound call already used aiohttp, which reads the OS store; the JWKS fetch was the lone exception and now matches. The metrics heartbeat moved off requests as well, and a Ruff banned-api rule now fails the build on import requests or import httpx, since the defect is invisible to commercial-cloud testing and would otherwise return unnoticed
  • Entra signing keys cached: Token validation fetched the tenant JWKS document and rebuilt the RSA public key from its JWK on every authenticated request, spending a TLS handshake, an HTTP round trip and a key conversion per API call. Keys are now cached in process following Microsoft's signing key rollover guidance — a one hour refresh with jitter, a twenty four hour hard expiry, and a forced refresh on an unknown kid floored at once per five minutes. Concurrent refreshes collapse to a single fetch, and a failed refresh serves the last known good keys rather than failing closed. The floor also closes an amplification vector, where a flood of tokens bearing bogus kid values previously produced one outbound fetch each on behalf of unauthenticated callers
  • Address space checks no longer pay for full network detail: Every occupancy and utilization check loaded each network in the tenant complete with its subnets, peerings and Cosmos DB parentage, then read nothing but the resource ID and address prefixes. On a large tenant that is a subnet mv-expand of up to 1,024 rows per virtual network, two peering joins, one ARM call per virtual hub and two Cosmos DB queries — on every request — and a direct contributor to the Resource Graph throttling that surfaces as failed allocations. A prefix-only query now backs the thirteen call sites that need nothing further. Most expand responses still use the detailed query because they return whole network objects, but the available-networks endpoint is an exception: its expanded response is modelled on NetworkExpand, whose six fields are exactly what the prefix-only query already returns, so every additional field the detailed query produced was discarded field-for-field. That endpoint backs the network association picker in the UI, which always requests expansion, so the most frequently exercised expanded path now costs one Resource Graph query rather than two. Nothing is cached: results are still read live on every request, so allocation decisions remain based on the current state of Azure. Verified against a live tenant of 158 networks as returning an identical network set, identical prefixes and an identical total address count, with utilization responses byte-for-byte unchanged
  • Azure Resource Graph failures are now classified instead of reported as access denied: Every non-success response from Resource Graph was caught as a bare HttpResponseError and re-raised as 403 Access denied, so throttling, upstream outages and expired credentials all presented as authorization failures — sending users to audit permissions for a fault that had nothing to do with them. Because ClientAuthenticationError subclasses HttpResponseError, the Token has expired handlers in the three Resource Graph wrappers were unreachable. Failures are now classified at a single choke point: throttling returns 429 with Retry-After, upstream and unexpected statuses return 502 rather than being blamed on the caller, authentication failures return 401 — or 500 when the Azure IPAM service principal itself failed to authenticate — and a genuine Forbidden still returns 403. The HTTP exception handler now propagates response headers so Retry-After reaches the caller. Percent-style placeholders in loguru calls, which silently dropped their arguments, were converted at the same time
  • Data Factory discovery now honours subscription exclusions: Data Factory lookup built its own Resource Graph client around a hardcoded query instead of going through the shared query path, so it never applied the configured subscription exclusions and returned factories from subscriptions an administrator had deliberately excluded. It now uses the shared path with a DATA_FACTORY query carrying the standard exclusion filter, which also removes a duplicated Resource Graph client and its per-call setup
  • Virtual Machine Scale Set discovery now skips excluded subscriptions: Scale Set discovery runs through the Azure SDK rather than Resource Graph, and enumerated every subscription visible to the credentials — so excluded subscriptions were still queried, returning resources the administrator had excluded and spending an SDK call per subscription to do it. Subscription enumeration now accepts the exclusion list and filters before any Scale Set call is made
  • Scale Set addresses no longer fail on the sovereign cloud discovery path: The Scale Set endpoint expanded each record over a private_ips array, which is the shape Resource Graph returns. The SDK path used where Resource Graph coverage is incomplete — sovereign clouds such as IL6 — returns a singular private_ip, and those records raised a KeyError instead of being returned. Both shapes are now handled
  • vWAN transit network matching no longer depends on a regular expression: When a Virtual Network connects to a vWAN hub, Azure peers it to a generated transit network named HV_<hub>_<suffix>, which Azure IPAM substitutes back to the hub itself for display. The hub name was interpolated into a pattern without escaping, and Azure permits periods in hub names — so a hub named my.hub produced a pattern in which the period matched any character, and myXhub would have matched too. The pattern was also case-sensitive against a Resource Graph identifier. The transit network is now matched as a lowercased substring, which is all the comparison ever needed and removes both problems
  • Azure resource IDs are now compared case-insensitively: Azure returns resource IDs with inconsistent casing — Microsoft documents this and directs callers to "always perform a case-insensitive comparison of names" — and Azure IPAM stores whichever casing the client supplied when a network was associated, since association validates case-insensitively but saves the request body verbatim. Fifteen comparisons matched those two values exactly, so a casing difference made a managed network invisible: utilization under-reported the address space in use, Block CIDR and external network validation skipped the network entirely, reconciliation marked it inactive, and the CIDR check reported it as belonging to no Space or Block. Every one of these failed in the same direction — occupied space reading as free, which is the failure mode an IPAM can least afford. One of them compared a Resource Graph ID against an ARM SDK ID, where the two APIs genuinely disagree on casing. Reservation identifiers, which Azure IPAM generates itself, remain case-sensitive
  • Scale Set discovery no longer depends on resource ID casing: Six expressions pulled the resource group, virtual network, subnet and instance number out of ARM resource IDs by matching the resourceGroups/, providers/, virtualNetworks/, subnets/ and virtualMachines/ segments literally. Unlike the comparison bugs above these failed loudly rather than silently — a non-matching expression returned no match and the immediate .group(0) raised, so the whole Scale Set request became a 500 — but the trigger is the same casing inconsistency, and the same resource can be returned with different casing by different Azure APIs. This is the SDK-based Scale Set path, which runs only outside Azure Public; commercial deployments answer the same request from Resource Graph and were never affected. That makes it a sovereign and air-gapped cloud fix, which is also where Azure IPAM depends on the SDK paths most, because Resource Graph coverage there is incomplete. The segment matches are now case-insensitive, which is what the equivalent expression in the Reservation path already did
  • Block utilization is computed in one place: The arithmetic that populates size and used on a Block, its networks and their subnets was duplicated across the four endpoints that report utilization — roughly thirty lines repeated four times, differing only in local variable names and in whether the running Space totals were accumulated alongside. That duplication is what allowed the Space utilization defect above to exist in exactly one of the four copies. The four copies collapse to a single add_block_utilization helper, with the Space totals now summed from each Block's own figures, which is arithmetically the same value since the Space totals were only ever the sum of their Blocks. Around seventy lines are removed. The behaviour is deliberately unchanged, including two long-standing quirks in the per-network figures that are documented and fixed separately, and equivalence was verified both against the original logic across every network shape — fully in-Block, partially in-Block, entirely outside the Block, virtual hubs, networks without subnets, unmatched associations and empty Blocks — and byte-for-byte against a live tenant across all twenty-eight utilization responses
  • External networks now report utilization the way Azure networks do: A Block counted each external network's full address range toward its own utilization but reported nothing about how that range was allocated internally, even though external subnets and their endpoints were added to the data model later. External networks and their subnets now carry size and used alongside the Azure networks they sit beside. An external network's used is the address space assigned to its subnets, mirroring how a Virtual Network's used is the space assigned to its subnets; an external subnet's used is the number of endpoints defined within it, mirroring the consumed addresses reported for an Azure subnet. Endpoints are counted as they are, without the five addresses Azure reserves in every subnet, because an external network is not Azure and reserves nothing. The Block's own used is deliberately unchanged and still counts each external network's full range exactly once, since its subnets sit inside that range and counting both would double count. This required utilization variants of the external network and subnet models, so the utilization responses gain two fields at each level and lose none
  • A network's utilization within a Block is measured on one basis: For an expanded network, size counted only the prefixes falling inside the Block while used counted the subnets of every prefix the network owns, including those outside it. The two figures answered different questions and were presented as a ratio, so a network whose address space straddles a Block boundary could report more space used than it has — on the verification tenant one reported size 256 against used 384. Separately, used was reset inside the loop over in-Block prefixes, so a network with no prefixes inside the Block never had it reset at all and reported the entire network's subnet total against a size of zero. used is now initialized once and counts only subnets that fall inside the Block, so both figures describe the same address space and used can no longer exceed size. Block and Space totals are unaffected because they never consulted subnets, and each subnet still reports its own size regardless of where it sits. The Azure IPAM interface never requested these figures, since it does not expand Spaces or Blocks, so this corrects the documented API for automation consumers rather than anything visible in the product
  • Unreadable Reservation tags are reported rather than silently ignored: The helper that splits a Reservation tag into IDs wrapped its work in a bare except returning an empty list, which conflated two unrelated situations — a network carrying no tag at all, which is the normal case for nearly every network, and a tag whose value could not be read. The Resource Graph query parses tag values as JSON so that an absent tag resolves to null, with the side effect that a value resembling JSON arrives as a list, object or number instead of text. Such a network was then treated as having no Reservation, so the Reservation waited to be fulfilled indefinitely and nothing was written to the log. The type check is now explicit, and reconciliation reports the network ID and the offending value once per run rather than once per comparison. Behaviour is otherwise unchanged and was verified identical across every input type the tag can produce: unreadable values are still skipped rather than guessed at, and no attempt is made to strip quotes or interpret JSON, since Resource Graph already removes surrounding double quotes and interpreting the rest would mean acting on a tag the user did not write as an ID
  • Azure Resource Graph throttling is now retried by the SDK rather than failing outright: Resource Graph reports throttling on a POST and states the wait in x-ms-user-quota-resets-after instead of the standard Retry-After. The Azure SDK's retry policy therefore never retried it — POST sits outside its method allowlist, and the Retry-After short circuit that would have overridden that never fires — so a throttled query failed on its first attempt even though 429 is in the SDK's own retryable set. The SDK's policy is now taught both facts rather than a second retry loop running alongside it, so it waits exactly the interval Resource Graph reported instead of backing off blindly, and a hand-picked backoff ceiling gives way to the SDK's own settings. One retry setting is deliberately overridden: Resource Graph quota is shared by every caller using the same credentials rather than allocated per user, and replenishes roughly every five seconds, so the SDK's default of three status retries could be spent in about fifteen seconds while a burst of requests was still draining. That allowance is raised to six, which bounds a throttled request at roughly thirty seconds of waiting rather than a failure. It is not raised further on purpose — beyond ten the total retry budget silently becomes the real limit, and a longer wait encourages users to reload the page, which adds load to the very quota the request is waiting on
  • Truncated Resource Graph results are no longer treated as complete: Resource Graph withholds the continuation token when it truncates a result set, so the missing rows cannot be paged for and the response is silently short while still reporting success. For an IPAM that is the most dangerous shape of failure, since an unseen network reads as free address space. The condition is now detected and reported as an error rather than returned as though whole
  • Space utilization no longer fails when a Block holds an external network: GET /api/spaces/{space}?utilization=true returned 500 for any Space with an external network in one of its Blocks. External address space was accumulated onto the space path parameter — a string — rather than the Space document, raising TypeError: string indices must be integers. The line was correct in the equivalent list handler, where space is the loop variable, and was copied into the single-Space handler without renaming. The list endpoint and the Block-level endpoints were never affected, which is why the fault went unseen
  • Virtual hubs now expand in Space and Block responses: Requesting expand=true on a Space or Block containing a vWAN hub silently returned unexpanded network references — just an ID and active flag — with an HTTP 200 and no error. The expanded network model requires a subnets field, which virtual networks carry and virtual hubs never had, so validation of the expanded shape failed and the Union response model quietly fell through to the plain reference shape. The effect was not limited to the hub: a single hub in a Block collapsed every network in that Block back to a reference. Virtual hubs now carry an empty subnet list, which is what they genuinely have, so the expanded shape validates. The Block networks and available-networks endpoints were never affected, as their response model does not require subnets
  • Virtual hub listing no longer fails once a hub is associated to a Block: GET /api/azure/vhub returned 500 as soon as any vWAN hub was associated to a Block, reporting Input should be a valid string for parent_block. A network can belong to Blocks in more than one Space, so the engine has always produced a list of Block names here; the response model declared a single string. The model was corrected to match the data rather than the reverse, since collapsing to one name would discard a real association. The defect was invisible on /api/azure/network, which returns the same hub data but declares no response model, and on the virtual network equivalent for the same reason

UI & UX Improvements

  • Drill-down navigation in Discover (fixes Navigation Improvements #183): Added hierarchical drill-down with multi-filter pass-through and hidden column auto-reveal when filters are active
  • Centralized authentication handling: New AuthHandler for MSAL error handling, centralized token acquisition via tokenService
  • Custom DraggablePaper component: Replaced react-draggable package with a purpose-built component
  • Associations UX: Shifted API work to Redux thunks, optimized data refresh, fixed infinite update loops by separating initial grid selection from user selection state
  • Planner data loading: Addressed issue where Planner would not load when data was incomplete (fixes Planner will not load #366)
  • Reservation UX: Fixed view reversion after cancelling a Reservation, fixed column sorting, and streamlined the interface
  • Unified grid experience: Centralized DataGrid and ConfigureGrid components with shared filter utilities, consistent loading overlays, and AG Grid custom styling (brightness filter for row hover)
  • Search bar: Fixed messaging when no resources are found in Azure IPAM scope
  • API errors without the standard envelope are now surfaced: The Axios response interceptor read error.response.data.error unconditionally, so any response that did not use the { error } envelope produced an Error with an empty message — and an error snackbar with no text at all. Unhandled engine exceptions return plain text, request validation failures return { detail }, and proxy errors return HTML; each is now resolved to a message, falling back to the Axios status text. This also fixes silent failures on any request rejected by model validation, which had always returned 422

Notifications & Service Management

  • New notification framework (GET /api/notifications, POST /api/notifications/{id}/resolve): A self-describing, API-first advisory system. Server-side detectors emit notifications that the UI — or any API / IaC consumer — can read, while remediation is resolve-by-reference: the client asks the backend to resolve a notification by id and the backend owns all the logic (admin-gated, with an active-notification guard). Adding a new advisory is a single detector module registered in one list
  • Registry-migration detector: Flags deployments still pulling container images from the legacy public registry (azureipam.azurecr.io, critical) or the development registry (azureipamdev.azurecr.io, warning) and offers a one-click remediation that repoints the App Service / Function LinuxFxVersion to registry.azureipam.com. The target image is validated as anonymously pullable before any change is applied, then the app restarts to pull it — complementing the update script's auto-migration with an in-app path
  • Update-available detector: Compares the running IPAM_VERSION against the latest published GitHub release and surfaces an informational notice linking to the update guide. Fails safe (no notification, no error) on network or rate-limit failures
  • Platform-migration detector: Flags deployments still running on the legacy Docker Compose App Service model (DEPLOYMENT_STACK == "LegacyCompose", critical) and links to the migration guide ahead of Microsoft's March 31, 2027 retirement of Docker Compose support for Azure App Service. Link-only guidance (no in-app remediation) since migration is a scripted, multi-step process run from the operator's workstation
  • Notification center (UI): AppBar bell with a severity-colored unread badge, a severity-ordered list (popover on desktop, full-screen sheet on mobile), a detail dialog with server-driven actions and early admin gating, per-item dismiss, mark-all-read, and a manual refresh that resyncs to the server's truth
  • Service restart gate (UI): A full-screen overlay shown while the backend restarts (e.g. after a registry switch). The live SPA stays in memory, polls /api/status, and confirms recovery via a changed service start time before reloading — with calm, time-keyed messaging and a deliberate manual-reload escape hatch if recovery runs long. Background polling is paused while the gate is up

Deployment, Build & Infrastructure

  • Configuration drift detection and remediation in update: The update script now compares an existing deployment against what a fresh deployment would produce today, presents the differences, and converges them on approval. Detection is based on live Azure resource state rather than the version originally deployed, so each difference is evaluated independently and a partially-current deployment only sees what it actually needs. Covers the container registry endpoint, Python runtime version, App Service startup command, health check, baseline app settings, creation of the staging slot, and Function App slot-sticky content-share settings. Existing values and user-added settings are never overwritten, and removal is restricted to settings Azure IPAM owns that no longer apply to the deployment's shape. Production is converged before the staging slot so a newly created slot inherits corrected values, and the two are kept in sync thereafter so a future swap cannot regress production
  • Version-aware update decisions: Version comparisons across all three update paths (public registry image label, private ACR repository build, and native GitHub release) now use semantic versioning rather than string equality, so prerelease suffixes order correctly. Downgrades are no longer silently applied and reported as updates — the update stops and points to -Force. Images built without a version stamp report the 0.0.0 Dockerfile default and are treated as an unknown version rather than a downgrade
  • Updates against stopped or unreachable deployments: The update script now detects application state up front and reports it. Configuration changes, image builds, and ZIP deployments all continue to work, restarts that would achieve nothing are skipped, and a stopped application is reported as remaining stopped. A new -ContainerType (Debian|RHEL) override on update matches the existing migrate switch, so a private ACR deployment can still be rebuilt when its container distro can't be probed. The distro probe is also now bounded by a timeout, so an application whose container fails to start no longer stalls the update indefinitely
  • Guidance instead of failures for anticipated conditions: update and migrate no longer raise exceptions for situations with a known remedy, such as an undetectable container distro. These now print actionable guidance under the relevant phase heading and exit cleanly, matching how legacy Compose deployments and out-of-resource-group registries were already handled. Genuinely unexpected errors now surface their message on screen rather than only in the log
  • DOCKER_REGISTRY_SERVER_URL removed from all deployment templates: This setting is only required for registries authenticating with stored credentials. Azure IPAM pulls anonymously from the public registry or with a managed identity from a private ACR, so it was never needed, and App Service removes it on its own whenever the container registry is reconfigured — which caused the update script to repeatedly report it as configuration drift
  • Azure PowerShell requirements aligned to 11.4.0: The deployment, update, and migration guides now state a single Azure PowerShell rollup version, and each script's granular #Requires module pins match the versions packaged in that rollup. Previously the update guide understated its requirement (Az.Resources 6.16.0 is only available from Az 11.4.0), and deploy.ps1 pinned the Az 10.3.0 module set while its documentation stated Az 11.0.0
  • Fixed Debian container build port: update.ps1 built Debian images with PORT=80 while deploy.ps1, migrate.ps1, and the Dockerfile default all use 8080
  • Public container registry migrated: The publicly hosted Azure IPAM container images have moved from azureipam.azurecr.io to the new registry.azureipam.com endpoint. The deployment and migration Bicep templates now reference the new registry by default, while the update script continues to recognize the legacy azureipam.azurecr.io endpoint so existing deployments keep working. The Docker Compose migration tooling intentionally still targets the legacy endpoint, as it only ever processes pre-existing legacy deployments
  • Migration robustness: The Docker Compose migrate script now halts on a non-standard or unresolvable container registry with guidance to re-run using the -JsonFile override, instead of silently falling back to the public registry. A new -ContainerType (Debian|RHEL) override handles cases where the source app is stopped/unreachable and its distro can't be auto-detected
  • RHEL images updated to UBI9: Including the container-image overrides in the deploy, update, and migrate PowerShell scripts
  • Node.js minimum raised to 22.22.0: Required by React Router v8; enforced in the build.ps1 version gate and the UI package.json engines field
  • Dockerfiles optimized: Improved layering, removed unneeded steps across all container images. Resolved Hadolint lint violations (ADDCOPY, JSON notation for CMD/ENTRYPOINT, pipefail for piped RUN commands). Added centralized .hadolint.yaml for rule suppressions
  • KeyVault Soft Delete added to deployment and migration Bicep templates (fixes Enable soft delete for Key Vault #373)
  • Build flexibility: Added support for building with either current or latest NPM/Python packages
  • CI/CD path exclusions: Added exclusions to avoid unnecessary test/build runs for documentation-only changes
  • GitHub Actions modernized: Bumped all workflow actions to their latest versions (checkout@v7, setup-node@v7, setup-python@v7, github-script@v9, create-github-app-token@v3, azure/login@v3, hadolint-action@v3.5.0), initially to clear the Node.js 20 runner deprecation and then to the current majors so the release does not ship already behind. The v7 line of the actions/* set is an ESM migration on the same node24 runtime, and none of the inputs these workflows pass were changed or removed. The fork-PR checkout restriction added in checkout@v7 does not apply here, as no workflow uses pull_request_target or workflow_run
  • GitHub Actions pinned to commit SHAs: Every action reference across all four workflows is now pinned to a full-length commit SHA with a trailing version comment, rather than a mutable tag such as @v7. A tag can be retargeted to arbitrary code by anyone able to push to the action's repository — the mechanism behind the tj-actions/changed-files and codfish/semantic-release-action compromises — and every one of these workflows holds credentials, whether an OIDC federated identity or a GitHub App private key. A new .github/dependabot.yml tracks the github-actions ecosystem weekly as a single grouped PR, with a seven day cooldown so a compromised release has a window to be reported before it is proposed here. Dependabot updates the SHA and the version comment together, so the pins do not go stale. Verified with the GitHub REST API that each pinned SHA is the commit the corresponding release tag resolves to
  • Azure PowerShell SDK v14 compatibility: Fixed Get-AzAccessToken breaking changes (fixes Breaking changes to Get-AzAccessToken #343); updated deploy & migrate scripts
  • OCI version labels on container images: All production images now carry standard OCI labels (org.opencontainers.image.version, .title, .source), stamped at build time via a new IPAM_VERSION build arg. This lets you determine the exact version behind a floating tag like latest by inspecting the registry — no pull or run required (e.g. docker buildx imagetools inspect or az acr manifest show). Covers all deb, rhel, and func variants across the root, engine, ui, and lb images, with the build workflow passing --build-arg IPAM_VERSION to every az acr build. Labels begin with the v4.0.0 release images
  • Release tag parsing hardened: Version extraction now strips only a leading v (^v) from the release tag, preserving suffixes such as -preview
  • Internet-restricted cloud deployments fixed: The ZIP-deploy path used by internet-restricted clouds (AZURE_US_GOV_SECRET) never resolved its bundled Python packages. init.sh referenced an APP_PATH variable that is defined nowhere in the repository — App Service exposes it only to interactive SSH sessions via ~/.bashrc, which a non-interactive startup command never reads — so PYTHONPATH expanded to a non-existent path at the filesystem root and no bundled dependency could be imported. The application root is now derived from the script's own location, matching what function_app.py already did for the Function App entry point. The accompanying PATH export was removed: it pointed at packages while pip installs console scripts to packages/bin, and nothing invokes them
  • Deployment archive now targets the App Service runtime rather than the build host: pip install --target resolves wheels for the machine running pip. When the GitHub Actions runner moved to Ubuntu 24.04, cryptography began resolving to a manylinux_2_34 wheel that cannot load on the App Service Python 3.11 image (Debian bullseye, glibc 2.31), producing GLIBC_2.33 not found at startup. Wheel resolution is now pinned explicitly (--only-binary=:all:, --platform manylinux2014_x86_64, --implementation cp, --python-version, --abi), with the Python tag and ABI derived from engine/app/version.json. manylinux2014 (glibc 2.17) is targeted deliberately, since sovereign clouds can run older stamps than commercial Azure
  • Native module verification gate: The build now fails if any bundled shared object targets the wrong Python ABI, requires a newer glibc than the platform tag allows, or is a Windows .pyd. Both defects above were invisible at build time and only surfaced after deployment — and the ABI mismatch presented as ModuleNotFoundError, which reads like a missing dependency rather than a build fault
  • Sovereign cloud root certificates now trusted: Secret cloud deployments (AZURE_US_GOV_SECRET, IL6) run against endpoints whose certificate chains are issued by that cloud's own roots, which are not present in the App Service image's default trust store. Every outbound TLS call the engine makes — Key Vault references, Cosmos DB, ARM, Microsoft Graph — therefore failed certificate validation. WEBSITES_INCLUDE_CLOUD_CERTS is now set for that cloud in the deployment, update, and migration templates, so the platform injects the cloud's root certificates. The update script's configuration drift detection also adds it to existing secret cloud deployments. This setting is necessary but was not sufficient on its own — see the TLS validation fix under Engine & Backend, without which the engine still could not validate tokens in that cloud
  • Managed identity role assignments no longer collide on redeploy: Deployments that pin resource names via -ResourceNames failed with RoleAssignmentUpdateNotPermittedTenant ID, application ID, principal ID, and scope are not allowed to be updated — whenever the managed identity had been deleted and recreated. Role assignment names are globally unique GUIDs, and the Contributor and Managed Identity Operator grants seeded theirs from the identity's resource ID, which survives a delete and recreate unchanged. The name therefore matched an existing assignment while the principal behind it had changed, which Azure treats as an illegal update rather than a create. Every other module already seeded from the principal ID and was never exposed; managedIdentity.bicep could not, because a role assignment name must resolve at the start of deployment (BCP120) and the principal ID does not exist until the identity is created. Both grants moved into a new managedIdentityRoles.bicep module that receives the principal ID as a parameter, following Microsoft's documented guid(scope, principalId, roleDefinitionId) pattern — a recreated identity now yields a new assignment name and deploys cleanly, with no manual cleanup. Because the names change, a first re-run of deploy.ps1 against an intact deployment created by an earlier version may report RoleAssignmentExists; removing the two superseded assignments clears it. Deployments that let Azure IPAM generate resource names are unaffected either way, as every run produces fresh names. update and migrate needed no equivalent change — update deploys no role assignments at all, and migrate references the identity as existing and already seeds from the principal ID
  • Build script failure reporting: Seven prerequisite checks ended in a bare exit, which returns 0. A build that could not find NodeJS, or that rejected the installed Python version, printed errors and then reported success to CI. All now exit non-zero. The exception handler also printed an unassigned variable in place of the log path
  • Faster archive creation: Compress-Archive took 353 seconds to package the 6,404-file deploy archive; ZipFile.CreateFromDirectory produces an identically sized result in 11 seconds. The new API also preserves Unix file modes, which Compress-Archive flattened to 0644
  • CI Python version single-sourced: All three workflows now read the interpreter version from engine/app/version.json instead of pinning 3.11 by hand. This matters most in the versioning workflow, which regenerates requirements.lock.txt — dependency resolution is Python-version sensitive, so a lock file produced on the wrong interpreter can omit packages the runtime requires

Documentation Overhaul

  • Massively revamped How-To docs: authentication, exclusions, Discover, Reservations, External Networks, and Virtual Network Associations sections
  • New documentation: Comprehensive automation docs, API docs for vNet Associations, initial External Networks docs, detailed Reservations feature docs
  • 50+ screenshots added or replaced to reflect the current UI
  • Fixed all markdown warnings/errors and added .markdownlint.json configuration
  • Doc cleanup: Fixed stale links, grammar, spelling issues, deprecated folder descriptions, and undocumented switches across all sections

Examples

  • Terraform example revamped: Migrated from Shell scripts to the official Azure IPAM Terraform provider
  • Azure ESLZ example modernized: Updated resource API references, modern coding standards, and improved parameterization
  • Script examples reorganized: PowerShell and Shell scripts moved into a dedicated examples/scripts/ folder with new helper scripts and README
  • Token helper function: Standardized access token generation; removed legacy Microsoft Graph SDK v1 support

Testing

  • Added tests for Virtual Network Association permutations
  • Expanded overall testing coverage with numerous additional Pester tests
  • Updated test expectations to align with additionally created resources
  • Added Tools coverage for Block list evaluation — falling through to a later Block when an earlier one cannot satisfy the requested size (with and without smallest_cidr), and the genuinely exhausted case
  • CI lint gate: Added pre-deployment lint job to the testing workflow (ESLint, Vite build verification, Ruff for Python, PSScriptAnalyzer for the PowerShell scripts, Bicep template validation, Hadolint for Dockerfiles). Deploy is skipped if any check fails, preventing wasted Azure resources
  • PowerShell scripts are analyzed in CI: No PowerShell file had ever been linted automatically, despite the repository carrying a PSScriptAnalyzerSettings.psd1, which is how the test suite came to be the only script not following the house conventions and the only one carrying an analyzer finding. The lint job now runs PSScriptAnalyzer over every script and fails on a finding at any severity, and also fails if it matches no scripts at all, so the check cannot pass green having analyzed nothing. The step is deliberately scoped to the four rules the settings file configures: running PSUseCorrectCasing alongside the full default rule set trips a thread-safety defect in the analyzer's command cache and aborts the run roughly a third of the time, which would fail the build on a clean tree. That is PSScriptAnalyzer #1708, open since 2021 and reproducible in both 1.24.0 and 1.25.0, so pinning an older analyzer does not avoid it. The scoping is annotated in the workflow with the single change needed to undo it once the defect is fixed
  • PowerShell conventions applied uniformly: The Pester suite used capitalized Function and Param, making it the only script not following the lowercase keyword style the settings file describes and every other script applies, and it carried the repository's only analyzer error — a false positive from normalizing an Azure access token to a SecureString, now suppressed inline with the same justification already used in the deployment scripts. version.ps1 declared pipeline binding on all four parameters without a process block, so piping several objects would have silently processed only the last; nothing pipes to it, and no comparable script declares that binding, so the unused bindings were removed. Every script also declared $logPath and then referenced $logpath on the following line, a mismatch copied into all five. Every PowerShell file in the repository now reports no analyzer findings at any severity
  • Transient failure handling in the API test suite: The integration suite failed outright on the first transient error, so an Azure Resource Graph throttle during a run looked like a product regression rather than an environmental hiccup. The shared request helpers now retry only 429 and 5xx responses, honouring Retry-After; the client errors the suite deliberately asserts are never retried, and transport failures are never replayed, since a POST may already have applied server-side. Access tokens are cached and refreshed ahead of expiry, so a long run no longer fails partway through on an expired token
  • Utilization and network expansion coverage: Added a Utilization & Expansion context exercising utilization=true and expand=true across the Spaces, Space, Blocks and Block endpoints. Neither parameter had any coverage at all, which is precisely why the Space utilization defect above survived. The context asserts that the utilization reported without expansion matches the expansion path exactly — the two are answered by different Resource Graph queries, so this guards the prefix-only query against divergence; that a Block's used address count equals its networks plus its external networks; and that expanded responses genuinely carry the expanded fields, since a Union response model degrades silently rather than erroring. Every assertion was confirmed to fail against the unfixed engine before the corresponding fix landed, so none of them can pass vacuously. All are read-only, leaving the suite's ordered state untouched
  • External network and per-network utilization coverage: The utilization tests asserted Block and Space arithmetic but nothing verified a network's own size and used, nor that external networks reported utilization at all, so both defects fixed in this release would have passed the suite. Four assertions were added: that an expanded network never reports more space used than its size, that an external network's used is the space assigned to its subnets, that an external subnet's used is its endpoint count, and that a Block counts an external network once rather than counting its subnets again. The first three were confirmed to fail against the code that carried each defect; the fourth is a guard that would catch an external network being double counted if the accounting were ever changed

Bug Fixes

[major]

…that were not being accounted for properly
…ue to improper handling of missing vNETs and vHUBs
DCMattyG added 21 commits July 23, 2026 16:13
…tion

Add a platform-migration detector that emits a critical, non-dismissible
notification when DEPLOYMENT_STACK == "LegacyCompose", ahead of Microsoft's
March 31, 2027 retirement of Docker Compose support for Azure App Service.

Link-only guidance (no server-owned resolve) pointing at the migration guide,
mirroring the update detector: migration is a scripted, multi-step process run
from the operator's workstation. Registered in the detector list and documented
in the v4.0.0 release notes.
…tput

Skip unnecessary work by checking whether a deployment is already current
before building, restarting, or running prerequisite checks.

- Private ACR: compare running vs. repo version first and short-circuit when
  up to date, before the ACR lookup and CLI/PowerShell context checks
- Public ACR: add Get-RegistryImageVersion to read the OCI version label via
  the Registry v2 API; skip the restart when current, fall back to
  restart-to-update when the label is missing (pre-OCI) or unreadable
- Public ACR: drop the redundant explicit restart after a registry repoint,
  since updating LinuxFxVersion already recycles the App Service
- -Force still deploys latest regardless
- migrate: handle a non-Compose deployment gracefully (guide to update.ps1)
  instead of throwing
- Align update/migrate output: inline "Success"/"Failed" status, per-attempt
  restart lines, and a dedicated "Fetching build logs" step

Affects: update/update.ps1, migrate/migrate.ps1
Pin azure-mgmt-resource-subscriptions to 1.0.0 in requirements.lock.txt
since the previously pinned 1.1.0 does not exist on PyPI and broke
dependency resolution.
Shorten the notification title from "Action required: Azure IPAM
container registry has moved" to "Azure IPAM container registry has
moved" for cleaner wording.
The docs state a rollup version while the scripts pin individual modules for
load-time performance, but the two had drifted apart:

- docs/update stated Az 10.3.0 while update.ps1 requires Az.Resources 6.16.0,
  which first ships in Az 11.4.0, so following the prerequisites exactly would
  fail the #Requires check.
- deploy.ps1 pinned the module set from Az 10.3.0 while its documentation
  stated Az 11.0.0.

Align every Az module pin with the versions packaged in Az 11.4.0 and state
that rollup version consistently across the deployment, update and migration
guides. migrate.ps1 already matched. The Microsoft.Graph modules are not part
of the Az rollup and are unchanged.
$APP_PATH was referenced in init.sh but never defined anywhere in the
repo. App Service does not export it to the startup command's
environment; Oryx only appends it to ~/.bashrc, which a non-interactive
shell never sources. PYTHONPATH therefore expanded to ":/packages", a
non-existent path at the filesystem root, leaving every bundled
dependency unimportable.

Derive the application root from the script's own location instead,
mirroring what engine/function_app.py already does for the Function App
entry point.

Also:

- Use ${PYTHONPATH:+...} so an unset PYTHONPATH no longer yields a
  leading colon, which silently placed the working directory on the
  import path.
- Drop the PATH export. It pointed at "packages" while pip installs
  console scripts to "packages/bin", those files ship non-executable
  (0644), and nothing invokes them since the app starts via
  "python -m uvicorn".

Only reachable when WEBSITE_RUN_FROM_PACKAGE is set, which is limited to
internet-restricted clouds (AZURE_US_GOV_SECRET). Every changed line is
inside that guard, so the Debian and RHEL single-container images are
unaffected.
pip install --target resolves wheels for the machine and interpreter
running pip, not for the deployment target. When the GitHub Actions
runner moved to Ubuntu 24.04, cryptography began resolving to a
manylinux_2_34 wheel, which cannot load on the App Service Python 3.11
image (Debian bullseye, glibc 2.31), and deployments failed with
"GLIBC_2.33 not found".

Passing --platform alone is not sufficient. pip still fills in the
Python version from the running interpreter, producing cpython-312
extension modules that Python 3.11 silently ignores, which surfaces as
an unrelated-looking ModuleNotFoundError.

Pin all four resolution inputs, deriving the Python tag and ABI from
engine/app/version.json so a runtime bump carries through without
further edits:

  --only-binary=:all: --platform manylinux2014_x86_64
  --implementation cp --python-version <ver> --abi cp<ver>

manylinux2014 (glibc 2.17) is chosen for maximum compatibility, since
sovereign clouds can run older stamps than commercial Azure. The
widening path is documented inline should a dependency ever drop those
wheels.

Add a post-install verification gate that fails the build when a native
module targets the wrong Python ABI, requires a newer glibc than the
platform tag allows, or is a Windows .pyd. Both defects above were
invisible at build time and only surfaced after deployment.

Container images are unaffected; they pip install inside the target
image, so their wheels match by construction.
Seven prerequisite checks ended in a bare `exit`, which returns 0. A
build that could not find NodeJS, or that rejected the installed Python
version, printed red errors and then reported success. In CI the step
went green, and the failure only surfaced later — if at all — when the
release upload found no artifact.

The exception handler also printed `$buildLog`, a variable that is never
assigned, so the "Build Log:" line showed an empty path at precisely the
moment it was needed. Point it at $transcriptLog.

Remove the unreachable condition guarding ZIP creation. $npmBuildErr and
$pipInstallErr are never assigned anywhere in the script, so the test was
always true and its else branch was dead code that also used a bare
`exit`. Control flow is unchanged: failures are caught by the
$LASTEXITCODE checks around npm and pip and by the native module
verification gate, all of which throw into the outer handler.

Also bring the script to zero PSScriptAnalyzer findings, matching
deploy.ps1, update.ps1 and migrate.ps1:

- Drop ValueFromPipelineByPropertyName from all three parameters. The
  script has no process block and is never pipeline-fed, so the binding
  was vestigial. Same fix previously applied to the other scripts.
- Correct Get-Date and -Format casing.
- Suppress PSAvoidUsingPositionalParameters at script scope; npm is an
  external executable, not a cmdlet, so its arguments are not
  PowerShell positional parameters.

Removing the if wrapper re-indents the archive assembly block. Review
with `git diff -w` to see the functional change in isolation.
All three workflows pinned python-version: '3.11' by hand, duplicating a
value that engine/app/version.json already owns. Nothing kept the two in
sync, so a runtime bump would silently leave CI on the previous
interpreter.

This matters most in azure-ipam-version.yml, which regenerates
requirements.lock.txt. Dependency resolution is Python-version sensitive,
so a lock file produced on the wrong interpreter can omit packages that
the target runtime requires, yielding an archive that fails only at
deploy time.

Read the value with jq after checkout and feed it to setup-python.
Reordering was required to make the file readable:

- azure-ipam-assets.yml: checkout promoted to the first step
- azure-ipam-version.yml: setup-python moved below checkout. Checkout
  cannot move earlier because it consumes the GitHub App token generated
  two steps prior.
- azure-ipam-testing.yml: already ordered correctly; read step only

No workflow changes what it does; only which interpreter is installed
and where two steps sit in the sequence. Container base images and the
python:3.11-slim serve images remain pinned separately and are unchanged.
uvicorn was started with --reload in both production init scripts. The
flag is documented as a development aid, and it caused two concrete
problems.

It masked crashes. Under --reload, uvicorn runs a supervisor plus a
child. When the child died during startup the supervisor stayed alive,
so the container kept running and both Docker and App Service saw a
healthy process serving nothing. Removing the flag makes the process
exit on failure, so the platform restarts it and the fault is visible.

It also polled the filesystem continuously. Under
WEBSITE_RUN_FROM_PACKAGE the watched directory is the read-only ZIP
mount, so the watcher tracked thousands of files that can never change.
Termination was slow as well: a signalled shutdown took roughly two
minutes with the reloader versus about one second without it, which
would surface as sluggish restarts and slot swaps.

Verified against the published container image and against a
run-from-package layout: no reloader process is started, and a startup
failure now exits non-zero instead of leaving the container up.

engine/Dockerfile.dev keeps --reload; hot reload is intended there.

Worker count is deliberately left at the default. The reservation
scheduler runs in-process, so additional workers would schedule it more
than once.
Compress-Archive took 353 seconds to package the 6,404-file deploy
archive; ZipFile.CreateFromDirectory produces a byte-identical 49.6 MB
result in 11 seconds. The cost grows with the dependency tree, so it is
paid on every release build.

Resolve the destination to an absolute path before handing it to .NET.
PowerShell cmdlets resolve relative paths against the session location,
whereas System.IO resolves against the process working directory, and
the two are not kept in sync. The release workflow invokes the script
from tools with -Path ../assets, so the archive path resolved outside
the repository entirely and the call failed.

Use -LiteralPath on both Convert-Path calls. With -Path, a directory
containing square brackets is treated as a wildcard character class,
matches nothing, and returns empty without raising an error, which would
silently produce a malformed destination.

The new API also preserves Unix file modes, where Compress-Archive
flattened everything to 0644. Console scripts under packages/bin have
been shipping non-executable as a result; they are now packaged with
their original permissions. Nothing in the deployment depended on the
old behaviour, since shared objects need read rather than execute
access and init.sh is invoked as an argument to bash.

Compress-Archive emitted nine additional directory records. Each has
descendant files, so every directory is still created on extraction and
archive contents are unchanged.
max() raised ValueError on an empty candidate list, aborting the block
list loop in nextAvailableVNet instead of evaluating the remaining
blocks. A request that only a later block could satisfy failed with
HTTP 500.

The same guard was missing in nextAvailableSubnet and in the single
block reservation path, where an unhandled exception replaced the
intended error detail.

Adds integration coverage for the fall-through and for the genuinely
exhausted cases.

Fixes #384
The response interceptor read error.response.data.error unconditionally,
so any response that did not use the { error } envelope produced an
Error with an empty message and an error snackbar with no text.

Unhandled engine exceptions return plain text, request validation
failures return { detail }, and proxy errors return HTML. Each of these
is now resolved to a message, falling back to the Axios status text.

Refs #384
Capacity exhaustion and malformed request bodies were reported as 500,
making them indistinguishable from genuine server faults, triggering
pointless client and proxy retries, and counting against server error
metrics.

Requests that cannot be satisfied by current allocation state now return
409, as does a JSON patch that conflicts with the stored document. A
malformed JSON patch now returns 400. Both already matched the codes
used by neighbouring checks in the same handlers.

Errors caused by Azure, Cosmos DB or the site configuration are
unchanged and still return 500.

BREAKING CHANGE: /api/tools/nextAvailableVNet, /api/tools/nextAvailableSubnet
and the reservation and external network creation endpoints return 409
rather than 500 when no network of the requested size is available. PATCH
endpoints return 400 rather than 500 for a malformed JSON patch. Clients
that branch on 500 for these conditions must be updated.
…deployments

Secret cloud (AZURE_US_GOV_SECRET / IL6) deployments run against endpoints whose
certificate chains are issued by that cloud's own roots, which are absent from the
App Service image's default trust store. Every outbound TLS call the engine makes -
Key Vault references, Cosmos DB, ARM, Microsoft Graph - failed certificate
validation as a result.

Set WEBSITES_INCLUDE_CLOUD_CERTS so the platform injects the cloud's root
certificates. The setting is added to the existing runFromPackage branch in the
deployment and slot templates, since that branch already identifies the
internet-restricted cloud shape. The Compose migration template carries its own
condition, as it has no shape branching to fold into.

The update script's configuration drift detection now converges existing secret
cloud deployments onto the setting, for both production and the staging slot, and
tracks it as shape-conditional so it is removed if a deployment is not on that cloud.
JWKS fetch and the metrics heartbeat now use aiohttp, which verifies
against the OS trust store where WEBSITES_INCLUDE_CLOUD_CERTS injects
sovereign cloud roots. requests/certifi cannot see those roots, so
token validation failed in secret clouds (IL6) while passing every
commercial-cloud test.

Adds a ruff banned-api rule so requests/httpx can't be reintroduced.
Every authenticated request fetched the tenant JWKS document and rebuilt
the RSA public key from its JWK, spending a TLS handshake, an HTTP round
trip and a key conversion on each API call. The cost is worst in
sovereign clouds, where egress is constrained and proxied.

Signing keys are now cached in process as PEM bytes keyed by kid,
following the algorithm in Microsoft's signing key rollover guidance: a
one hour refresh interval with jitter, a twenty four hour hard expiry,
and a forced refresh when a token presents an unknown kid, floored at
once per five minutes. Refreshes are serialised behind a lock, so a
burst of concurrent requests produces a single fetch, and a failed
refresh keeps serving the last known good keys until the hard expiry
rather than failing closed.

The unknown kid floor also closes an amplification vector. A flood of
tokens bearing bogus kid values previously produced one outbound fetch
each, on behalf of unauthenticated callers.

Cache lifetime does not affect rotation safety. Entra derives kid from
the certificate thumbprint, so a new signing key always arrives under a
new kid and takes the forced refresh path however fresh the cache is. A
longer lifetime trades only against how long a revoked key stays
honoured.

The attempt timestamp driving the five minute floor is recorded before
the fetch rather than after a success. Recording it on success alone
would leave a failing endpoint unthrottled, since each failure would
leave the timestamp untouched and every request would retry.

No background refresh job is used. The scheduler runs only on the App
Service production slot, so it would never fire for the Function App,
which serves the same ASGI application and validates tokens too. The
cache is per process, which suits every hosting shape: App Service and
container deploys run one uvicorn worker per instance, and the Function
App runs one language worker.
@jagiraud

Copy link
Copy Markdown

DCMattyG any idea when this will be merged/released?

@DCMattyG

Copy link
Copy Markdown
Contributor Author

Jacob Giraud (@jagiraud), the release will be extremely soon, I'm just working through one final issue with a customer to make sure the fixes make it into the v4.0.0 release. Hoping by the end of this month, if not sooner.

If you need to test any of these new features sooner, please let me know and I can provide you a way to test out the current development code/container.

Hope that helps!

DCMattyG and others added 7 commits August 24, 2026 10:12
Redeploying with -ResourceNames failed on the managed identity module with
"Tenant ID, application ID, principal ID, and scope are not allowed to be
updated" whenever the identity had been deleted and recreated.

Role assignment names are globally unique GUIDs, and the Contributor and
Managed Identity Operator grants seeded theirs from the identity's resource
ID. That ID is stable across a delete and recreate, so the name collided
while the principal behind it had changed, which Azure treats as an illegal
update rather than a create.

Every other module already seeds from principalId and was never exposed.
managedIdentity.bicep could not do the same, because a role assignment name
must resolve at the start of deployment (BCP120) and principalId does not
exist until the identity is created.

Move both grants into managedIdentityRoles.bicep, which receives principalId
as a parameter and can therefore use it in the name, following the documented
guid(scope, principalId, roleDefinitionId) pattern. A recreated identity now
produces a new name and creates cleanly, with no manual cleanup required.

Assignment names change as a result. A first re-run of deploy.ps1 against an
intact deployment created by the previous template may report
RoleAssignmentExists; removing the two superseded assignments clears it.
Deployments without -ResourceNames generate fresh names per run and are
unaffected.

update and migrate need no equivalent change: update deploys no role
assignments at all, and migrate references the identity as existing and
already seeds from principalId. The Key Vault and Cosmos seeds are
deliberately left untouched, since they are identical across deploy and
migrate and must stay that way.
Upated NPM packages to latest versions.
Pin every action reference to an immutable commit SHA with a trailing
version comment, and add a Dependabot configuration for the
github-actions ecosystem so the pins stay current. Pinning by SHA
prevents supply-chain attacks in which a mutable tag is retargeted to
malicious code, per the GitHub Actions security hardening guidance. The
7 day Dependabot cooldown leaves a window for compromised releases to be
reported before they are proposed as updates.

Also move every action to its latest GA release so the branch does not
ship already behind: checkout v7.0.1, setup-node v7.0.0 and setup-python
v7.0.0 (all ESM migrations, both majors already on the node24 runtime),
plus hadolint-action v3.5.0. The inputs in use are unchanged across the
major boundary, and the fork-PR checkout restriction added in checkout
v7 does not apply here as no workflow uses pull_request_target or
workflow_run.

hadolint-action v3.5.0 upgrades the linter from v2.14.0 to v2.15.1,
which adds DL3066 and flags the intentional `USER root` directives in
the RHEL images. Ignore DL3066 so the lint job stays green; this affects
linting only and leaves the built images unchanged.

Supersedes #386, which was authored against main and pinned to action
versions that predate the OIDC migration and action bumps on this
branch.

Co-authored-by: Dan Fiedler <151573964+danfiedler-msft@users.noreply.github.com>
Handle both aggregated ARG private_ips arrays and singular private_ip
records returned by the sovereign-cloud SDK path.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

4 participants