Microsoft will start switching on memory integrity by default for eligible Windows 11 devices from October 2026, arriving through that month’s Patch Tuesday update on October 13.
The feature has existed for years, but plenty of PCs that qualify for it have not had it turned on, and this change closes that difference without anyone needing to flip a setting. Here is what you need to know:
What is Memory Integrity and why does Microsoft want it enabled?
Memory integrity is also called Hypervisor-protected Code Integrity, or HVCI. It runs in addition to Virtualization-based Security, using your CPU’s virtualization hardware to make an isolated environment that the rest of Windows cannot reach directly.
Memory Integrity turned on in Windows Security. Credit: Windows Latest
Kernel-mode drivers, the lowest-level software running on your PC, have to pass through this isolated checker before Windows lets them execute. Malware that tries to sneak in through a vulnerable or unsigned driver gets blocked at this layer instead of reaching the kernel.
Which PCs get Memory Integrity automatically enabled in October?
Not every Windows 11 PC qualifies, but Microsoft’s hardware requirements for automatic enablement are pretty lenient:
Intel 8th-generation processor or newer,
AMD Zen 2 or newer,
Qualcomm Snapdragon 8180 or newer,
at least 8GB of RAM on x64 systems,
a 64GB SSD, virtualization enabled in firmware,
and drivers that are already confirmed compatible with memory integrity.
Secured-core PCs, a certification tier Microsoft and OEMs apply to business and enterprise hardware, already ship with this turned on today.
Microsoft has also built in a safety check before flipping the switch. Windows will run a readiness assessment on each device first, instead of pushing the change to every PC that meets the minimum hardware.
If your device already has memory integrity turned off on purpose, whether through Group Policy, Intune, or a manual registry change, Windows Update will respect that and leave it alone. The rollout works as an opt-out by policy, and not a forced override.
The Memory integrity toggle shows On or Off directly on that page.
IT admins can pre-configure the feature through registry keys under HKLM\System\CurrentControlSet\Control\DeviceGuard\Scenarios\HypervisorEnforcedCodeIntegrity, or manage it at scale through Group Policy and Intune.
Incompatible drivers are the main reason this has not been on for everyone already. A driver written years ago that pokes at memory in ways memory integrity does not allow will get blocked from loading, and in bad cases, that has caused boot failures on older hardware in the past.
Windows logs these conflicts under Event Viewer, in Applications and Services Logs > Microsoft > Windows > CodeIntegrity > Operational, tagged with Event ID 3087 when a driver gets flagged as incompatible. Check that log first if a PC starts acting up after this rolls out.
Why Microsoft is pushing this now
Windows security has been under a different kind of pressure lately. AI-assisted vulnerability research has gotten fast enough that researchers, and presumably attackers, can find kernel-level bugs in Windows far quicker than before.
Core isolation in Windows Security. Credit: Windows Latest
Turning on a moderation that was already unused on millions of eligible PCs is a cheap way to close off a whole class of attack without waiting on a slower fix somewhere else in the OS.
Memory integrity was not a Windows 11 exclusive to begin with. It shipped as an opt-in VBS feature going back to Windows 10, and it became the default only on clean installs of Windows 10 in S mode and, later, Windows 11 on hardware that met the requirements.
Millions of PCs upgraded instead of being clean installed over the years, which are the kind of devices Microsoft is targeting with this October update. All these are hardware that has always qualified but had the protection switched off.
Should you turn Memory Integrity off? Usually, no
If your PC does get memory integrity switched on and something breaks, usually an older driver failing to load, Windows Security will flag it under Core isolation, and the CodeIntegrity event log will name the driver responsible.
From there, you just have to wait for an updated driver from the manufacturer or turn the feature off manually while you wait, which is a trade-off PCs with memory integrity already enabled have lived with for years.
As we enter the next era, what will be the defining measure of our progress?
Every industry has a word that shapes how it thinks. For pilots, it’s safety. For insurers, it’s risk.
For the semiconductor industry, it’s yield.
Yield does not ask how elegant the solution is, how many years it took or what the roadmap promised. Rather, it asks one simple question: What useful output did we produce?
For more than 60 years, the semiconductor industry has asked that question, relentlessly maximizing the number of usable chips produced from every wafer. Generation after generation, wafer after wafer, it is precisely that discipline that turned the transistor from a laboratory curiosity into the foundation of modern life.
Today, we need to apply the same principle to the unprecedented resources that the world is pouring into AI: capital on a scale once reserved for nations, gigawatts of power and record-breaking fabs and datacenters.
The question that will define this decade is the same one this industry has always asked: What actually comes out? Not just chips and tokens, but as affordable intelligence, as work that matters and as outcomes that improve lives.
That is the yield imperative.
YouTube Video
The limit of more
AI has proliferated with remarkable speed, at a rate of adoption faster than the internet, the PC or even the smartphone. Yet, global penetration still stands at just 18% of the working population, and the vast majority of that usage is chat-based. As systems move from answering individual prompts to reasoning, planning, using tools and executing longer agentic workflows, the infrastructure equation changes dramatically. A single agentic task can use more than 3,400 times as many tokens as a typical chat interaction.
Source: Microsoft AI Diffusion Report
We are only in the early innings of agentic adoption, and the infrastructure is already strained. Power is setting the limits on what we can build and when. Packages and racks are growing larger and denser. Memory is becoming an even tighter constraint.
For years, the industry’s rational answer to each new requirement resulted in more: more silicon in the package, more memory beside it, more power to feed it and more fiber to connect it. Each generation delivered meaningful progress. But when each new gain requires more input than the one before it, we are on a treadmill. It moves only as long as we keep adding to it.
I believe we need to pursue two paths forward. The first is evolutionary: we continue improving the architectures we have today, driving incremental efficiency, utilization and economics within each generation. The second is transformational: changing the curve itself with innovation in new architectures, new materials and new approaches to system and model design.
The history of our industry is defined by transformations like these. When increasing CPU clock speeds ran into the power wall, we moved to multicore processors. When planar NAND reached its limits, memory went vertical.
And now, once again, we have an opportunity to challenge our assumptions and rethink the fundamentals. Because the next chapter of AI won’t be defined simply by how much infrastructure we build, it will be defined by how much intelligence we can create from it.
Engineering useful yield
For decades, the computing industry has optimized yield in the context of manufacturing. Today, that discipline has to extend across layers, from datacenters and silicon through models and the agentic harnesses that orchestrate them. And the work does not stop once the technology is built. We must then deploy and scale it faster, while developing new tools and systems to maximize utilization across our fleet.
From development through execution, each layer has a yield of its own, and losses and gains compound across them. Capacity at any one layer is only a starting point. The real measure is how effectively those layers work together to produce useful output from the system as a whole.
Our experience at Microsoft building and operating AI infrastructure at scale has reinforced two key lessons. First, the biggest constraints are rarely solved in the layer where they appear. Second, when we attack a constraint across the whole stack, tradeoffs that seemed inherent to the problem often turn out to be artifacts of the architecture.
The greatest advances often come when we apply these learnings through co-design, working across layers to turn apparent limits into solvable system constraints. Innovations in memory, networking and power show what this approach looks like in practice.
Memory: More intelligence from every byte
Today, memory is viewed as a supply problem or a component problem. In reality, it is a system problem.
In AI inference, memory is now setting the limits on system performance. It must hold larger models, preserve longer contexts and deliver data fast enough to keep the compute fed. And agents raise the bar even further. Generation, retrieval, tool use and persistent memory run together in loops that can last minutes or hours. The result is a much longer memory horizon, with far more information kept close to the compute and available across an expanding sequence of turns. Doing that efficiently at scale will define the next generation of AI infrastructure.
Our experience building the Azure Maia platform demonstrates that memory bottlenecks are not resolved by a single layer.
Model architecture, data science and compression can reduce the amount of KV cache, which stores the model’s working context during generation. Software can manage memory hierarchies more effectively, silicon can be optimized for data movement efficiency and compilers can place data closer to compute. No one change removes the constraint. Together, they increase the useful intelligence the system can deliver from the same memory resources.
That is useful yield: not simply adding bytes but getting more useful intelligence from every byte we already have.
Networking: Designing across layers
As we zoom out to the cluster level, we see that intelligence does not come from one chip. It comes from thousands of chips operating as one system. Faster links matter, but the productivity of the system also depends on congestion management, failure recovery, workload placement, programming complexity and the boundaries between silicon, system and software. Together, they determine whether expensive compute is producing intelligence or sitting idle.
When architecting the platform for Maia, we did not begin with an existing networking design. We began with the outcome we wanted to deliver: efficient inference at fleet scale, designing across silicon, networking and system software. Instead of separate scale-up and scale-out fabrics, we built a two-tier scale-up network, integrated the NIC functionality directly into the chip and developed a custom transport layer.
The result is scalable, consistent performance across dense inference clusters, with a unified fabric that simplifies programming, improves workload flexibility and makes better use of available capacity. And with less network hardware needed to deliver this performance, we also lowered the cost of running the entire system.
Our objective is not merely to move data faster. It is to keep more compute productive and deliver more tokens from every watt and every dollar.
Power: Co-designing for efficiency, from grid to chip
Moving from the cluster to the grid, AI has introduced new challenges around power availability, distribution and utilization. Racks have gone from tens of kilowatts to hundreds of kilowatts, and datacenter campuses can operate on the scale of gigawatts. Power used to be something the system simply plugged into. Now, it is something we design around, from the grid to the chip.
That is why the industry is rethinking power across the system. Solid-state transformers and 800-volt direct current power delivery can reduce distribution losses as power moves through infrastructure. Power and cooling are no longer downstream of the design, they are part of the product definition from the start. And increasingly, that co-design is needed all the way into the silicon.
Azure Cobalt 200, our Arm-based server CPU, shows what this looks like in practice. We designed Cobalt so that every core has its own voltage and frequency controls, paired with software-based, per-virtual-machine power capping. This finer-grained control enables targeted power adjustments while protecting the performance of critical workloads, allowing us to run more servers within the same power envelope.
As Cobalt demonstrates, hardware-software co-design enables us to more effectively turn every megawatt into customer value.
The breakthroughs between boundaries
The pattern we see across memory, networking and power extends throughout the system: start with the useful output, then optimize the whole rather than any one layer. This first requires us to be precise about the output we are optimizing for and which design constraints are truly fixed. Are we optimizing for peak performance or sustained system throughput? Would the workload benefit from significantly more capacity with marginally less redundancy? What creates more value: a broader set of capabilities or significantly earlier customer deployment?
Not every constraint in today’s systems is a law of physics. Some are inherited from decisions made elsewhere in the system and can change only when we work across traditional boundaries. Evolution comes from the steady gains each company drives within its own domain, but transformation comes when we challenge those assumptions together and redesign the system as a whole.
The breakthroughs ahead will emerge from collaboration across the ecosystem, spanning hyperscalers and silicon providers, equipment makers and materials innovators, utilities and datacenter operators, model builders and software developers.
Toward full yield
But tokens and intelligence are not the finish line. What we produce becomes the input for someone else’s work.
What matters next is how broadly that input translates into productivity across the economy and value in people’s lives, whether it helps a scientist accelerate discovery, a clinician identify a signal earlier, a student get help at the right time or a small business find a new path to growth.
That happens as AI becomes part of everyday work across industries and around the world. People build on it, new uses emerge, the tools improve and the value compounds.
For that cycle to spread, AI must be broadly accessible. And at scale, accessibility depends on efficiency. It is the only way to deploy enough intelligence, and at a cost that allows it to reach everyone.
When every person and every company can access intelligence, build on it and create value of their own, that is full yield.
We will continue to build capacity because the world will need it. But our defining measure of progress must be what comes out: not only chips or tokens, but useful intelligence translated into empowerment, opportunity and human achievement.
That is the yield imperative. And it is work our entire industry must take on together.
Rani Borkar leads the core organizations responsible for planning, architecting, developing and deploying hardware and infrastructure for Microsoft’s leading cloud computing platform — from silicon, to systems, to supply chain.
Learn more about Microsoft’s silicon to systems approach to Azure infrastructure.
JavaScript's native `Date` object internally tracks UTC instants but defaults to local time for display, leading to common timezone pitfalls. In this post, I describe how to manage and display timezones using various JavaScript APIs old and new from the old and trusty `Date` object, `INTL.DateTimeFormat()` and the newish `Temporal` class.
All Azure Key Vault control plane (management) API versions released before 2026-02-01 will stop working on February 27, 2027.
This does NOT affect the data plane (getting/setting secrets, keys, and certificates) — only vault and access-control management operations are impacted.
TL;DR — What you need to do
Upgrade Azure CLI to 2.90.0 or later (run az upgrade).
Upgrade Az PowerShell's Az.KeyVault module to 6.7.0 or later, or upgrade the Az package to 16.3.0 or later.
Do this before February 27, 2027 to avoid any interruption to vault management operations.
No data plane changes are required — reading/writing secrets, keys, and certificates is unaffected.
Current releases of Azure CLI 2.90.0 or later and Az PowerShell 16.3.0 or later (including Az.KeyVault 6.7.0 or later) already use Key Vault control plane API version 2026-02-01. Upgrade to these versions before February 27, 2027; after that date, earlier control plane API versions will no longer be served.
Who is affected
You are affected if you use either of the following:
Azure CLI scripts or automation that call az keyvault management commands (create, update, network-rule, access-policy, etc.).
Az PowerShell scripts that use the Az.KeyVault module's management cmdlets (New-AzKeyVault, Set-AzKeyVault, Update-AzKeyVault, etc.).
What breaks on and after February 27, 2027 if you have not upgraded
Key Vault management operations (create/update/delete vault, change network rules, change access policies or RBAC settings, etc.) issued with a pre-2026-02-01 API version will fail.
You’ll get the following error message: “Due to an RBAC security default, Key Vault's Resource Manager API versions older than 2026-02-01 will retire on 2027-04-01.”
Your vaults are NOT deleted and existing secrets/keys/certificates remain fully accessible via the data plane — only management (control plane) operations are blocked.
There is no way to opt out or request an extension once the retirement takes effect on February 27, 2027.
How to check your installed versions
Check your Azure CLI version
Run the following command:
az version
Check the azure-cli value. If the installed version is earlier than 2.90.0, upgrade Azure CLI before February 27, 2027 by running az upgrade.
Check your Az PowerShell version
Run the following commands:
Get-InstalledModule Az, Az.KeyVault | Select-Object Name, Version
If Az.KeyVault is earlier than 6.7.0, or the installed Az package is earlier than 16.3.0, upgrade before February 27, 2027. To upgrade the complete Az package, run:
Update-Module Az -Force
Alternatively, to update only the Az.KeyVault module to the latest version, run:
Update-Module Az.KeyVault -Force
Timeline
Date
Milestone
March 2026
Control plane API version 2026-02-01 generally available across public Azure regions, Azure Government, and Azure operated by 21Vianet.
Today → February 27, 2027
Preparation window. Upgrade to Azure CLI 2.90.0 or later, Az.KeyVault 6.7.0 or later, or Az 16.3.0 or later before February 27, 2027.
February 27, 2027
All control plane API versions released before 2026-02-01 are retired and stop being served.
Questions?
For any questions about Azure CLI or Azure PowerShell, please submit feedback through GitHub:
In today’s episode, Cassidy, Marlene, and GPS tackle the ever-growing pile of new AI and agent terminology that's popped up seemingly overnight. They break down loop engineering (designing systems that prompt your agents instead of prompting them yourself), the Ralph loop that helped popularize it, and how squads and fleets let multiple agents split up planning, testing, and review work in parallel. They also dig into harness engineering (the scaffolding around a model that turns raw text generation into something useful), the newly coined “hill climbing,” and why forward deployed engineer might just be a solutions architect with better AI tooling. They wrap with a look at the spectrum between closed, open weight, and open source models, plus each host's open source pick of the week: Astro 7, the RAE API project, and Tau.
A researcher at a pharmaceutical company deploys a productivity assistant built on a well-regarded open-weight model. The base model is clean—it's been evaluated, its weights are public, its behavior is understood. A low-rank adaptation (LoRA) adapter is added to tune it for the lab's domain: it understands drug-discovery terminology, knows the team's workflows, and uses the right tools. The assistant is useful. Nobody looks too hard at the adapter file. Why would they?