TLDR
Microsoft Execution Containers is an open source containment layer from Microsoft that reached general availability on October 7 with the sentence the agent-safety world has been waiting for: the policy lives outside the agent, and generated code cannot grant itself access. It enables IT to give an agent the repository and the compiler and deny it your SSH keys, on four escalating isolation tiers, with evidence-first policy authoring and SIEM-grade audit export through NVIDIA’s OpenShell. The difference between the announcement and a security boundary you can trust is one open bug: with PowerShell 7 on the PATH, the SDK’s own helper hands the sandbox a read of the whole C drive, and the fix is still unmerged.
Caption: the four MXC backends: only the process tier ships cross-platform today, and it is where Copilot, OpenShell and OpenClaw already run.
What shipped
Microsoft’s post opens with a two-bad-options problem: an agent either gets the user’s full authority or gets blocked. MXC answers it with a declaration. A developer lists what a workload may touch, covering files at read, read-write, or nothing, network destinations including whether loopback counts, the process command and environment, and whether the desktop exists. The OS enforces that boundary no matter what the model, the generated code, or a plugin decides. The post’s own example is a coding agent with access to a website repository: it may read the server configuration and may not touch it, and “the containment environment is designed to prevent the operation” if it tries.
Three operating modes make the policy writable in practice. Learning mode does the useful work: ungranted access is blocked and recorded into a JSON activity report, so the least-privilege policy gets authored from evidence of what the agent actually tried to touch and then flipped to Enforcement. Permissive mode inverts that for rollout, allowing and recording. GitHub’s post walks the result on a real workload: a daily automation that reads public issue metadata, runs in a sandboxed working directory, and writes a dashboard on a schedule, with the model choice (local MAI Code 1.1 Flash versus cloud) independent of the sandbox.
Caption: the policy loop: author from evidence (Learning), enforce (Enforcement), roll out (Permissive), with OpenShell adding formal verification and SIEM export on top.
NVIDIA’s layer: credentials the agent never sees
OpenShell is the piece that turns OS containment into an agent runtime, and its design answers the question MXC alone leaves open: how an agent calls an inference API without holding the API key. Credential providers inject secrets only into requests bound for approved endpoints, so the agent’s sandbox can route to a model and still never see the key. Policy changes pass formal verification before they apply, with risky new grants (a fresh host reachable with live credentials, a new API method) deferred for human review. The promise is stronger than Microsoft’s Learning mode: MXC records what the agent tried, and OpenShell proves what the change would allow before it ships. The audit trail exits as OCSF JSON with named support for AWS Security Lake, Splunk, and CrowdStrike, the language enterprise security teams already run.
The same October 7 event carried the memory numbers that make local agentic inference work on these machines. GitHub’s Copilot Auto orchestration (coming by month end under the HydraFusion name) routes tasks between local and cloud, and the local side is MAI Code 1.1 Flash at 53GB quantized: 70.8 percent on SWE-bench Verified, within 1.8 points of its own cloud bf16 build and 38.8 points above GPT-OSS-120B’s GGUF on the same benchmark, at a peak of 75.5GB with a 256k context on a machine whose Task Manager shows about 110GB visible memory out of the 128GB headline.
Caption: what 128GB becomes: 74.9GB dedicated plus 35.5GB shared on one physical LPDDR5X bank; the 53GB quantized model plus cache and peak live inside it, and score 70.8 on SWE-bench Verified.
The hole, and why it belongs in the lede
Issue #1455 was filed the day after GA. Microsoft’s own getAvailableToolsPolicy helper, which the README recommends as the baseline, adds the drive root to the sandbox’s read-only paths whenever PowerShell 7 is on the machine, documented behavior meant to make interactive PowerShell work. On BaseContainer, a drive-root path grants the whole subtree. A sandbox built exactly as documented reads everything the user can read, .aws credentials, .ssh keys, browser profiles, other projects’ source, and nothing in the logs or the SDK’s return values signals that read limits no longer apply. The maintainer-commented fix makes whole-drive read opt-in, and it was still unmerged two days later.
The bug matters more than its severity because of what it says about the design process. The Hacker News thread (197 points) carries a kernel-adjacent contributor’s read that the backends are an “awkward abstraction” over projects with different security models, a tester’s verdict that this is “an early tech preview” wearing a 1.0 version, and the useful observation that the actual opening happened earlier: Windows 11 25H2’s August cumulative update made AppContainers creatable without administrator privileges, and MXC is the policy layer on top. The policy model is sound, and the 1.0.0 release shipped with an unmerged fix for a helper that grants the whole drive.
What to watch
- The #1463 merge and whether the SDK’s result object ever surfaces the effective drive grant; FilesystemPolicyResult carrying only path lists is the design flaw underneath the bug.
- Copilot Auto shipping at month end: the orchestrator’s routing decisions become the highest-volume consumer of the decision-model category Microsoft launched the same week.
- Entra agent identity and Agent 365 policy controls landing on-device: containment without identity leaves no accountability trail.
- Whether Apple follows: Seatbelt is already inside the process tier, and a macOS-native containment product aimed at agents is the obvious countermove.
- The enterprise SIEM path: OCSF export plus formal policy verification is the pitch that gets MXC into banks and hospitals, and it is NVIDIA’s, not Microsoft’s, layer.
Sources: Windows Developer Blog - GitHub Command Line - HN thread - Issue #1455 - PR #1463 - June partner post - NVIDIA developer blog - OpenShell OCSF docs - NVIDIA OpenShell
Related on this site: AI agent liability: contracts, insurance, and the gap - Opus 5 Ultracode database wipe postmortem - Self-hosting AI: DGX Spark vs RTX vs Mac - Serverless alarm loop: the $10,811 Cloudflare bill
Discussion
Be the first to commentStart a discussion
Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.