Microsoft MXC Gives Agents One Sandbox Call, and a Different Box on Every OS
MXC SDK 1.0 makes "run this untrusted code" a single function in Rust, .NET or Node. What it actually isolates depends on where it runs, and the SDK will tell you if you ask.
Every coding agent eventually writes a script and wants to run it. The usual answer is a Docker container someone set up once and nobody audits, or worse, the developer's own shell with a confirmation prompt everyone learned to click through.
Microsoft's answer arrived on October 7 as MXC SDK v1.0.0, and it was posted to Hacker News on October 9. MXC (Microsoft eXecution Container) is an MIT-licensed SDK for running untrusted code, the README names "model output, plugins, and tools," inside an isolated container on Windows, Linux or macOS. Rust, .NET and Node packages all expose the same V1 API.
Here is the part worth slowing down for. One API sits on top of very different machinery. The README says the backends range "from OS-native process sandboxes to full VMs." Same call, different box. Which box you get depends on which laptop or server your agent happens to run on.
What MXC is, in one snippet
The README's Node sample is short enough to quote whole:
import { spawn, type ContainerRequest } from '@microsoft/mxc-sdk/v1';
const request: ContainerRequest = {
command: 'node -e "console.log(\'hello from container\')"',
network: { egress: { default: 'deny' } },
timeoutMs: 30_000,
};
const child = await spawn(request);
That's the pitch. Describe the command, deny outbound network, set a timeout, run. The 1.0 release notes describe three execution modes (run to completion with captured output, spawn with pipes, spawn with a PTY), lifecycle calls for long-lived containers (provision, start, stop, deprovision), and APIs for backend discovery and filesystem policy. Packages are mxc-sdk on crates.io, Microsoft.Mxc.Sdk on NuGet and @microsoft/mxc-sdk on npm.
For an agent harness, that's a real step up. A sandboxed exec tool used to mean writing three integrations or giving up on two platforms.
The box changes with the host
Now the table that matters, straight from the README:
| Platform | Default backend | Other backends |
|---|---|---|
| Windows 11 (x64, ARM64) | processcontainer |
windows_sandbox*, wslc, microvm*, hyperlight*, isolation_session |
| Linux (x64, ARM64) | bubblewrap |
lxc, microvm, hyperlight |
| macOS (ARM64, x64) | seatbelt |
none listed |
* marked experimental
Three different defaults on three operating systems. Bubblewrap is a Linux namespace sandbox. Seatbelt is Apple's sandbox profile system. Windows process containers have their own model, with a support matrix pinned to specific Windows 11 builds (24H2 needs build 26100.9278 or later, for example). The VM-grade options, microvm and hyperlight, are experimental on Windows, and macOS has no alternative to Seatbelt at all.
The network story varies too. The README lists "Proxy support, outbound controls, and backend-dependent host filtering." That phrase, backend-dependent, is the one to remember. A policy that says "only talk to pypi.org" may be enforceable on one backend and not on another.
Microsoft isn't hiding this. The API reference says "Containment selects the backend" and "Native validation decides which policies the backend can enforce." The design assumes you'll ask. The risk is that the snippet above doesn't.
The SDK will tell you what you're getting
This is where MXC earns some trust. It ships the calls you need to stop guessing.
getAvailableBackends() returns each backend the host offers along with its isolation tier, capabilities and warnings, according to the Node API reference. The docs call host availability "advisory," so treat the list as a starting point. getPlatformSupport() reports whether the SDK can launch on this host at all.
Then there's a family of dry runs: validateProvision, validateStart, validateStop, validateDeprovision and validateProcess. Each one runs native validation without creating anything and returns warnings. On Windows, probe gives "request-specific compatibility diagnostics" without creating a container. The reference adds an honest caveat: validation "does not guarantee that a later operation will succeed on a changed host."
What the docs I read don't settle is the case you most care about: when a backend can't enforce a policy you asked for, does the request fail, or does the policy drop silently? The API reference documents a few specific rejections (Seatbelt PTY rejects guiAccess, for one) but not a general rule. Until you've tested it on your own hosts, assume nothing.
Two warnings in the README worth taking literally
The first is about audit mode. MXC can record what a workload touches and turn that into a starting policy, which is a smart way to write containment rules. The README is blunt about the cost: "--audit turns off all sandbox security for the workload being analyzed." Never point audit mode at code you don't trust. Run it on your own known-good tool, then apply the generated policy to the untrusted one.
The second is about friction. "Your application will hit access issues when running in a sandbox, until you've had time to tune your containment rules." That's the right expectation to set. The failure mode it invites is the developer who hits five access-denied errors in a row and widens the policy until the errors stop. MXC ships helpers like getUserProfilePolicy() (read-only access to standard user-profile app data) and getAvailableToolsPolicy() (tool and SDK directories found in the environment). They're convenient. Read-only access to profile app data is still read access to wherever your other tools keep their state, so add these on purpose, not as the first fix for an error.
Put this into practice
Here's the lowest-friction way to put MXC behind an agent's exec tool without fooling yourself about what it does.
1. Inventory the box on every host you'll run. Call getAvailableBackends() on your dev machine, your CI runner and whatever box hosts the agent in production. Log the isolation tier and warnings. If those three lists differ, your agent's safety differs across them too.
2. Choose the backend explicitly. Containment selects the backend. Set it in your request instead of inheriting the platform default, so a move from a Linux runner to a Mac doesn't silently change what "sandboxed" means.
3. Validate every policy you rely on. Run the matching validate* call (or probe on Windows) for each request shape your agent uses, and fail closed in your own code if validation returns a warning about network or filesystem rules.
4. Start from deny. Copy the sample's network: { egress: { default: 'deny' } } and a short timeoutMs, then add filesystem paths one at a time as access-denied errors show you what the workload really needs. The repo's docs/logging-access-denied.md covers the diagnostics.
5. Write a test that proves the box holds. Have the sandboxed command try to reach the network and read a file outside its allowed paths. Assert both fail. Run that test on every platform you ship to. It takes ten minutes and it's the only evidence that counts.
Honest limitations
MXC is a 1.0, and a young one. The release notes call v1.0.0 the first stable release, and they list breaking changes that move everything to the V1 entry points. Expect more churn in the backend-specific *Config types.
macOS gets one backend. If Seatbelt can't express a policy you need, there's no VM option to fall back to on a Mac today.
The strongest isolation is the least finished. On Windows the microVM and Hyperlight backends carry the experimental mark, and the default is a process-level container. For hostile code, a process sandbox and a VM are not the same promise, and MXC doesn't pretend otherwise.
Network filtering beyond plain deny depends on the backend. If your agent needs "deny everything except one registry," verify that rule on each platform rather than assuming the request field means the same thing everywhere.
And a note on reading the repo itself: when I checked, the rendered GitHub page served a stale view with no releases and preview-era wording that the current raw README no longer contains, while the raw README and release feed showed 1.0. Check the release feed directly before deciding what version you're looking at.
The decision is yours to make on purpose
MXC turns sandboxing from a project into a function call, and that's worth celebrating. It also makes it easy to believe you've made a security decision when the platform made it for you.
The SDK gives you getAvailableBackends(), the validation calls and an explicit containment choice. Use all three, write the escape test, and you'll know exactly what box your agent's code runs in on every machine you own. Skip them, and "sandboxed" means whatever the host happened to default to that day.
Sources: microsoft/mxc README, MXC releases, MXC API reference, MXC Node V1 API, processcontainer OS version support, Hacker News thread.
Medium metadata
- Title: Microsoft MXC Gives Agents One Sandbox Call, and a Different Box on Every OS
- Subtitle: MXC SDK 1.0 makes "run this untrusted code" a single function in Rust, .NET or Node. What it actually isolates depends on where it runs, and the SDK will tell you if you ask.
- Tags: AI Agents, Security, Microsoft, Sandboxing, Software Development
- Canonical URL: set to the fervorai.dev post on import