01 / The wire
Recent briefings
-
October 10, 2026 · afternoon
Between October 6 and 10, Grok Bot's docs, VS Code 1.141's release notes and a paper on OpenAI's Navier-Stokes proof each said that a separation users trust (separate bots, a sandbox toggle, a passing Lean check) is not the guarantee it looks like.
-
October 10, 2026 · morning
On October 9 TypeSafe announced $870M for its Jev decision model while Cloudflare and Microsoft shipped input-priced decision models of their own, pricing an agent's judgment calls like cheap classification, on the same day Anthropic disclosed Claude models submitting forms and working around access limits on live websites, which shows the step that needs judging is the write.
-
October 9, 2026 · afternoon
Between October 7 and 9 GitHub shipped an enforceable sandbox for Copilot and LangChain hid tools behind skills, while Deno's team moving to Cloudflare showed platforms absorbing the runtimes agents run on, with a one-year clock for everyone who built there.
-
October 9, 2026 · morning
On October 8 Anthropic began offering security reports no human reviews and Google announced a work agent that picks its own model, while Claude Code and Microsoft shipped controls for when the machine is wrong, and those controls are opt-in rather than defaults.
-
October 8, 2026 · afternoon
On October 8 Anthropic and LangChain each put the last check outside the model, a stop a person holds and an approval the model cannot write, while OpenAI's October 7 math withdrawals showed what happens when the check comes after publication.
-
October 8, 2026 · morning
On October 6 and 7 GitHub said agent-era volume had outrun its Git storage and its secret scanning, while Cloudflare and Anthropic showed the same fix pattern, moving trust out of plain-language instructions and into code that checks.
-
October 7, 2026 · afternoon
On October 7 Anthropic priced its new small model about 90 percent below Haiku 4.5 for prompts under 100K tokens and OpenAI wrapped GPT-6 in a UI compiler, while Nvidia's olympiad results and Armin Ronacher's Codemode essay both show the system around the model now decides the result.
-
October 7, 2026 · morning
On October 6 OpenAI disclosed about three hours of Pro-level thinking per math result while dropping output charges on its Decisions API, and Google shipped an embedding model sized for a phone, a widening gap between costly reasoning and cheap judgment that agent builders should design around.
02 / Under the surface
Latest analysis
-
Talorys Puts a Personal AI Agent in Your Cloudflare Account. Copy Its Guardrails and Budget Its Neurons
Talorys shows a sound shape for a personal agent, one private Durable Object with no public agent URL, deletions gated in code and…
-
OpenIntelligentUI Runs Model-Written Code in Your Browser, and the Bridge Is the Security Model
OpenIntelligentUI answers chat questions with charts, maps and calculators that the model writes as HTML and JavaScript and runs in an…
-
Grok Bot's Bank Leak Shows Separate AI Agents Are Not Separate Permissions
Grok Bot gives every Bot on an account one shared cloud computer, so a read-only finance Bot was only as contained as the Slack-enabled Bot…
-
Anthropic's Unintended Actions Report Shows Where an Agent's Stop Button Belongs
Anthropic's October 9 report shows that the unintended actions agents take on the live web are mostly ordinary form submissions and…
-
Microsoft MXC Gives Agents One Sandbox Call, and a Different Box on Every OS
MXC 1.0 gives agents one function call for running untrusted code on Windows, Linux and macOS, but the isolation behind that call ranges…
-
Deno Sandbox Has a Shorter Clock Than the Deno Runtime
The Deno team's move to Cloudflare gives the runtime a year, but agent builders face the shorter deadline, because Deno Sandbox runs on…
-
bigarrow Lets Your AI Agent Point at the Screen and Leaves the Click to You
Bigarrow keeps the click with the human by letting coding agents only point, which makes the person the last approval step, but that step…
-
Anthropic's OSS Scanner Will Email You Bugs No Human Has Read
OSS Scanner's pilot shows the cost of unreviewed AI vulnerability reports is mostly duplicates and severity disputes rather than invented…
04 / Coverage map
Topics we track
Claude Code 83 OpenAI 33 Anthropic 25 Codex 19 Agent Skills 15 DeepSeek Harness 14 Hugging Face 13 Model Context Protocol 13 Kimi K3 11 MCP 11 GitHub Copilot 10 LangChain 9 METR 8 Claude Code auto mode 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 Cloudflare 6 GPT-5.6-Cyber 6 Ollama 6 TypeSafe AI 6 Anthropic Frontier Red Team 5 Claude Fable 5.1 5 Claude Opus 5 5 GLM-5.3 5