01 / The wire
Recent briefings
-
September 16, 2026 · afternoon
The day's launches stopped asking whether an agent can do the task and started asking whether it will do it again, so the new products are measurements of repeatability and deterministic rails around the model rather than smarter models.
-
September 16, 2026 · morning
Tuesday's launches from Cloudflare, Anthropic, and TypeSafe all replace an all-or-nothing switch with a typed, scoped control, while the day's security story shows the old switches still leaking.
-
September 15, 2026 · afternoon
The same week the labs escalated the story that agents are becoming dangerous threat actors, three independent outside reads pushed back, and the gap between the catastrophe framing and the agents you can actually observe got wide enough to see through.
-
September 15, 2026 · morning
Supervision of agents is leaving the prompt and becoming a separate runtime component with its own veto, whether that is a per-command network allowlist, a guard model trained on execution events, or a validating agent that is never the one that found the bug.
-
September 14, 2026 · afternoon
Four desktop and OS vendors shipped in five days, and every one of them kept the harness and made the model the swappable part.
-
September 14, 2026 · morning
The product shipping this weekend was the workspace around the model, its files, memory, tools and provenance, and not the model itself.
-
September 13, 2026 · afternoon
Every argument that mattered this weekend arrived at the same place, that a rule written in prose is not a boundary, and the only things holding were enforced by a runtime.
-
September 13, 2026 · morning
The measuring instruments went under the microscope this weekend, because a benchmark, a safety harness and a scoring program are all things an optimizing agent can read.
02 / Under the surface
Latest analysis
-
TypeSafe Jev Is a Frontier Model That Cannot Write a Sentence, and That Is the Point
Most decisions inside an agent loop are enums, and a model that returns typed probabilities instead of text turns the hallucinated tool…
-
An AI Agent Found a Live Admin Token in a 2023 Docker Image in 25 Minutes. Yours Is Probably Still There.
This week's scoped-token launches govern credentials minted from now on, but the credential that gets you was baked into a container's…
-
Open Code Review: Alibaba Put the Parts That Must Not Fail in Code, Not in the Prompt
Open Code Review's value is the deterministic layer that fences the LLM into writing comments, and its self-run benchmark, lower recall by…
-
The Consistency Gap: Your Agent's 77% Is Hiding a 24-Point Reliability Problem
An agent's average pass rate hides a consistency gap caused by near-tied token decisions flipping under platform noise, so the metric to…
-
dbt Charts Puts the Dashboard in the Pull Request, and Keeps the Language Behind a Mirror
Dbt Charts matters because it makes an agent's dashboard a file that CI can fail, and the read-only mirror plus the Fivetran copyright line…
-
Claude Code 2.1.271 Makes Network Permission a Property of the Command, Not the Session
Claude Code's per-command alloweddomains turns a network approval from a session-wide grant into a per-command review, and the transferable…
-
Busbar Puts a Governed Boundary Between Your Agents and Everything They Can Reach
A single enforcement point in front of model calls, MCP tools, and agent-to-agent delegation is the concrete form of this month's…
-
Your Coding Agent Gets Worse the Longer It Runs, and a New Paper Measures Exactly When
Coding agents convert tokens into quality faster than random search at first, then their marginal gains fall below it, so past a measurable…
04 / Coverage map
Topics we track
Claude Code 55 OpenAI 24 Anthropic 20 Codex 17 Agent Skills 15 Hugging Face 13 DeepSeek Harness 12 Model Context Protocol 12 Kimi K3 11 MCP 8 METR 8 Claude Code auto mode 7 GitHub Copilot 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 GPT-5.6-Cyber 6 Ollama 6 Anthropic Frontier Red Team 5 Claude Fable 5.1 5 Claude Opus 5 5 GLM-5.3 5 GPT-6 Astra 5 OpenAI Presence 5 alibaba/open-code-review 4