Latest signal September 8, 2026 · afternoon edition
Ten thousand agents can now be pointed at one problem, and the only thing that makes their output checkable is a formal certificate rather than the fleet that produced it.
01 / The wire
Recent briefings
-
September 8, 2026 · afternoon
Ten thousand agents can now be pointed at one problem, and the only thing that makes their output checkable is a formal certificate rather than the fleet that produced it.
-
September 8, 2026 · morning
The industry stopped arguing about whether agents work and started publishing what they cost, in dollars per researcher per day, in context tokens per skill, and in the gap between cheap findings and expensive fixes.
-
September 7, 2026 · afternoon
What an agent loads has become the thing worth managing, and the week's launches are almost all knobs on that inventory rather than new capability.
-
September 7, 2026 · morning
Last week's shipping was almost entirely about approval gates, machinery deciding what an agent may read and what it may finalize, and GitHub handed an agent the approval bit in the middle of it.
-
September 6, 2026 · afternoon
OpenAI spent Sunday publishing its own evidence that the layer watching AI work is falling behind the layer doing it, and two independent pieces from the same week describe the identical failure at human scale.
-
September 6, 2026 · morning
Agent capability work has moved from the model to the box the model runs in, and this week showed both halves of that shift at once, labs industrializing the manufacture of training environments while the sandboxes already in production kept failing to hold.
-
September 5, 2026 · morning
Three separate shippers landed systems this week whose load-bearing part is a checker that sits outside the model and that the model cannot talk its way past.
-
September 4, 2026 · afternoon
Four launches in four days all moved the same piece, the control point sitting between an agent and everything it can touch, and each one moved it somewhere different.
02 / Under the surface
Latest analysis
-
Claude Code's /skill-doctor Prices Your Skills in Context Tokens. The Price Is Not a Verdict.
A skill that never fires is usually a description problem rather than a useless skill, so the right response to a cheap unused skill is to…
-
ripwire Hands Coding Agents a Repo Map Instead of grep. Its Most Convincing Number Is the One That Got Worse.
Ripwire earns trust not with its 52x headline but by re-running its own head-to-head, publishing a corrected margin of 1.46x instead of the…
-
Telling Your Coding Agent to Use Property-Based Testing Probably Makes It Worse
Verification instructions in a system prompt only change outcomes when they move the agent off a specific default behavior, and describing…
-
AutoHedge Asks for Your Wallet Private Key, and Four Fields Tell You Whether to Give It
Whether an agent repo is safe to run is decided by its credential surface, its reversibility path, its maintenance recency and its…
-
sv-number/skills and the Agent Skills Supply Chain: What a SKILL.md Actually Installs
A SKILL.md installs a vendor's judgment about when to use its product directly into an agent's startup context, and sv-number/skills is the…
-
Over-Editing Is Why Your Coding Agent's Diffs Are Unreviewable
Edit fidelity is a quality axis separate from correctness, and a preservation instruction in the prompt moves it further than a larger…
-
npm Staged Publishing, Copilot PR Approvals, and the Rule That Decides Which Way the Gate Swings
The variable that decides whether an agent gets the approval bit is the reversibility of the action, not the competence of the agent, and…
-
AutoHarness Lets Claude Code Skills Die of Disuse, and Only the Ones It Wrote
AutoHarness bounds a skill library by adherence in live use rather than a benchmark score, which is the right signal, and its scope limit…
04 / Coverage map
Topics we track
Claude Code 46 OpenAI 22 Codex 17 Anthropic 16 Agent Skills 15 DeepSeek Harness 12 Hugging Face 12 Model Context Protocol 12 Kimi K3 9 MCP 8 Claude Code auto mode 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 METR 7 GPT-5.6-Cyber 6 Ollama 6 Anthropic Frontier Red Team 5 Claude Opus 5 5 GitHub Copilot 5 GLM-5.3 5 OpenAI Presence 5 Claude Code self-hosted environments 4 Claude Fable 5.1 4 Cloudflare Vulnerability Discovery and Remediation 4