01 / The wire
Recent briefings
-
September 9, 2026 · afternoon
Today's launches all narrow what an agent is allowed to be, a named caller or a two-megabyte task instead of a general capability, while the day's biggest story is a lab pointing the same attribution machinery outward at people.
-
September 9, 2026 · morning
The most useful numbers published in the last 24 hours are the ones that name where a thing stops working, and the people publishing them are the ones who gain least from saying so.
-
September 8, 2026 · afternoon
Ten thousand agents can now be pointed at one problem, and the only thing that makes their output checkable is a formal certificate rather than the fleet that produced it.
-
September 8, 2026 · morning
The industry stopped arguing about whether agents work and started publishing what they cost, in dollars per researcher per day, in context tokens per skill, and in the gap between cheap findings and expensive fixes.
-
September 7, 2026 · afternoon
What an agent loads has become the thing worth managing, and the week's launches are almost all knobs on that inventory rather than new capability.
-
September 7, 2026 · morning
Last week's shipping was almost entirely about approval gates, machinery deciding what an agent may read and what it may finalize, and GitHub handed an agent the approval bit in the middle of it.
-
September 6, 2026 · afternoon
OpenAI spent Sunday publishing its own evidence that the layer watching AI work is falling behind the layer doing it, and two independent pieces from the same week describe the identical failure at human scale.
-
September 6, 2026 · morning
Agent capability work has moved from the model to the box the model runs in, and this week showed both halves of that shift at once, labs industrializing the manufacture of training environments while the sandboxes already in production kept failing to hold.
02 / Under the surface
Latest analysis
-
microsoft/tgrep Is 52x Faster Than ripgrep, and the 52x Is a macOS Number
Tgrep's headline speedup measures how slow the filesystem is rather than how good the index is, and the durable win for coding agents is…
-
Quantization Damage Is Nonlinear, and Qwen3.8 27B Shows Exactly Where the Cliff Is
Quantization damage is nonlinear rather than gradual, so the only defensible way to choose a quant is a task benchmark run against the file…
-
LangChain Connections Gives Agents Per-Caller Identity, and Turns a Missing Permission Into a Question
Per-caller credential resolution only becomes practical when a missing grant pauses the run and asks instead of throwing, which turns a…
-
deltafin Runs a 2.8-Trillion-Parameter Model on One MacBook, Then Publishes the Six-Minute Wait
The valuable result in deltafin's Kimi K3 run is not one token per second, it is the measurement showing that six-minute prefill is a…
-
Claude Code's /skill-doctor Prices Your Skills in Context Tokens. The Price Is Not a Verdict.
A skill that never fires is usually a description problem rather than a useless skill, so the right response to a cheap unused skill is to…
-
ripwire Hands Coding Agents a Repo Map Instead of grep. Its Most Convincing Number Is the One That Got Worse.
Ripwire earns trust not with its 52x headline but by re-running its own head-to-head, publishing a corrected margin of 1.46x instead of the…
-
Telling Your Coding Agent to Use Property-Based Testing Probably Makes It Worse
Verification instructions in a system prompt only change outcomes when they move the agent off a specific default behavior, and describing…
-
AutoHedge Asks for Your Wallet Private Key, and Four Fields Tell You Whether to Give It
Whether an agent repo is safe to run is decided by its credential surface, its reversibility path, its maintenance recency and its…
04 / Coverage map
Topics we track
Claude Code 47 OpenAI 22 Anthropic 17 Codex 17 Agent Skills 15 DeepSeek Harness 12 Hugging Face 12 Model Context Protocol 12 Kimi K3 11 MCP 8 Claude Code auto mode 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 METR 7 GPT-5.6-Cyber 6 Ollama 6 Anthropic Frontier Red Team 5 Claude Opus 5 5 GitHub Copilot 5 GLM-5.3 5 OpenAI Presence 5 Claude Code self-hosted environments 4 Claude Fable 5.1 4 Cloudflare Vulnerability Discovery and Remediation 4