Latest signal September 10, 2026 · afternoon edition
Cheap capability has retired every control that was secretly a bet on scarcity, and most of today's launches are replacements for one of those bets.
01 / The wire
Recent briefings
-
September 10, 2026 · afternoon
Cheap capability has retired every control that was secretly a bet on scarcity, and most of today's launches are replacements for one of those bets.
-
September 10, 2026 · morning
Every headline number this morning is a price, and in each case the party quoting it is the party with the most to gain from it sounding small.
-
September 9, 2026 · afternoon
Today's launches all narrow what an agent is allowed to be, a named caller or a two-megabyte task instead of a general capability, while the day's biggest story is a lab pointing the same attribution machinery outward at people.
-
September 9, 2026 · morning
The most useful numbers published in the last 24 hours are the ones that name where a thing stops working, and the people publishing them are the ones who gain least from saying so.
-
September 8, 2026 · afternoon
Ten thousand agents can now be pointed at one problem, and the only thing that makes their output checkable is a formal certificate rather than the fleet that produced it.
-
September 8, 2026 · morning
The industry stopped arguing about whether agents work and started publishing what they cost, in dollars per researcher per day, in context tokens per skill, and in the gap between cheap findings and expensive fixes.
-
September 7, 2026 · afternoon
What an agent loads has become the thing worth managing, and the week's launches are almost all knobs on that inventory rather than new capability.
-
September 7, 2026 · morning
Last week's shipping was almost entirely about approval gates, machinery deciding what an agent may read and what it may finalize, and GitHub handed an agent the approval bit in the middle of it.
02 / Under the surface
Latest analysis
-
Terminal-Bench 2.1 vs 4.0: The Benchmark Version Number Is Now More Informative Than the Score
When a model beats a rival on one generation of a benchmark and loses to it by twenty points on the next, the gap is the training target…
-
OpenConnector Moves Provider Secrets Out of Your Agent, and the 1,490-Provider Catalog Is Not What You Self-Host
OpenConnector's real product is moving provider secrets out of the agent process and behind a runtime boundary you own, and the catalog…
-
openai/NavierStokesAndEuler: A Lean Certificate Proves the Logic and Leaves the Authorship Blank
A Lean certificate settles whether a proof term satisfies a formal statement and settles nothing about whether that statement is the…
-
Your Claude API Key Is Now Three Things to an Attacker: Loot, Compute, and Cover
Anthropic's threat report reclassifies a leaked model credential from a billing problem into an attribution problem, and the only control…
-
microsoft/tgrep Is 52x Faster Than ripgrep, and the 52x Is a macOS Number
Tgrep's headline speedup measures how slow the filesystem is rather than how good the index is, and the durable win for coding agents is…
-
Quantization Damage Is Nonlinear, and Qwen3.8 27B Shows Exactly Where the Cliff Is
Quantization damage is nonlinear rather than gradual, so the only defensible way to choose a quant is a task benchmark run against the file…
-
LangChain Connections Gives Agents Per-Caller Identity, and Turns a Missing Permission Into a Question
Per-caller credential resolution only becomes practical when a missing grant pauses the run and asks instead of throwing, which turns a…
-
deltafin Runs a 2.8-Trillion-Parameter Model on One MacBook, Then Publishes the Six-Minute Wait
The valuable result in deltafin's Kimi K3 run is not one token per second, it is the measurement showing that six-minute prefill is a…
04 / Coverage map
Topics we track
Claude Code 47 OpenAI 22 Anthropic 17 Codex 17 Agent Skills 15 DeepSeek Harness 12 Hugging Face 12 Model Context Protocol 12 Kimi K3 11 MCP 8 Claude Code auto mode 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 METR 7 GPT-5.6-Cyber 6 Ollama 6 Anthropic Frontier Red Team 5 Claude Opus 5 5 GitHub Copilot 5 GLM-5.3 5 OpenAI Presence 5 Claude Code self-hosted environments 4 Claude Fable 5.1 4 Cloudflare Vulnerability Discovery and Remediation 4