Latest signal September 13, 2026 · afternoon edition
Every argument that mattered this weekend arrived at the same place, that a rule written in prose is not a boundary, and the only things holding were enforced by a runtime.
01 / The wire
Recent briefings
-
September 13, 2026 · afternoon
Every argument that mattered this weekend arrived at the same place, that a rule written in prose is not a boundary, and the only things holding were enforced by a runtime.
-
September 13, 2026 · morning
The measuring instruments went under the microscope this weekend, because a benchmark, a safety harness and a scoring program are all things an optimizing agent can read.
-
September 12, 2026 · afternoon
Every story worth reading this afternoon is about who keeps the record of what a company did, and in each case the only copy belonged to the party being judged.
-
September 12, 2026 · morning
Shared developer infrastructure has become the surface agents act on, and the people who operate that infrastructure keep finding out last.
-
September 11, 2026 · afternoon
Today's loudest AI stories are all outside parties re-running somebody else's headline number and publishing the smaller one.
-
September 11, 2026 · morning
The harness stopped being scaffolding you write and became the product vendors sell, which means the layer that decides how your agent behaves is now the layer you no longer control.
-
September 10, 2026 · afternoon
Cheap capability has retired every control that was secretly a bet on scarcity, and most of today's launches are replacements for one of those bets.
-
September 10, 2026 · morning
Every headline number this morning is a price, and in each case the party quoting it is the party with the most to gain from it sounding small.
02 / Under the surface
Latest analysis
-
Webcmd Promises to Cut Browser Agent Token Spend by 90%. Its Own Benchmark Says 0.09%.
Webcmd's README headline and Webcmd's own published benchmark describe two different products, and the benchmark is the more useful one…
-
Real-SWE Is a Coding Benchmark You Are Not Allowed to See, and That Is the Point
A coding-agent benchmark is only uncontaminated while it stays private, so the field now has to choose between numbers it can audit and…
-
i-have-adhd Has 43,000 Stars, One Markdown File, and Zero Enforcement
The skill works because it governs output form, which the model has no incentive to defend, and the same week's alignment research shows…
-
Homebrew 7.0.0 Shipped the Agent Security Fix Everyone Else Is Still Writing Prompts For
The fix for untrusted code inside an automated loop is replacing evaluated strings with declared calls over literal arguments, and Homebrew…
-
Ship an Attestation, Not a Benchmark: What litelm's README Does That a Pass Rate Cannot
A scope-limited attestation naming the commit range reviewed, the tests run, and the compatibility explicitly not claimed tells an adopter…
-
RubyGems, OpenAI's Agents, and the Missing Disclosure Channel
Package registries have no inbound channel for an operator to report that its agents caused an incident, so attribution is being done four…
-
hyperresearch Blocks a Hallucinated Quote From Shipping. Its Own Headline Number Ships Unchecked.
Hyperresearch enforces unusually hard verification gates on the reports it produces, and applies none of that machinery to the benchmark…
-
Google's Artemis Now Credits the Project It Copied. The Credit Arrived in a Commit, Not in the Issue Asking for It.
Artemis now carries the attribution Apache 2.0 conditions redistribution on, added within about a day of a public blog post and never…
04 / Coverage map
Topics we track
Claude Code 51 OpenAI 24 Anthropic 18 Codex 17 Agent Skills 15 Hugging Face 13 DeepSeek Harness 12 Model Context Protocol 12 Kimi K3 11 MCP 8 METR 8 Claude Code auto mode 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 GPT-5.6-Cyber 6 Ollama 6 Anthropic Frontier Red Team 5 Claude Fable 5.1 5 Claude Opus 5 5 GitHub Copilot 5 GLM-5.3 5 GPT-6 Astra 5 OpenAI Presence 5 Claude Code self-hosted environments 4