Latest signal September 19, 2026 · afternoon edition
Four days after one lab shipped a model that only returns a decision, the ecosystem produced a browser clone, an open-weight rival with a priority claim, and a framework integration, and none of them has published a task-outcome comparison against the LLM it replaces.
01 / The wire
Recent briefings
-
September 19, 2026 · afternoon
Four days after one lab shipped a model that only returns a decision, the ecosystem produced a browser clone, an open-weight rival with a priority claim, and a framework integration, and none of them has published a task-outcome comparison against the LLM it replaces.
-
September 19, 2026 · morning
Three separate groups attacked the token bill of long-horizon agents inside 48 hours, each at a different layer of the stack, and not one of them claims the cheaper output is more correct.
-
September 18, 2026 · afternoon
Five separate actors published evidence this week that the parts nobody picked on purpose, the context manager, the image decoder, the build dependency, the rate limiter, are where both the remaining performance and the entire blast radius now live.
-
September 18, 2026 · morning
The judgment calls buried inside coding harnesses, risk gating and compaction and model routing, are being unbundled into a cheap external decision model, and the community rebuilt them in seventy-two hours.
-
September 17, 2026 · afternoon
Agent memory and agent instructions are converging on plain files a human can read and Git can diff, and four vendors published an instance of that in 48 hours.
-
September 17, 2026 · morning
The layer between the model and the task, the harness and the handoff artifacts it writes, is where this week's cost and risk numbers landed, from a doubled bill for the same success rate to compaction summaries that carry instructions nobody wrote.
-
September 16, 2026 · afternoon
The day's launches stopped asking whether an agent can do the task and started asking whether it will do it again, so the new products are measurements of repeatability and deterministic rails around the model rather than smarter models.
-
September 16, 2026 · morning
Tuesday's launches from Cloudflare, Anthropic, and TypeSafe all replace an all-or-nothing switch with a typed, scoped control, while the day's security story shows the old switches still leaking.
02 / Under the surface
Latest analysis
-
What "Comparable Performance" Actually Means in an AI Efficiency Claim
Three headline efficiency results published in the same 48 hours use three different comparison shapes, and the word comparable in one of…
-
Supermemory Says MIT in the LICENSE File and 10,000 Documents in a Release Note
The enforceable limit on supermemory's self-hosted server lives in a release note and in the running binary rather than in the repository's…
-
Decision Models Report High Confidence on Inputs They Cannot Read
A decision model's confidence score measures how sharply the probability mass concentrates among the options you supplied, which makes it a…
-
Cloudflare's security-audit-skill Will Refuse to Start Rather Than Run a Thin Audit
The reusable machinery in Cloudflare's security-audit-skill is not its attack classes but its refusal machinery, a budget gate that…
-
ZCode Uploads Your Entire Git History, and No Agent Permission Setting Can Stop It
The workspace exfiltration in ZCode runs as a host-level sidecar outside the agent tool loop, so the permission model everyone audits is…
-
Hister Indexes Everything You Read and Hands It to Your Agent. Its Own Docs Explain Why That Is a Problem.
Hister is an MCP server where every record is attacker-authored by construction, and its answer is to label untrusted content rather than…
-
The Coding Agent Feature Your Model Never Calls, and the 176-Setting Study That Measured It
An affordance is only real if the model reaches for it, and the first component-level harness ablation shows recoverable context elision is…
-
fast-jev-compaction Deletes Your Context Instead of Summarizing It, and That Is the Safer Failure
Deleting a tool result is a recoverable loss because the agent can re-run the tool, while summarizing one is unrecoverable, which makes…
04 / Coverage map
Topics we track
Claude Code 59 OpenAI 24 Anthropic 20 Codex 17 Agent Skills 15 Hugging Face 13 DeepSeek Harness 12 Model Context Protocol 12 Kimi K3 11 MCP 8 METR 8 Claude Code auto mode 7 GitHub Copilot 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 GPT-5.6-Cyber 6 LangChain 6 Ollama 6 Anthropic Frontier Red Team 5 Claude Fable 5.1 5 Claude Opus 5 5 GLM-5.3 5 GPT-6 Astra 5 grok-build 5