01 / The wire
Recent briefings
-
September 22, 2026 · afternoon
Two frontier labs shipped within ninety minutes of each other and neither led with a capability claim, they led with cost per finished task, and the lever both of them pulled was the price of a cache read.
-
September 22, 2026 · morning
Capability keeps arriving free and openly licensed while the thing that actually decides whether you can use it has moved into the plumbing around the model, node counts, compaction, CI throughput, and the identifiers riding invisibly inside the output.
-
September 21, 2026 · afternoon
Every launch on today's board is gated by an ordinary software-supply fact rather than a model capability, a package platform, a retention line, a hand-written enum, a non-commercial license.
-
September 21, 2026 · morning
The agent stack spent the weekend arguing about where a decision gets computed, pushing it up into a cluster runtime or down onto a laptop, while the first shared test set for typed decision models showed most of them score below a baseline that reads nothing but answer length and formatting.
-
September 20, 2026 · afternoon
Four of today's stories turn on the same gap, that deleting a thing and isolating a thing both take far more machinery than the button offering them implies.
-
September 20, 2026 · morning
Every performance number published about System One decision models this week came from a party selling the conclusion, and two of the primary artifacts say in their own text that the numbers do not transfer.
-
September 19, 2026 · afternoon
Four days after one lab shipped a model that only returns a decision, the ecosystem produced a browser clone, an open-weight rival with a priority claim, and a framework integration, and none of them has published a task-outcome comparison against the LLM it replaces.
-
September 19, 2026 · morning
Three separate groups attacked the token bill of long-horizon agents inside 48 hours, each at a different layer of the stack, and not one of them claims the cheaper output is more correct.
02 / Under the surface
Latest analysis
-
Claude Opus 5.5 and GPT-6 Sol Both Cut the Same Price, and It Wasn't the Headline One
Both labs cut cache reads harder than they cut input and output tokens and shipped cache measurement tooling alongside, which is a joint…
-
jev-chat-JARVIS Reads Your Screen by Pretending to Be a System Service, and Says So in the README
Screen reading through the accessibility API has become the universal integration layer for on-device agents, and because that API was…
-
Atlas Is Source Control for Coding Agents, and Its Best Idea Lives in a Gitignored Folder
Atlas records which agent session produced which commit, and that checkpoint record is the one artifact in the product that is neither…
-
Agent-Written Code Made CI the Bottleneck, and Per-Job Setup Is Where the Money Goes
When agents write most of the code, the dominant CI cost stops being test runtime and becomes per-job fixed setup, which is what decides…
-
The Typed Decision Leaderboard Put a Length-and-Formatting Baseline Above Eight of Thirteen Answer Verifiers
A verifier that cannot beat a surface-feature baseline on your own data is sorting answers by shape rather than correctness, and the AUC…
-
The Cheap AI Judge Has a Retention Bill: What LangSmith's Jev Integration Actually Turns On
Switching an eval judge for a cheaper one changes two retention defaults you did not choose, one inside LangSmith that moves every scored…
-
Kev Speaks a Paid API's Protocol From localhost, and Its Own Table Says Don't Run It on a Mac
Kev commoditizes a hosted decision API by cloning its request surface rather than its model, and the same Qwen3.5 upgrade that bought it…
-
Heretic Turns Alignment Removal Into a Pareto Front, and That Is the Uncomfortable Part
Heretic's real contribution is not that it strips refusals from open-weight models but that it converts the operation into a two-objective…
04 / Coverage map
Topics we track
Claude Code 60 OpenAI 25 Anthropic 20 Codex 17 Agent Skills 15 Hugging Face 13 DeepSeek Harness 12 Model Context Protocol 12 Kimi K3 11 MCP 9 LangChain 8 METR 8 Claude Code auto mode 7 GitHub Copilot 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 GPT-5.6-Cyber 6 Ollama 6 Anthropic Frontier Red Team 5 Claude Fable 5.1 5 Claude Opus 5 5 GLM-5.3 5 GPT-6 Astra 5 grok-build 5