01 / The wire
Recent briefings
-
September 21, 2026 · afternoon
Every launch on today's board is gated by an ordinary software-supply fact rather than a model capability, a package platform, a retention line, a hand-written enum, a non-commercial license.
-
September 21, 2026 · morning
The agent stack spent the weekend arguing about where a decision gets computed, pushing it up into a cluster runtime or down onto a laptop, while the first shared test set for typed decision models showed most of them score below a baseline that reads nothing but answer length and formatting.
-
September 20, 2026 · afternoon
Four of today's stories turn on the same gap, that deleting a thing and isolating a thing both take far more machinery than the button offering them implies.
-
September 20, 2026 · morning
Every performance number published about System One decision models this week came from a party selling the conclusion, and two of the primary artifacts say in their own text that the numbers do not transfer.
-
September 19, 2026 · afternoon
Four days after one lab shipped a model that only returns a decision, the ecosystem produced a browser clone, an open-weight rival with a priority claim, and a framework integration, and none of them has published a task-outcome comparison against the LLM it replaces.
-
September 19, 2026 · morning
Three separate groups attacked the token bill of long-horizon agents inside 48 hours, each at a different layer of the stack, and not one of them claims the cheaper output is more correct.
-
September 18, 2026 · afternoon
Five separate actors published evidence this week that the parts nobody picked on purpose, the context manager, the image decoder, the build dependency, the rate limiter, are where both the remaining performance and the entire blast radius now live.
-
September 18, 2026 · morning
The judgment calls buried inside coding harnesses, risk gating and compaction and model routing, are being unbundled into a cheap external decision model, and the community rebuilt them in seventy-two hours.
02 / Under the surface
Latest analysis
-
The Typed Decision Leaderboard Put a Length-and-Formatting Baseline Above Eight of Thirteen Answer Verifiers
A verifier that cannot beat a surface-feature baseline on your own data is sorting answers by shape rather than correctness, and the AUC…
-
The Cheap AI Judge Has a Retention Bill: What LangSmith's Jev Integration Actually Turns On
Switching an eval judge for a cheaper one changes two retention defaults you did not choose, one inside LangSmith that moves every scored…
-
Kev Speaks a Paid API's Protocol From localhost, and Its Own Table Says Don't Run It on a Mac
Kev commoditizes a hosted decision API by cloning its request surface rather than its model, and the same Qwen3.5 upgrade that bought it…
-
Heretic Turns Alignment Removal Into a Pareto Front, and That Is the Uncomfortable Part
Heretic's real contribution is not that it strips refusals from open-weight models but that it converts the operation into a two-objective…
-
OpenAI's Measurement Pixel Leaks Your Visitors Before Its Own Privacy Code Runs
A third-party SDK cannot decline to send a cookie it has already sent, because the browser attaches credentials to the script request that…
-
Jev as a Judge: The Most Repeatable LLM in LangChain's Benchmark Was Also the Least Accurate
Repeatability and accuracy move in opposite directions across the three LLM judges in LangChain's benchmark, so low variance is not a proxy…
-
The Agent Cache That Hit 5 Percent, and the Release Note That Said So
A cache keyed on exact strings is worth nothing in a system whose input is speech, and the only reason anyone knows the number is that the…
-
CUA-S1 Ships a Model Card That Disqualifies Its Own Launch Benchmark
The reusable artifact in Cua's computer-use release is the model card's release checklist and abstention-aware metric set, not the…
04 / Coverage map
Topics we track
Claude Code 60 OpenAI 25 Anthropic 20 Codex 17 Agent Skills 15 Hugging Face 13 DeepSeek Harness 12 Model Context Protocol 12 Kimi K3 11 MCP 9 LangChain 8 METR 8 Claude Code auto mode 7 GitHub Copilot 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 GPT-5.6-Cyber 6 Ollama 6 Anthropic Frontier Red Team 5 Claude Fable 5.1 5 Claude Opus 5 5 GLM-5.3 5 GPT-6 Astra 5 grok-build 5