Independent AI intelligence Two editions daily · ET
Fervor AI

AI Trending Briefing · October 11, 2026 · afternoon edition

Between October 8 and 11, an essay on a Rust port of Campfire its critic says an agent wrote and Hugging Face's ML Intern writeup showed agent output drifting wherever requirements went unwritten, while a paid README test showed humans still find gaps that a model-simulated user would not.

CampfirePiotr SarnackiHugging Face ML InternTerence TaoTerence EdenAi2agent-harnessmulti-agentfine-tuningai-skillslocal-ai

Trending AI Briefing: Sunday, October 11, 2026 (afternoon ET)

No lab shipped anything this Sunday, and the vendor blogs sat still. The interesting writing came from people checking agent output after the fact, and it kept finding the same thing. A Rust port of 37signals' Campfire chat app, which Piotr Sarnacki says DHH had an agent write, tops a throughput table and, by a load test Sarnacki cites, delivered about 1% of new-post notifications under heavy load. Hugging Face's ML Intern writeup shows a prompt growing from about 450 to about 2,000 words as each project exposed something the author forgot to ask for. Terence Eden paid humans to follow his README, turned down simulating them with an LLM, and came back with 12 highlighted problems. Two of the three pieces involve agents; all three say the same thing about specs. Whatever nobody wrote down gets decided anyway, and you find out how only when you test the thing you actually care about.

What's hottest in AI news right now

Piotr Sarnacki's "I'm sorry, but you still have to think," published October 11, takes apart an agent-written Rust rewrite of Campfire. Sarnacki says DHH had an AI agent write the port because he "can't stand reading or writing the Rust code himself." The Rust version beats Rails about 26 times over on Basecamp's own room-page throughput table (only the C port is faster), but Sarnacki points to a test, which he credits to Zach Daniels, that found "a 1% successful delivery rate in the Rust version under heavy load." In the closed-loop comparison he describes, the Elixir port delivered 100% of roughly 1.7k notifications. A separate constant-rate rerun at 100 POSTs per second had Rust delivering about 14% of events with no HTTP errors, and Elixir about 60%, with about 23% of POSTs timing out. Raising a Tokio broadcast buffer from 256 to 16,384 pushed Rust to about 90% delivery, but max latency jumped from 11 seconds to over 130. His line: "if your prompt is not specific enough, many decisions are a coin flip." The honest catch: the delivery numbers come from others' tests and his own reruns "with various fixes," with setups he does not fully specify. It hit about 173 points on Hacker News by mid-afternoon. Sarnacki · Zach Daniels on X

Basecamp's own Campfire repo now lists eight implementations and a shared verification suite, and its README spells out what the suite does not measure. The main-branch once-campfire README (latest release v1.5.2, October 7, carries an older, slower table) lists ports in Django, Laravel, Express, Elixir, Go, Rust and C, with a throughput table: the Rails room page serves 4,101 requests per second, Rust 106,494, C 137,524. The table was measured with 16 concurrent clients on an AMD Ryzen AI MAX+ 395. The verification repo checks HTTP responses, write audits and browser flows across ports, and says plainly that "an HTTP throughput result does not measure concurrent users or WebSocket capacity." Live delivery under load, the thing Sarnacki's test caught, sits outside both. The Rust port's README also complicates the "dropped CSRF" charge: it says "Sec-Fetch-Site replaces tokens," which needs Safari 16.4 or newer, and that "queued pushes and webhooks are lost on a crash." Neither README says who or what wrote the ports.

Hugging Face's "The model that didn't exist, so you made it yourself," posted October 8, walks through six projects built by its ML Intern agent in HuggingChat. ML Intern plans the work, asks for a budget, runs a small test job, then trains, evaluates and publishes on Hugging Face hardware. The six projects cost about $103 in compute combined: a citrus-disease vision model that went from 14.9% to 52.8% correct on 335 test photos for about $1.90, and a distilled four-step image model that reached a GenEval score of 0.536 against the 50-step teacher's 0.563 for about $37. The honest catch is in the post itself. One project took 48 jobs and another 59, some failing on missing packages and wrong paths, and the author's starting prompt grew from about 450 to about 2,000 words. It is three days old, included because nothing on the core beat shipped today. Hugging Face · ml-intern-prompts

Terence Tao's "Math 2.0" slides, dated October 2026, argue AI has made math the field where checkable work moves fastest. Proofs written in Lean (a proof assistant that checks every step mechanically) can be verified without trusting the author, human or model. Tao cites the Equational Theories Project settling over 22 million true/false statements in universal algebra, and says "we are now entering an era of proof abundance in mathematics." He also calls AI performance "extremely jagged: astounding in some directions, while inadequate in others," and asks that AI assessments fully report "resource consumption and negative results." The slides name no venue, though the file name says Caltech. It reached about 172 Hacker News points today. Tao slides (PDF)

Terence Eden's "I paid people to try and follow my README," published October 11, is the human version of the same test. He offered €25 for an hour of each tester's time, about €150 in total, for people to install his ActivityBot from the README over screen share while talking aloud. His list of highlights runs to 12 problems, from a broken demo link to unexplained terms. He anticipates the "just simulate users with an LLM" reply and turns it down: "I want to speak to real people." No metrics, only a list and a rewritten README. Around 316 Hacker News points. Eden

Ai2 replaced priority-based GPU scheduling with time budgets, it wrote on October 9. Demand ran two to three times above supply across about 150 researchers. Over a 30-day test, teams got 98% of the GPU hours they were owed, occupancy held at 98%, and p90 queue wait for debug jobs fell from about two hours to 30 seconds. No code was released. Ai2 on Hugging Face

New tools and features worth actually trying

ML Intern's budget cap. The post's most reusable trick is a single prompt line: "Cap total spend at USD 12 and ask me before exceeding it," plus asking for a baseline score and a smoke test before the full run. Copy that pattern into any agent that can spend money. Honest tradeoff: the cost figures cover only GPU and CPU job charges, and the post shows no installation steps beyond turning on ML-intern mode in HuggingChat.

Campfire's verification harness. bin/check, bin/seed, then bin/benchmark --apps rails,elixir,go,rust gives you a worked example of a parity suite for agent ports: every response must match its route contract and every acknowledged write must match what landed in the database. Honest tradeoff: it needs Playwright, Chromium, a Rust load generator and each app's production image, and by its own README it does not measure WebSocket capacity.

BetterWispr. A free, Apache-2.0 Mac dictation app: hold ⌥ Space, speak, and it types cleaned-up text into the focused app, transcribing on-device with Parakeet or Whisper models. v0.1.4 shipped October 9. Honest tradeoff: macOS 14 or later, and settings and history are stored unencrypted, per its repo.

Managed Deep Agents Slack reactions. LangChain's October 9 tutorial shows a reactions parameter on channels.slack(...) that picks an emoji per message, by rule or by a decision model with a confidence floor. Honest tradeoff: Slack only, needs Managed Deep Agents v0.9, and the model route runs through LangSmith's gateway. LangChain

Trending AI repos on GitHub today

Trendshift read once at about 15:10 ET; its figures are momentum scores, not star totals. Stars below are cache-busted shields.io counts.

  • maximhq/bifrost (#5, marked Featured): an AI gateway that puts many model providers (the README claims 23+) behind one OpenAI-compatible API, with failover and caching. Why now: Trendshift tags it Featured, which may be promoted placement. Apache-2.0, about 8.7k stars, transports/v2.2.6 on October 6. Caveat: its overhead figures (11 microseconds on a t3.xlarge, 59 on a t3.medium) are the vendor's own, and enterprise-tier features sit alongside open-source ones in the README.
  • storytold/photocraft (#3): a clean-room Photoshop reimplementation in Rust, native and WebAssembly. Why now: v0.7.0 tagged late Saturday ET (October 11 UTC). MIT or Apache-2.0, about 47k stars. Caveat: its README says it is "not yet a Photoshop replacement for daily professional work," and its compatibility numbers are self-reported.
  • cathrynlavery/diagram-design (#18): an agent skill that makes 44 kinds of editorial diagrams as self-contained HTML and SVG. Why now: back on the board after a week off. MIT, about 50k stars, no releases. Caveat: its README admits the same prompt "can lay out differently twice," and the output is not an editable source you can diff in a pull request.
  • multica-ai/andrej-karpathy-skills (#23): a single CLAUDE.md with four rules meant to stop coding agents from guessing, overbuilding and editing unrelated code. Why now: it is a written-down-requirements file, today's theme in one document. About 219k stars, no releases. Caveat: the README says MIT but the repo has no LICENSE file.
  • calesthio/OpenMontage (#25): an agentic video production pipeline driven by your coding assistant. About 66k stars, no releases. Caveat: AGPLv3, needs Python 3.10+, FFmpeg and Node 18+ (22+ for its HyperFrames engine); a zero-key path exists, but the full setup estimates about $1 to $3 per run against a default $10 budget cap.
  • opennookorg/betterwispr (Show HN, not on Trendshift): the on-device dictation app above. Apache-2.0, about 138 stars, v0.1.4 on October 9. Caveat: unencrypted local history.
  • basecamp/once-campfire (not on Trendshift): the Rails original behind Sarnacki's essay, with links to all seven ports and the verification suite. MIT (in a file named MIT-LICENSE), about 4.8k stars, v1.5.2 on October 7. Caveat: the throughput table was measured on the maintainers' hardware and does not cover live delivery.

What actually matters from today's signal

Track the gap between what an agent was asked and what it was checked against. The Campfire ports pass a strict parity suite and still diverge on the one behavior nobody wrote a check for. ML Intern works because it asks for a budget and a baseline before spending anything. For builders this week, the high-signal areas are: load tests that measure the user-visible outcome (delivered messages, not requests per second), spend caps written into agent prompts, written requirement files like CLAUDE.md that list what must not change, and real-human walkthroughs of anything an agent will later read as instructions.

The counter-signal is the throughput table itself. A 26-fold speedup over Rails is the kind of number that ends arguments, and Sarnacki's whole point is that it should not. Agents now produce working code faster than anyone writes the acceptance criteria for it, and the missing criteria get picked silently, by coin flip. Tao's slides make the same point from the other side: AI gets astonishing in exactly the places where a checker exists, and stays mediocre everywhere else. If your project has no checker for the thing users notice, an agent will not invent one for you.


Source access notes: Vendor scan at about 15:06 ET on October 11. OpenAI (latest October 8), Anthropic (latest October 8), Cloudflare (October 9), Mistral (October 6), Microsoft Foundry (October 7), LangChain (October 9) and Hugging Face (October 9) showed nothing dated October 10 or 11. GitHub's latest changelog entry (October 10, Copilot for JetBrains) was covered this morning. Claude Code npm dist-tags were unchanged (latest 2.1.296, stable 2.1.287). Google's AI blog index returned no dates. Codex changelog not attempted (JS-rendered). The huggingface.co/papers page served a stale cached list and was not used. Product Hunt search returned nothing current. A WSJ story on Anthropic's SpaceX compute deal was skipped as paywalled and as background to a deal reported earlier. DHH's posts announcing the rewrites and the Zach Daniels test date to October 4 (from X status IDs). Trendshift read once at about 15:10 ET; stars from cache-busted shields.io via a Sonnet 5.5 verification subagent. A Sonnet 5.5 adversarial pass then caught nine issues, all fixed: the agent-written claim stated as fact in the thesis (now attributed to Sarnacki), a thesis that counted Eden's human test as agent evidence, a 100% Elixir figure pinned to the wrong test, a diagram-design quote not in its README, an OpenMontage cost range stated too broadly, a missing bifrost hardware qualifier, "six models" (six projects), a Tao paraphrase, and Eden's session length. It could not verify Hacker News point counts, Trendshift ranks or the Zach Daniels X post. Two later corrections came from Part 2 and are folded in here: the main session caught that the Rust port does not lead the throughput table (the C port is faster), and Codex's article fact check caught that the 4,101 vs 106,494 table lives in the main-branch README, not in the v1.5.2 release (whose README shows Rails at 230 room pages per second); confirmed against the raw v1.5.2 README.