Campfire's Rust Port Is 26x Faster Than Rails. Check What Its Tests Never Measured
Basecamp built a strict parity suite for eight versions of its chat app. The most useful part is the paragraph that admits what it skips.
A chat app has one job that users notice: when someone posts a message, everyone in the room sees it. Page load speed is nice. Search is nice. A message that never arrives is a bug report.
Now picture a version of that chat app that serves pages 26 times faster than the original, passes a strict suite of checks against the original, and, under heavy load in one outside test, delivered about 1% of its new-post notifications.
That is roughly where the Rust port of Campfire, 37signals' self-hosted chat app, stood this week. The speed is in Basecamp's own README. The delivery number comes from a test that Piotr Sarnacki cites in his October 11 essay. And the reason both can be true sits in plain sight in Basecamp's verification repo, which says: "An HTTP throughput result does not measure concurrent users or WebSocket capacity."
That one sentence is the lesson for anyone porting code with AI agents. A parity suite (a set of tests that runs the old and new versions side by side and checks they behave the same) proves the new version does what the old one did on the paths you checked. It says nothing about the paths you didn't. And agents are very good at making the checked paths fast.
What Basecamp built
Campfire is a Rails app: chat rooms, direct messages, search, file attachments, a bot API. Earlier this month, DHH posted on X about rewrites of it in other languages, and the main repo now links seven of them: Django, Laravel, Express, Elixir, Go, Rust and C.
Sarnacki says DHH had an AI agent write the Rust one, because DHH "can't stand reading or writing the Rust code himself." The repos themselves do not say who or what wrote each port. The Rust repo does ship an AGENTS.md, a rules file written for coding agents, so agents clearly work in it.
The main README, as it stands on the main branch on October 11, carries a throughput table measured with 16 concurrent clients on an AMD Ryzen AI MAX+ 395. (The last tagged release, v1.5.2 from October 7, carries a different, older table with Rails at 230 room pages per second, so these figures are moving.) For the room page:
- Rails: 4,101 requests per second
- Elixir: 5,350
- Rust: 106,494
- C: 137,524
The Rust port serves about 26 times more room pages per second than Rails.
The interesting artifact is the verification repo. Think of it as the exam every port has to pass. It runs the same checks against each implementation:
- Every HTTP response must match. The README says "every measured HTTP response must match its route contract," down to status, headers and the full decoded body.
- Every write must land. "Every acknowledged message write must match its exact persisted ID, body, room and search-index entry." If the app says it saved your message, the suite checks the database row, the room and the search index.
- Browser flows must work. Installing, live messages, editing, search, permissions and invitations, driven through Playwright (a tool that clicks through a real browser).
That is a serious suite. Most teams porting code with agents have nothing close to it.
The gap the suite names
Then the docs say what it does not prove.
The README: "An HTTP throughput result does not measure concurrent users or WebSocket capacity." WebSockets are the open connections that push new messages to everyone in a room without them refreshing.
The October 8 performance review goes further: HTTP throughput "establishes no concurrent-people count, WebSocket fanout, Internet performance, notification delivery." Fanout means one message being copied out to every connected reader.
So the suite checks that a posted message is saved correctly. It does not check how many of the people watching that room actually got it, at speed, under load.
That gap is where Sarnacki's numbers land. He links a load test by Zach Daniels that found "a 1% successful delivery rate in the Rust version under heavy load." In that closed-loop comparison (where each simulated client waits for a reply before sending again), per Sarnacki, "Rust delivered 1% of notifications out of 6-7k, whereas Elixir delivered 100% of notifications out of ~1.7k."
Notice the second half. Rust handled three to four times as many events and delivered almost none of them. Elixir handled fewer and delivered all of them. Requests per second rewards the first. Users reward the second.
Sarnacki then ran a constant-rate test at 100 posts per second, using k6's constant-arrival-rate mode, which sends requests at a fixed pace no matter how fast the server answers. Rust delivered about 14% of events with no HTTP errors. Elixir delivered about 60%, but timed out on about 23% of posts, hit a worst-case delivery delay near 180 seconds, and used 1.8GB of memory. Neither port looks good. They fail differently.
Why the Rust port drops messages
The cause is a design choice, not a typo.
Picture a mailroom with a shelf that holds 256 envelopes. When envelopes arrive faster than readers collect them, the shelf fills. The Rust port's choice is to drop slow readers rather than let the shelf grow. Sarnacki: "In Rust when the deliveries are lagging, clients get disconnected." The buffer, a Tokio broadcast channel (Rust's standard tool for sending one message to many listeners), was set to 256, according to his essay.
Sarnacki reports raising it to 16,384. Delivery at 100 posts per second jumped from about 14% to about 90%. Maximum latency went from 11 seconds to over 130.
So there is no free fix. A small buffer fails fast and drops people. A big buffer delivers more, late enough that a chat message arriving two minutes after it was sent is arguably worse. Sarnacki's view: "If you want a reliable system you should know when to fail."
The Rust port's own README lists other choices of the same kind, in plain words. Redis and Resque, the job queue the Rails app uses, "are replaced by in-process queues," and "queued pushes and webhooks are lost on a crash." For CSRF protection (the guard against other websites submitting forms as you), "Sec-Fetch-Site replaces tokens," which needs a browser that sends that header, Safari 16.4 or newer for HTTPS forms. Sarnacki describes this as dropping the CSRF token "so that caching is easier." The README frames it as a replacement. Both are true descriptions of a deliberate tradeoff.
None of these choices is crazy. Each one makes the app faster and simpler. Each one is also a requirement somebody decided without a test that would have pushed back.
The rule that invited it
The Rust repo's AGENTS.md explains the shape of the project in two sentences. The port "had to be indistinguishable from the Rails app," and "that parity is done." Then: "The app may now diverge from Rails where that makes it faster or better."
That line gives an agent permission to trade behavior for speed. It also says to "list every deliberate divergence under 'Known differences'" and that "performance changes need before-and-after measurements." Good rules. But "measurements" here means the shared benchmark, and the shared benchmark measures HTTP.
So the instruction file pointed the work at the number the suite could see. Sarnacki's essay names the general problem: "if your prompt is not specific enough, many decisions are a coin flip."
This is the insight worth taking home. An agent working against a parity suite will make everything the suite checks correct and everything it measures faster. Whatever the suite does not check is up for grabs, and "faster" is the tiebreaker you wrote down.
Put this into practice
If you are porting or rewriting anything with a coding agent, copy Basecamp's structure, then add the test they skipped.
1. Run the real thing first. The verification repo is public and MIT-licensed, like the Campfire repos it tests. Its README lists the prerequisites (Ruby with Minitest, Rust 1.98.1, Node 22.18+, SQLite CLI, FFmpeg, Git and Docker). It expects the implementation checkouts to sit beside it, with each app's production image built first. Then the flow:
npm ci
npx playwright install chromium
bin/check
bin/seed
cargo build --release --locked --manifest-path loadgen/Cargo.toml
bin/benchmark --apps rails,elixir,go,rust
Read its checks as a template for your own: match every response, audit every acknowledged write, click through the main flows.
2. Write down the outcome users feel. For a chat app, that is "every connected reader gets the message within N seconds." For a payments service, "every charge the API confirms exists exactly once." Write it as one sentence in your agent's instruction file, next to the rule about speed. If your AGENTS.md says "faster is fine," it should also say what faster is not allowed to cost.
3. Test that outcome at a fixed rate. Use an open-model load test, one that keeps sending at a set pace instead of waiting for the server. In k6, the posting side looks like this:
export const options = {
scenarios: {
posts: {
executor: 'constant-arrival-rate',
rate: 100,
timeUnit: '1s',
duration: '2m',
preAllocatedVUs: 50,
maxVUs: 200,
},
},
};
The maxVUs line matters: k6 can only hold the rate while it has virtual users free, and it reports dropped iterations when it runs out. Then count what subscribers actually received against what the server acknowledged. That ratio is the number Sarnacki's test found and the throughput table could not.
4. Decide how it should fail. Pick on purpose: drop slow readers, or deliver late. Put the choice and the limit (say, "disconnect clients more than 10 seconds behind") in the requirements, and test it.
5. Keep a known-differences list. The Rust port's README does this well. Every change an agent makes from the original goes in one list a human reads. If the list is empty after a big port, nobody looked.
Honest limitations
The 1% figure is not Basecamp's. It comes from a test Sarnacki attributes to Zach Daniels and links on X, and Sarnacki's own reruns came "with various fixes," with setups he does not fully specify (no test duration or subscriber count is given). Treat the exact percentages as one person's measurement, and the direction as the point.
Who wrote the Rust port is Sarnacki's account, not the repo's. The repos are silent on it.
Basecamp's throughput table comes from the maintainers' own hardware, with four cores per app. The performance review notes that the Elixir, Go, Rust and C figures reuse earlier published runs while Rails, Express, Laravel and Django were remeasured.
The verification suite is not wrong for skipping delivery. It says so openly, which is more than most benchmarks do. The problem only appears when someone reads the table without reading the caveats.
And the k6 snippet above covers only the sending side. Counting deliveries needs WebSocket subscribers of your own; the verification repo's loadgen cable profile opens sessions but, by its own docs, counts sessions, not verified people.
What to take from it
Basecamp did more than most teams ever will: eight implementations, one exam, and a written list of what the exam does not cover. Copy all of it.
Then read that list as your to-do list. The checks you have are the ones an agent will satisfy. The outcome you care about needs its own test, written down before the agent starts, or the agent will pick the answer for you.
Sources: basecamp/once-campfire README, main branch · basecamp/once-campfire-verification README · Performance review, October 8, 2026 · basecamp/once-campfire-rust README · once-campfire-rust AGENTS.md · Piotr Sarnacki, "I'm sorry, but you still have to think" (October 11, 2026) · Zach Daniels on X · k6 constant-arrival-rate docs. Repos reviewed on main as of October 11, 2026.
Medium metadata
Title: Campfire's Rust Port Is 26x Faster Than Rails. Check What Its Tests Never Measured
Alternates to test:
- AI Code Ports Aren't Correct Just Because Their Tests Pass
- The Fastest Campfire Delivered 1% of Its Messages
Subtitle options:
- Basecamp built a strict parity suite for eight versions of its chat app. The most useful part is the paragraph that admits what it skips. (picked)
- Agents make the checked paths fast. The unchecked ones are a coin flip.
- What a 26x speedup and a 1% delivery rate teach about porting code with AI.
Description: Basecamp's Campfire verification suite checks eight ports for parity. Its own docs say it skips live delivery, exactly where the fastest port fell down.
Tags:
- Software Engineering (porting and testing)
- AI Agents (agent-built ports)
- Rust (the port at the center)
- Programming (reach)
- Testing (the takeaway)
Article type: A (spike: a real repo with a clear payoff). Publish at 7 to 8 AM ET, let Medium distribute for 2 to 6 hours, send to subscribers only if it is performing.