Is Your AI Model Getting Worse? livenerf Shows How to Actually Prove It
Everyone swears their model got dumber a week after launch. Almost nobody can prove it. A small open-source monitor that hit the Hacker News front page this week shows what proof would take, and why most complaints will never clear the bar.
The complaint is older than any single model. A frontier model launches, people spend a few days amazed, and then the threads start: it feels slower, it feels lazier, it used to one-shot this and now it argues with me. The vendor says nothing changed. The users are sure something did. Both sides have exactly the same evidence, which is to say a feeling and some screenshots.
A GitHub project called livenerf, which reached the Hacker News front page on September 29 and drew hundreds of points overnight, is interesting less for what it finds and more for what it refuses to accept as an answer. It is a monitor for one question: has Claude Opus 5.5 gotten worse since it launched on September 22? And its whole design is built around the fact that almost every way you would naively try to answer that question is wrong.
I want to walk through how it works, because the method is portable to any model your product depends on, and because understanding it tells you something uncomfortable about most of the "the model got nerfed" posts you have ever read.
The thing that makes this hard
Silent degradation is plausible. Serving a frontier model is expensive, and there are real levers a provider can pull that a user would never see: quantizing the weights to save memory, swapping in a smaller distilled variant under load, cutting the reasoning budget, routing some requests to cheaper hardware. Anthropic's own postmortem from September 2025 documented three separate infrastructure bugs that degraded real responses, including a routing error that sent some Claude Code requests to the wrong server type; roughly 30% of Claude Code users had at least one request hit by it over the period. So the concern is not paranoid. Models served over an API genuinely do change without announcement, sometimes by accident.
The problem is that human perception is a terrible instrument for detecting it. You get used to a model. The tasks you throw at it get harder as you trust it more. You remember the one brilliant response and forget the ten mediocre ones. And the model itself is nondeterministic, so the same prompt gives different answers on different days regardless of whether anything upstream changed. Every one of those effects points in the same direction: toward feeling like the model got worse, whether or not it did.
So a real monitor has to beat four things at once: your memory, your shifting expectations, the model's own randomness, and your motivation to find a regression you already believe is there. livenerf's design is basically a checklist of defenses against each.
How livenerf answers it
The README lays out the mechanics, and each choice maps to one of those traps.
It only uses questions the model is unsure about. The panel keeps items the model answers correctly somewhere between 30% and 70% of the time. That sounds backwards until you think about the statistics. A question the model always gets right, or always gets wrong, carries no information about small capability changes, because there is no room for the score to move. The questions in the middle band are the only ones sensitive enough to register a shift. The current panel is 78 such questions, drawn from GPQA Diamond, MMLU-Pro, and competition math.
It compares the model to itself, not to a memory. Each question runs repeatedly against a baseline built from the monitor's own first 10 days of collection. This is paired-differences analysis, and it is the same technique Anthropic recommended in its 2024 paper on evaluating models statistically: comparing scores on shared questions cancels out the noise from question difficulty, so you measure the change and nothing else.
It runs a control. A second arm tests Opus 5, the previous model, on GPQA items every day. If both arms drop at once, the cause is something platform-wide, a bad day on the servers, not a change to Opus 5.5 specifically. Without a control, every provider-side hiccup looks like your model getting nerfed.
It sets the bar before looking. A change only counts if the 99% confidence interval excludes zero across two consecutive 10-day windows, and the effect is at least 3 points, and the control arm stayed stable. Pre-registering the threshold is what stops you from finding a regression by squinting at noise, which is the failure mode of nearly every informal complaint.
Six days into a 30-day run, livenerf has no verdict, and it is honest that it cannot have one yet. The baseline period alone is 10 days.
What this tells you about the complaints
Here is the part that stings. Run the "my model got worse" discourse through this filter and almost none of it survives.
The typical complaint is one person, comparing today against a fuzzy memory of last week, on tasks that have changed, with no control for a bad server day, no baseline, no correction for the model's own variance, and a strong prior that the regression is real. That is not a measurement. It is the exact set of biases livenerf spends its entire design budget trying to cancel. This does not mean the complaints are wrong. It means they are unfalsifiable, which for a builder is almost as useless as being wrong, because you cannot make a decision on them.
The people arguing the other side on the Hacker News thread made a fair point too: benchmarks like this have existed for a while, and if large silent degradation were routine, one of them would have caught it by now. That is also evidence. The absence of a clean signal, from a method designed to find one, is worth more than a hundred anecdotes.
Put this into practice
You do not need to care about Opus 5.5 specifically. You need to care whether the model your product rides on is the same today as the one you tested. Here is the lowest-friction path to actually knowing.
Start with a calibrated panel, not a benchmark. Take 50 to 80 tasks that look like your real workload, run each 10 or 20 times against the model as you call it today, and keep only the ones landing in the 30 to 70% success band. Throw away the ones it always passes or always fails. This is an afternoon of work and it is the single highest-value step, because it is what makes the panel sensitive.
Freeze a baseline in the first days after you adopt a model. The window right after you start trusting a model in production is the only time you can capture what "good" looked like. Miss it and you have nothing to compare against later, which is why most teams can only argue about drift instead of measuring it.
Run it on the exact path you ship on. livenerf measures Opus 5.5 as served through Claude Code on a paid subscription, on purpose, because that is a different path from the raw API and could degrade independently. Test the endpoint, the tier, and the settings your users actually hit, not a clean lab call.
Add a control arm. Run an older, stable model through the same panel on the same schedule. When your numbers move, the control tells you instantly whether the cause is your model or the whole platform having a bad day.
Decide the threshold before you look. Write down what size drop, sustained over how long, would make you switch providers or open a ticket. Deciding after you see the data is how you talk yourself into or out of a regression that is really just noise.
Where this breaks
livenerf is honest about its limits, and any monitor you build will inherit them.
It cannot reliably tell apart two models in the same family. If a provider swaps Opus 5.5 for a slightly smaller Opus variant, the panel may not have the resolution to see it. A statistical monitor detects a change in the score distribution, not the mechanism behind it, so it tells you that something moved, never what.
It depends on the quality of your panel, and quality is hard. livenerf's own README admits an audit found 8 wrong answer keys and 30 ambiguous items, which were kept to preserve the pre-registered design. Bad items add noise and cost you sensitivity, and you will have them too.
It is slow by construction. The two-window requirement means a real regression takes 20 days to confirm. That is the price of not crying wolf, and it means a monitor is a way to hold a provider accountable over time, not an alarm that fires the afternoon something changes.
And it costs money and tokens to run daily, which for a subscription-metered path is a real constraint.
None of that makes the method optional. It makes it the floor. The alternative is not a faster answer, it is no answer, which is where the entire "is it nerfed" argument has been stuck for two years. A model you depend on deserves better than a feeling, and the tools to replace the feeling with a number are now sitting in a public repo you can read in an evening. The question worth asking is not whether your provider would ever change the model under you without saying so. It is whether you would be able to prove it if they did.
Sources: livenerf on GitHub, the Hacker News thread, Anthropic's postmortem of three infrastructure issues, and Anthropic's statistical approach to model evaluations.
Medium metadata
- Suggested tags: AI, Machine Learning, LLM, Benchmarking, Software Engineering
- Suggested subtitle: A small open-source monitor for Claude Opus 5.5 shows the minimum bar for proving a served model degraded, and why almost no complaint clears it.
- Canonical: publish on fervorai.dev first, import to Medium from that URL.