AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

An open-source benchmark called livenerf is running a pre-registered, 30-day daily test of Claude Opus 5.5 against a launch-week baseline to detect whether Anthropic quietly degrades the model after release. Six of 30 days are collected; the first verdict cannot come before roughly October 24, 2026, and validation shows the tool cannot detect a same-family model swap.

An open-source benchmark project called livenerf is six days into a pre-registered, 30-day daily testing campaign to determine whether Claude Opus 5.5, released by Anthropic on September 22, 2026, gets quietly worse after launch. As of September 29, 2026, the project reports 6 of 30 days collected with none missed — all baseline days so far — meaning no verdict on the model’s stability exists yet, and none is possible before roughly October 24, 2026.

The project, hosted on GitHub by the user ninjahawk, was designed to answer one narrow question: does a frontier model degrade after it ships? The project’s documentation states that for months there have been reports that Anthropic “nerfs” models days or weeks after release — through quantization, a smaller model behind the same name, reduced effort, or routing changes — but that these reports could equally be people “pattern-matching on noise.” Because no one had a clean day-0 baseline, the project argues, every prior argument amounted to “vibes versus vibes.”

The methodology is deliberately conservative. The panel consists of 78 questions drawn from 2,336 screened GPQA Diamond, MMLU-Pro, competition-math and AIME 2025–26 items. According to the project’s calibration notes, Opus 5.5 answers about 93% of screened questions correctly on the first try, and 97% of items were either always right or always wrong; only the 78 “sometimes right” questions have enough signal to detect drift. The harness runs on a pinned Claude Code CLI (version 2.1.280) with frozen prompts, and the statistical approach follows Anthropic’s own published guidance, “Adding Error Bars to Evals,” built on the UK AI Security Institute’s open-source Inspect framework.

Validation results published by the project quantify what the instrument can and cannot see. One daily run of the full panel can detect an accuracy change of about 7.5 percentage points per 10-day window. Deliberately lowering effort showed up strongly in token counts — “effort low” cut output tokens by 62% and cost 8.3 ± 4.5 points of accuracy. However, the project reports a key limitation: swapping in Opus 5 for Opus 5.5 was not distinguishable at 99% confidence (−3.8 ± 6.3 points), meaning the tool cannot reliably detect a same-family model swap.

At a glance
reportWhen: ongoing; progress reported 2026-09-29;…
The developmentThe livenerf project reported on 2026-09-29 that it has collected 6 of 30 planned daily runs of Claude Opus 5.5, with no missed days, as it builds the first clean baseline needed to test post-launch ‘nerf’ claims.

Why a Day-0 Baseline Changes the Debate

The project matters because it converts a recurring, unverifiable community complaint into a measurable statistical question. If Anthropic or any frontier lab changed serving behavior post-launch — through quantization, routing, or reduced reasoning effort — users paying subscription rates would have no way to notice until aggregated evidence emerged. livenerf’s contribution is a baseline fixed on launch day, before any change could occur, with a pre-registered analysis plan that prevents after-the-fact cherry-picking.

The token-count metric is the project’s most sensitive instrument: its documentation notes that if a model “quietly starts thinking less,” output tokens drop before accuracy moves at all. The validation data supports this — effort reductions produced token changes of −62% and −26% against accuracy changes of roughly 8 and 4 points.

The stated limits are equally consequential. Because a same-family swap (Opus 5 for Opus 5.5) was not detectable in validation-sized samples, even a clean 30-day result will not prove nothing happened — only that nothing happened above the benchmark’s detection threshold. The project itself frames the result table as reporting “improvements just as loudly as regressions.”

Amazon

Top picks for "livenerf opus nerf"

As an affiliate, we earn on qualifying purchases.

The Nerf Rumors and Prior Evidence

: “

Community claims that Anthropic degrades models after release have circulated for months, according to the project’s README, but lacked a controlled baseline to test against. livenerf began collecting data on 2026-09-24 at 22:10 UTC, roughly 2.5 days after Opus 5.5’s launch on September 22. The schedule is 30 daily runs: days 1–10 form the baseline, followed by two 10-day comparison windows, with the first possible verdict around October 24, 2026 and the first results row published after day 20.

The panel and protocol were chosen, locked and validated under a pre-registered protocol (v2) before collection began. A report-only audit of the 78 panel questions found 8 answer keys that look wrong and 30 ambiguous questions — which the project says is expected for questions a strong model only sometimes answers correctly. Nothing was dropped; a pre-registered sensitivity analysis reruns results excluding them. The project also documents one deviation: day 5 ran with the budget guard overridden once, recorded in a deviations log. The benchmark runs on a personal Claude Max subscription rather than an API key, and rejects and excludes samples touched by the safety classifier.

“Nobody has had a clean day-0 baseline to check against, so every argument ends up as vibes versus vibes.”

— livenerf project README (GitHub, ninjahawk/livenerf)

No Verdict Until Late October

No finding about Opus 5.5 exists yet. All six collected days are baseline days (days 1–10 of 30), and the primary comparison only begins after the baseline window closes. The first results row appears after day 20, and the first formal decision is expected around October 24, 2026.

It also remains unclear whether the benchmark’s sensitivity is sufficient for the most commonly rumored scenario: the validation data shows a same-family model swap was not detectable at 99% confidence, and while a 10-day window contains about 2.5 times as many samples, the project states this “hasn’t been shown to be enough.” Additionally, the run depends on a single Max subscription and a pinned CLI; Anthropic could change the serving path or Claude Code in ways that interrupt the series, and the safety-classifier exclusions introduce another moving part.

The Road to October 24

The next milestones are mechanical: the project must complete baseline days 7–10 (through roughly October 4, 2026), then run the first 10-day comparison window, with the first results table row landing after day 20 (around mid-October). The first formal nerf-or-no-nerf decision under the pre-registered protocol is expected around October 24, 2026. The repo will maintain a running 10-day table reporting score deltas, standard errors, output token medians and control deltas versus the launch-week baseline, and a sensitivity analysis excluding the flagged questions will accompany the main result. Readers can follow the results in the GitHub repository’s results documentation.

Key Questions

Has Claude Opus 5.5 been nerfed according to livenerf?

There is no answer yet. As of September 29, 2026, the project has collected only baseline days (6 of 10). The first possible verdict is expected around October 24, 2026.

How would livenerf detect a nerf?

By comparing daily scores on a fixed 78-question panel against the launch-week baseline, using paired per-item differences with clustered standard errors. A secondary signal — median output token count — is expected to reveal reduced reasoning effort before accuracy drops.

What changes can the benchmark actually detect?

Per the project’s validation, one daily run can detect an accuracy change of about 7.5 points per 10-day window, and effort reductions clearly appear in token counts. A same-family model swap (Opus 5 for Opus 5.5) was not distinguishable at 99% confidence in validation.

Who runs livenerf and is it affiliated with Anthropic?

It is an independent, open-source project on GitHub under the user ninjahawk, running on a personal Claude Max subscription. It is not affiliated with Anthropic, though its statistical methods follow Anthropic’s published eval guidance.

When will the first results be published?

The first results table row appears after day 20 of collection, around mid-October 2026, and the first formal decision under the pre-registered protocol is expected around October 24, 2026.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Anthropic’s New Auto Mode For Claude: What It Means For AI Enthusiasts

Anthropic announces auto mode will become Claude’s default starting August 14, but details on operation and affected products remain unclear.

AWS Continuum Integrates OpenAI Codex And Anthropic Claude To Push AI Security Boundaries

AWS Continuum has integrated with OpenAI Codex and Anthropic Claude to enhance AI coding security, though technical details and deployment specifics remain undisclosed.

Build More Engaging Voice Interactions Using GPT‑Live‑1 Technology

OpenAI announces GPT-Live-1, a new API model enabling developers to create more natural, real-time voice experiences in applications and devices.

GPT-6 Astra On OpenRouter

OpenRouter integrates GPT-6 Astra, marking a significant step in open AI deployment. Details remain limited, but interest is surging.