Static

Livenerf: Has Opus 5.5 been nerfed yet?

First reported by Github ·

The signal ●○○○ Compiled by AI from Github and Hacker News
Why you might care

Your access to Claude Opus 5.5 may become more reliable or less reliable without prior notice.

What happened

The "livenerf" project on GitHub, developed by ninjahawk, aims to provide a consistent benchmark for detecting subtle degradations in the performance of large language models after their initial release. The project was initiated to address anecdotal reports of Anthropic's Claude Opus model being "nerfed" post-launch, a claim previously difficult to substantiate due to a lack of a stable baseline. Livenerf establishes this baseline by running a panel of carefully selected questions, which are designed to be "sometimes right" for a strong model, across thousands of samples. It standardizes factors like prompts, the command-line interface, and grading to ensure determinism, measuring statistical drift in accuracy and output token count over time. The benchmark's methodology is built on the UK AI Security Institute's Inspect framework, utilizing methods outlined in Anthropic's own research for error bar calculation.

What it means

The livenerf benchmark is designed to detect accuracy changes of approximately 7.5 percentage points over a 10-day window, with a 3.6% margin of error per week. It also monitors output token count as an early indicator of reduced model "thinking." However, the project acknowledges that it cannot reliably detect subtle changes, such as swapping one version of Opus for another very similar one, with a less than 99% confidence interval. This suggests that while users might notice significant performance drops, minor or incremental changes might go unnoticed by both the benchmark and users.

The benchmark's sensitivity analysis indicates that the selection of 'sometimes right' questions can introduce bias, potentially overestimating a model's stable performance. Additionally, issues with answer keys or ambiguity in the selected questions were identified, though the pre-registered protocol accounts for this. The project's reliance on a specific CLI version and the potential for safety classifiers to interfere with certain question types highlight the complexities in ensuring true model-agnostic evaluation. Future observations will reveal whether these factors significantly impact the observed performance drift of Claude Opus 5.5 or other frontier models.

AI-written summary. May contain errors.

Livenerf