The ‘they nerfed it’ argument never ends because nobody has a day-0 baseline, so it’s vibes versus vibes. What I like here is the discipline, not the question: frozen prompts, a pinned CLI, exact-match grading, no LLM judge, raw logs kept forever.
The line I’m stealing for my own Claude Code setup: a changed harness looks exactly like a changed model. Before I blame the model for a bad week, I should check what updated underneath it. And I’d watch output tokens, not just accuracy: in validation, low effort cut tokens 62% but accuracy only about 8 points. No verdict yet, and I’ll take that over vibes.
The story — livenerf is an append-only benchmark tracking whether Claude Opus 5.5 gets worse after launch. It runs a locked panel of 78 sometimes-right questions once a day for 30 days through headless Claude Code on a Max subscription, comparing two 10-day windows against a 10-day baseline. Seven days are collected; the first possible call is around 2026-10-24. Validation could not distinguish Opus 5 from Opus 5.5. (Source)