Our team found significant issues with Opus 5, enough that we've rolled back to Opus 4.8. Here's what we identified.
- Short-term memory loss
- Path confusion
- Disregards simple imperatives
- Overly aggressive compaction
- Confidently wrong (the scariest one)
We haven't seen issues like these since Sonnet 4.x. Don't get us wrong: Opus 5 is a better model than Sonnet 4.x. It's intelligent. But intelligence alone does not make a better model, and that's why it's a clear regression from Opus 4.8.
In our line of business, assessing models quickly matters. Every day spent on the wrong model costs our customers time, energy, and money.
New releases regress sometimes. That is exactly why you need a system for catching it.
How we grade models
Two things anchor every assessment.
- Aligned benchmarks
- A standard harness
Aligned benchmarks
Benchmarks serve as an early warning system for detecting variances in models. Not the public leaderboards, but benchmarks aligned to our own systems and workflows, scored against our own baselines.
Every process that touches AI gets one or more benchmark tests. When a new model lands, we know quickly whether it alters our outputs, for better or worse.
More on how we build these in a later post.
A standard harness
Benchmarks don't catch everything. Your second line of defense is a standard harness: one uniform set of tools, skills, and prompt patterns that everyone on the team uses to address the model.
Hold the harness constant, and the model becomes the only variable.
When outputs drift, you know it was the model, not the phrasing, not the tooling, not the person driving. Whether we're planning a new feature or researching a concept, we address the model the same way, every time. Research and discovery, code review, or anything in between: same formula, every time.
If everyone on your team prompts with "whatever comes to mind," outcomes vary by person and by attempt. If everyone uses the same structure, patterns, and tools, consistent results follow, and they're comparable. Team members can compare notes because they're running the same experiment.
We can't stress enough what a difference this has made.
This is the standard harness from the article: the deep engineering skills our team trusts every day to roughly 7x our productivity with confidence. Same tools, same patterns, same structure, every model, every time. Install via the plugin marketplace and you'll get every future update too.

