Opus 5 Is Fable 5 for Half the Money?
Prompt Engineering
4,540 views • 10 hours ago
Video Summary
Anthropic has released Opus 5, a new AI model that significantly outperforms previous versions on major benchmarks, including the Arc GAIT, and is priced at half the cost of comparable models like Fable 5. This release is notable for its advanced capabilities in agentic coding, demonstrated by its ability to generate a 3D model from a drawing without direct visual input, and its strong performance in front-end design and web development.
Despite its strengths, Anthropic intentionally reduced Opus 5's cybersecurity capabilities, potentially to expedite its release. The model's performance on third-party benchmarks like Cursor and Frontier Code is highly competitive, even surpassing some aspects of GPT-5. Early tests show impressive results in creative tasks such as generating a Pokémon encyclopedia website and an animation of text formed by a crowd, though some tasks, like the ISS tracker, were notably slow to complete on claude.ai.
Short Highlights
- Opus 5 significantly outperforms previous models on benchmarks like Arc GAIT.
- It offers Fable-level performance at half the price.
- Demonstrates advanced agentic coding and creative web design skills.
- Cybersecurity capabilities were intentionally reduced.
Key Details
Opus 5 Release and Benchmark Performance [00:00]
- Opus 5 is Anthropic's latest release, noted for being on a Friday and outperforming Fable 5 on major benchmarks.
- It achieves this performance at half the price of Fable 5.
- The model shows exceptional scores on third-party benchmarks like Arc GAIT, reaching 30%.
You can say that they really cooked with this model.
Agentic Coding Capabilities [01:01]
- Opus 5 excels in agentic coding, demonstrated by a task to rebuild a machine part from a drawing into a 3D model.
- The model developed its own computer vision pipeline to interpret the drawing without direct visual input.
- This showcases the model's agency and ability to handle complex tasks.
So, the model has agency.
Third-Party Benchmark Comparisons [01:31]
- Cursor Benchmark shows Opus 5 on max settings is close to Fable 5 on max settings at a lower price.
- On high settings, Opus 5 surpasses Fable 5.
- It also surpasses GPT-5 6 solve on extra high settings.
And if you look at high, it actually surpasses Fable 5.
Frontier Code and Thinking Budget [02:17]
- Frontier Code benchmark places Opus 5 on par with Claude 3 5.
- Increasing the thinking budget improves performance, though it increases cost.
- Opus 5 surpasses GPT-5 6 solve across different thinking settings, especially in agentic coding.
Specifically for agentic coding.
Availability and Pricing [03:20]
- Opus 5 is available on claude.ai and Claude console.
- Pricing remains the same as Opus 4 8, with performance improvements at the same cost.
- Reduced cybersecurity capabilities may have aided timely release.
They're keeping the same price level, but with every we are getting performance improvements and intelligence increase.
Token Usage and Performance [04:19]
- Reports of Opus 5 being token-hungry are not strongly supported by performance plots.
- For similar performance levels, it does not appear significantly more costly than Opus 4 8.
- Higher efforts and performance do increase cost, but hopefully not substantially for real-world tasks.
For the same level of performance, and actually or even higher level of performance, you are getting better pricing.
Creative Web Design Tests [05:33]
- The first prompt asked for an encyclopedia of the first 25 legendary Pokémon, resulting in a well-designed website.
- The output showed significant creative freedom and a nice aesthetic, surpassing typical card-like LLM outputs.
- Opus models are noted for their web design capabilities.
And it's a really, really nice-looking website. I'm kind of impressed.
Animation and Spatial Understanding Test [06:23]
- A second prompt involved creating an animation of a crowd forming text, requiring spatial tracking.
- The model successfully created the animation, forming "Hello World, I'm Opus."
- Later, it was updated to include camera controls, allowing zoom and angle adjustments.
Oh, wow. All right, awesome. This is pretty neat, right?
ISS Tracker Task and Performance [08:06]
- The ISS tracker task took over 50% of the time and was still working on claude.ai, indicating slowness.
- The final output was impressive, showing status levels, tracking, and a potential path.
- The detail and accuracy of the ISS tracker were highly praised.
This thing is really, really good.
Final Thoughts and Future Plans [11:00]
- The speaker concludes that Anthropic "really cooked" with Opus 5.
- A more detailed video with production code testing is planned.
- The current first impression suggests Opus 5 is a very impressive model.
So far, I am going to say that on topic really cooked with this model.
Other People Also See