Menu
ByteDance Seed Code (Fully Tested): ANTHROPIC is OFFICIALLY SCARED of this MODEL!

ByteDance Seed Code (Fully Tested): ANTHROPIC is OFFICIALLY SCARED of this MODEL!

AICodeKing

16,747 views 10 months ago Save 2 min 6 min read

Video Summary

A new Chinese AI model from ByteDance, integrated into the Trey code editor, is reportedly outperforming Claude and GPT-5 on coding benchmarks at a significantly lower cost. This performance has allegedly led Anthropic to revoke Trey's access to its models. The new model, D**ubau seed code, costs $17 and $12 per million tokens and supports image and video, challenging established players in the AI coding space. Despite some limitations in agentic benchmarks, its raw code generation and cost-effectiveness are highlighted as major advantages, suggesting a shift towards more affordable yet capable AI solutions.

An interesting fact is that this new Chinese model is estimated to be 15 times cheaper than its competitors, offering a compelling alternative for developers.

Short Highlights

  • A new AI model from ByteDance, powering the Trey code editor, reportedly outperforms Claude and GPT-5 in coding tasks.
  • This model is significantly cheaper, costing $17 and $12 per million tokens, approximately 15 times less than competitors.
  • Anthropic has allegedly cut off Trey's access to its models, possibly due to competitive pressure from the new Chinese model.
  • The model, identified as Dubau seed code, ranks well on the SWEBench verified leaderboard, even surpassing Anthropic's Sonnet.
  • It supports images and videos, with impressive performance on visual prompts, though agentic benchmarks show some inconsistencies.

Key Details

New Chinese AI Model Emerges [00:04]

  • A new AI model by ByteDance, the company behind TikTok, is making waves in the AI coding space.
  • This model is integrated into an AI code editor called Trey, which has been previously discussed.
  • It is reportedly outperforming established models like Claude and GPT-5 on coding benchmarks.
  • The new model achieves this performance at a fraction of the cost of its competitors.
  • One interesting fact is that this model is estimated to be 15 times cheaper than comparable models.

    "There's a new Chinese model by Bite Dance that is apparently beating Claude and GPT5 on coding for a fraction of the cost."

Anthropic's Response and Benchmark Performance [00:17]

  • Anthropic has reportedly cut off Trey's access to its models, a move described as "scummy."
  • This action is speculated to be a reaction to the competitive threat posed by the new Chinese model.
  • The SWEBench verified leaderboard is highlighted as a key indicator of AI model performance, which Anthropic actively monitors.
  • The combination of Trey and Dubau seed code is currently topping this benchmark.
  • This benchmark performance might be causing insecurity for Anthropic, potentially delaying their own model releases.
  • There is an approximately 8% difference observed between Claude's implementation and the new top-performing model.

Dubau Seed Code: Capabilities and Cost [02:23]

  • The model in question was launched publicly just two days prior to this video.
  • It has beaten Sonnet on SWEBench verified and other benchmarks.
  • This is not an open model and is primarily available in China, but it is noted for being extremely fast and cheap.
  • The cost is $17 and $12 per million tokens, making it incredibly affordable.
  • It also supports image and video inputs, adding significant versatility.
  • The low cost combined with advanced capabilities is seen as a major threat to Anthropic.

Accessing the Model [03:14]

  • The model is currently China-only and accessing its Volcano API requires a Chinese mobile number.
  • However, it is accessible through platforms like ZenMox, which functions similarly to OpenRouter and offers a variety of models.
  • ZenMox provides free credits to try out the Dubau seed code model.
  • It also offers an Anthropic-compatible API, allowing integration with tools like Claude Code.

Non-Agentic Benchmark Results [03:53]

  • On non-agentic benchmarks, the model showed mixed but promising results.
  • A floor plan generation was not perfect but produced correct code.
  • SVG generation for a panda and burger was recognizable and good, though the interaction was weak.
  • A 3JS Pokeball generation was fine, with accurate colors, though a button element was missing.
  • An autoplay chessboard failed to work, which was a disappointment.
  • However, Minecraft and Kandinsky style generations were described as "awesome" and better than Sonnet.
  • The depth of the map and random generation worked well, demonstrating impressive capability for a small, cheap model.
  • A butterfly flying in a garden animation was smooth and visually appealing.
  • A CLI tool in Rust worked fine, but a Blender script did not.
  • These results placed the model 15th on the leaderboard, which is considered good given its cost.

    "The raw code it writes is really good, but it's just the tool calling that seems to affect it negatively."

Agentic Benchmark Results and Limitations [05:50]

  • On agentic benchmarks, using Claude Code for testing revealed further limitations.
  • A movie tracker app and a God game failed to work properly, exhibiting bugs and errors.
  • However, a Go Tui calculator was a standout success, with excellent UI and functionality.
  • Other applications like Spelt app, Nux app, and an Open Code repo question also did not work.
  • This resulted in a 12th position on the agentic leaderboard, outperforming some tools but lagging behind others.
  • It is suggested that issues might stem from the model's interaction with Claude Code, specifically its use of terminal commands instead of edit diff tools.
  • The model is believed to be specifically trained for usage with Trey, and its tool-calling capabilities seem to be a weak point when used with other frameworks.
  • The raw code generated by the model is praised, but its integration with external tools is where it falters.

Future Potential and Market Impact [07:24]

  • There is hope that Trey will make this model generally available for broader testing.
  • Its strong performance on SWEBench Verified and its extremely low cost are significant advantages.
  • The model is described as "dirt cheap" with "super good capabilities."
  • The philosophy of achieving 80% of Sonnet's performance for a fraction of the cost is seen as a paradigm shift.
  • This approach is attributed to Chinese companies like Minimax, GLM, and this new model, indicating a trend towards democratizing AI.
  • The overall assessment is that this is a very cool and impactful model that warrants attention.

    "I want a model that is 80% as good as Sonnet for super cheap and it seems that Chinese companies get it because Miniax, GLM, and now this aimed to do just that."

Other People Also See