Sage Advice Hub

DeepSeek readies the next AI disruption with self-improving models

Barely a few months ago, Wall Street’s big bet on generative AI had a moment of reckoning when DeepSeek arrived on the scene. Despite its heavily censored nature, the open source DeepSeek proved that a frontier reasoning AI model doesn’t necessarily require billions of dollars and can be pulled off on modest resources.

It quickly found commercial adoption by giants such as Huawei, Oppo, and Vivo, while the likes of Microsoft, Alibaba, and Tencent quickly gave it a spot on their platforms. Now, the buzzy Chinese company’s next target is self-improving AI models that use a looping judge-reward approach to improve themselves.

Self-improving AI?

The topic of AI that can improve itself has drawn some ambitious and controversial remarks. Former Google CEO, Eric Schmidt, argued that we might need a kill switch for such systems. “When the system can self-improve, we need to seriously think about unplugging it,” Schmidt was quoted as saying by Fortune.

The concept of a recursively self-improving AI is not exactly a novel concept. The idea of an ultra-intelligent machine, which is subsequently capable of making even better machines, actually traces all the way back to mathematician I.J. Good back in 1965. In 2007, AI expert Eliezer Yudkowsky hypothesized about Seed AI, an AI “designed for self-understanding, self-modification, and recursive self-improvement.”

In 2024, Japan’s Sakana AI detailed the concept of an “AI Scientist” about a system capable of passing the whole pipeline of a research paper from beginning to end. In a research paper published in March this year, Meta’s experts revealed self-rewarding language models where the AI itself acts as a judge to provide rewards during training.

Microsoft CEO Satya Nadella says AI development is being optimized by OpenAI’s o1 model and has entered a recursive phase: “we are using AI to build AI tools to build better AI” pic.twitter.com/IHuFIpQl2C
— Tsarathustra (@tsarnick) October 21, 2024

Meta’s internal tests on its Llama 2 AI model using the novel self-rewarding technique saw it outperform rivals such as Anthropic’s Claude 2, Google’s Gemini Pro, and OpenAI’s GPT-4 models. Amazon-backed Anthropic detailed what they called reward-tampering, an unexpected process “where a model directly modifies its own reward mechanism.”

Google is not too far behind on the idea. In a study published in the Nature journal earlier this month, experts at Google DeepMind showcased an AI algorithm called Dreamer that can self-improve, using the Minecraft game as an exercise example.

Experts at IBM are working on their own approach called deductive closure training, where an AI model uses its own responses and evaluates them against the training data to improve itself. The whole premise, however, isn’t all sunshine and rainbows.

Research suggests that when AI models try to train themselves on self-generated synthetic data, it leads to defects colloquially known as “model collapse.” It would be interesting to see just how DeepSeek executes the idea, and whether it can do it in a more frugal fashion than its rivals from the West.

About Us

We are a comprehensive and trusted information platform dedicated to delivering high-quality content across a wide range of topics, including society, technology, business, health, culture, and entertainment.

From breaking news to in-depth reports, we adhere to the principles of accuracy and diverse perspectives, helping readers find clarity and reliability in today’s fast-paced information landscape.

Our goal is to be a dependable source of knowledge for every reader—making information not only accessible but truly trustworthy. Looking ahead, we will continue to enhance our content and services, connecting the world and delivering value.

Sage Advice Hub

DeepSeek readies the next AI disruption with self-improving models

Self-improving AI?

Recommended Articles

The hottest new ChatGPT trend is disturbingly morbid

Expert reveals the phones AI fans need to push Gemini & ChatGPT to the limit

Opera One puts an AI in control of browser tabs, and it’s pretty smart

Clinical test says AI can offer therapy as good as a certified expert

The delay is over — you can now generate images with ChatGPT for free

ChatGPT app could soon generate AI videos with Sora

T-Mobile’s parent company is making an AI Phone with Perplexity

OpenAI’s rebrand is meant to make the company appear ‘more human’

About Us