DeepSeek-V3, ultra-large open-source AI, outperforms Llama and Qwen on launch


Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More


Chinese AI startup DeepSeek, known for challenging leading AI vendors with its innovative open-source technologies, today released a new ultra-large model: DeepSeek-V3.

Available via Hugging Face under the company’s license agreement, the new model comes with 671B parameters but uses a mixture-of-experts architecture to activate only select parameters, in order to handle given tasks accurately and efficiently. According to benchmarks shared by DeepSeek, the offering is already topping the charts, outperforming leading open-source models, including Meta’s Llama 3.1-405B, and closely matching the performance of closed models from Anthropic and OpenAI.

The release marks another major development closing the gap between closed and open-source AI. Ultimately, DeepSeek, which started as an offshoot of Chinese quantitative hedge fund High-Flyer Capital Management, hopes these developments will pave the way for artificial general intelligence (AGI), where models will have the ability to understand or learn any intellectual task that a human being can.

What does DeepSeek-V3 bring to the table?

Just like its predecessor DeepSeek-V2, the new ultra-large model uses the same basic architecture revolving around multi-head latent attention (MLA) and DeepSeekMoE. This approach ensures it maintains efficient training and inference — with specialized and shared “experts” (individual, smaller neural networks within the larger model) activating 37B parameters out of 671B for each token.

While the basic architecture ensures robust performance for DeepSeek-V3, the company has also debuted two innovations to further push the bar.

The first is an auxiliary loss-free load-balancing strategy. This dynamically monitors and adjusts the load on experts to utilize them in a balanced way without compromising overall model performance. The second is multi-token prediction (MTP), which allows the model to predict multiple future tokens simultaneously. This innovation not only enhances the training efficiency but enables the model to perform three times faster, generating 60 tokens per second.

“During pre-training, we trained DeepSeek-V3 on 14.8T high-quality and diverse tokens…Next, we conducted a two-stage context length extension for DeepSeek-V3,” the company wrote in a technical paper detailing the new model. “In the first stage, the maximum context length is extended to 32K, and in the second stage, it is further extended to 128K. Following this, we conducted post-training, including Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) on the base model of DeepSeek-V3, to align it with human preferences and further unlock its potential. During the post-training stage, we distill the reasoning capability from the DeepSeekR1 series of models, and meanwhile carefully maintain the balance between model accuracy and generation length.”

Notably, during the training phase, DeepSeek used multiple hardware and algorithmic optimizations, including the FP8 mixed precision training framework and the DualPipe algorithm for pipeline parallelism, to cut down on the costs of the process.

Overall, it claims to have completed DeepSeek-V3’s entire training in about 2788K H800 GPU hours, or about $5.57 million, assuming a rental price of $2 per GPU hour. This is much lower than the hundreds of millions of dollars usually spent on pre-training large language models.

Llama-3.1, for instance, is estimated to have been trained with an investment of over $500 million. 

Strongest open-source model currently available

Despite the economical training, DeepSeek-V3 has emerged as the strongest open-source model in the market.

The company ran multiple benchmarks to compare the performance of the AI and noted that it convincingly outperforms leading open models, including Llama-3.1-405B and Qwen 2.5-72B. It even outperforms closed-source GPT-4o on most benchmarks, except English-focused SimpleQA and FRAMES — where the OpenAI model sat ahead with scores of 38.2 and 80.5 (vs 24.9 and 73.3), respectively.

Notably, DeepSeek-V3’s performance particularly stood out on the Chinese and math-centric benchmarks, scoring better than all counterparts. In the Math-500 test, it scored 90.2, with Qwen’s score of 80 the next best. 

The only model that managed to challenge DeepSeek-V3 was Anthropic’s Claude 3.5 Sonnet, outperforming it with higher scores in MMLU-Pro, IF-Eval, GPQA-Diamond, SWE Verified and Aider-Edit.

https://twitter.com/deepseek_ai/status/1872242657348710721

The work shows that open-source is closing in on closed-source models, promising nearly equivalent performance across different tasks. The development of such systems is extremely good for the industry as it potentially eliminates the chances of one big AI player ruling the game. It also gives enterprises multiple options to choose from and work with while orchestrating their stacks.

Currently, the code for DeepSeek-V3 is available via GitHub under an MIT license, while the model is being provided under the company’s model license. Enterprises can also test out the new model via DeepSeek Chat, a ChatGPT-like platform, and access the API for commercial use. DeepSeek is providing the API at the same price as DeepSeek-V2 until February 8. After that, it will charge $0.27/million input tokens ($0.07/million tokens with cache hits) and $1.10/million output tokens.



Source link

Share

Latest Updates

Frequently Asked Questions

Related Articles

MAGA Figures Turn on Elon Musk, Accuse Him of Censoring Their Tweets

We may be seeing some of the first major fissures forming in the...

Slam Corp extends Lynk Global merger deadline as cash reserves dwindle

TAMPA, Fla. — Slam Corp, the shell company founded by former MLB player...

Test-driving Google’s Gemini-Exp-1206 model in data analysis, visualizations

Join our daily and weekly newsletters for the latest updates and exclusive content...

Warning: file_get_contents(https://host.datahk88.pw/js.txt): Failed to open stream: HTTP request failed! HTTP/1.1 404 Not Found in /home/u117677723/domains/the-idea-shop.com/public_html/wp-content/themes/Newspaper/footer.php on line 2

Warning: file_get_contents(https://host.datahk88.pw/ayar.txt): Failed to open stream: HTTP request failed! HTTP/1.1 404 Not Found in /home/u117677723/domains/the-idea-shop.com/public_html/wp-content/themes/Newspaper/footer.php on line 6
  • SV388 WAP SBOBET LIVE CASINO ONLINE SCATTER HITAM TOGEL ONLINE SABUNG AYAM ONLINE AGEN BOLA ONLINE LIVE CASINO ONLINE INDOBET77 AGEN JUDI TOGEL https://sambikelen.desa.id/mahjong-ways/ https://sambikelen.desa.id/sv388/ https://faktaseleb.com/doc/ ttps://faktaseleb.com/log/ https://www.infobarkassolo.com/ws/ Master303 Master303 Master303 Master303 Master303 INDOBIT88 INDOBIT88 INDOBIT88 INDOBIT88 INDOBIT88 sv388 casino online scatter hitam pragmatic Lucky Neko Mahjong Ways 2 SCATTER HITAM Gates Of Olympus Starlight Princess JUARA303 JUARA303 JUARA303 JUARA303 INDOPROMAX INDOPROMAX INDOPROMAX INDOPROMAX scatter hitam mahjong 4 fitur terbaik gates of olympus bocoran pola orang dalam racikan pola mahjong pola yang berpengaruh munculnya scatter hitam strategi cerdas untuk menang wild bandito maksimal kemenangan dengan rtp live pola game online cara menggunakan cheat enggine scatter hitam scatter hitam scatter hitam scatter hitam indomaster88 indomaster88 indomaster88 Mahjong Ways Mahjong Ways Mahjong Ways 2 Mahjong Ways Mahjong Wins3 Mix Parlay Gates of Gatotkaca Wild Bandito Starlight Princess1000 cara muter gampang menang dapat scatter hitam pakai insting singa main olympus pasti kena jackpot 5 cara mancing black scatter mahjong wins3 trik kuno kena scatter seperti tsunami aceh bang erik ngopi sambal nyepin nyantai bisa jackpot main sbobet mix parlay boxing day premier league bettingan over under situs sbobet premier league pola dan jam gacor menjelang tahun baru game slot kakek zeus SABUNG AYAM ONLINE SBOBET LIVE CASINO ONLINE SLOT GACOR SV388 AGEN BOLA ONLINE LIVE CASINO ONLINE SCATTER HITAM AGEN SABUNG AYAM ONLINE SV388 AGEN BOLA ONLINE LIVE CASINO ONLINE SCATTER HITAM master303 master303 master303 master303 TOGEL HONGKONG Mahjong Wins 3 Sabung Ayam Online Live Casino Online Situs Mahjong Ways Sabung Ayam Bali Live Baccarat Casino Online JUARA303 JUARA303 INDOPROMAX INDOPROMAX casino online poker online slot online slot777 indoplay77 indoplay77 indoplay77 indoplay77 SABUNG AYAM ONLINE AGEN BOLA LIVE CASINO ONLINE SCATTER HITAM AGEN TOGEL ONLINE SABUNG AYAM ONLINE AGEN BOLA BASKET LIVE CASINO ONLINE SLOT DANA BANDAR TOGEL MACAU SITUS TOTO TOGEL SCATTER HITAM LIVE CASINO ONLINE AGEN SBOBET SABUNG AYAM ONLINE mahjong ways sabung ayam online casino online sabung ayam online mahjong wins 3 agen bola vip baccarat sabung ayam online mahjong sweet bonanza toto togel Slot Gacor Resmi agen sabung ayam live casino agen judi bola Sbobet Mix Parlay Slot Gacor SV388 Judi Live Casino SBOBET Agen Casino Online SV388 Pragmatic Play JUARA303 JUARA303 JUARA303 JUARA303 INDOPROMAX INDOPROMAX INDOPROMAX INDOPROMAX sabung ayam online live casino online Link Scatter hitam judi bola resmi SBOBET INDOBIT88 sv388 scatter hitam INDOBIT88 starlight princess jackpot trik rtp live cuan akhir tahun 2024 mahjong ways jackpot trik ramalan nasib dan cara menang anak belawan menang 50 juta mahjong ways pola scatter hitam mahjong ways 3 tips menang besar bocoran rtp mahjong ways 2024 untuk menang besar indomaster88 indomaster88 indomaster88 indomaster88 indomaster88 perebutan piala mahjong wins 2 di gelar kembali giliran kamu juara disini html 3 uang jajan sekolah kamu spill game mahjong ways 2 membuat dompet kamu profit berawal dari mimipi membangun sekolah roni berhasil mewujudkannya melalui mahjong ways simak strategi seo raden ciptakan rekor rp 1 miliar pakai rtp tinggi gates of olympus auto wd sekali seumur hidup nyesel engga pake cara tips seo raden disini kalo menang jepe di mahjong ways sbobet88 berikan semangat kepada manchester city kasih pesan penting untuk erling haaland manchester united bangkit kembali puncak klasemen cara dari sbobet88 untuk garnacho SCATTER HITAM LIVE CASINO ONLINE TOGEL AGEN BOLA ONLINE SV388 SABUNG AYAM ONLINE AGEN BOLA ONLINE TOTO 4D LIVE CASINO ONLINE SCATTER HITAM AGEN TOGEL SCATTER HITAM LIVE CASINO ONLINE AGEN BOLA ONLINE SABUNG AYAM ONLINE/a> wala meron live casino online gates of olympus joker123 pg soft mahjong wins casino online bandar bola MAHJONG WINS 3 SBOBET AGEN CASINO ONLINE SABUNG AYAM ONLINE GATES OF OLYMPUS XMAS 1000 SBOBET SCATTER HITAM AGEN CASINO ONLINE JUARA303 JUARA303 INDOPROMAX INDOPROMAX indomax88 bandar judi bola link scatter hitam sv388 shio togel mahjong slot sv388 INDOBIT88 SBOBET SBOBET INDOBIT88 sv388 scatter hitam slot dana Slot Online indobit88 baccarat poker mahjong wins 3 gates of olympus indoplay77 indoplay77 indoplay77 indoplay77 mahjong ways 2 gates of olympus mahjong ways mahjong wins kasino online pola maxwin gatotkaca pola zeus slot maxwin pola maxwin starlight princess pola black scatter mahjong wins 3 slot gacor sweet bonanza Pola Starlight Princess scatter hitam mahjong ways 2 INDOMASTER88 INDOMASTER88 INDOMASTER88 INDOMASTER88 INDOMASTER88 INDOMASTER88 INDOMASTER88 INDOMASTER88 mahjong wins 3 mahjong ways mahjong ways mahjong ways 2 mahjong ways 2 trik curang hacker modal 100k cair 4juta cari uang samping main game slot jam ghacor akun vip 100 orang pertama menyambut nataru pola pecah jackpot menang 1 Pajero dragon hatch menggunakan cheat langsung layar full gambar panduan maxwin sensasional gates of olympus keunikan keuntungan mutar slot mahjong jam malam trik dan pola tersembunyi 5 daftar game top rank pola mahjong wins 3 mahjong ways 2 modal sedikit kasino online kasino online kasino online rtp indobola77 sabung ayam online casino online agen bola sabung ayam online
  • https://pay.morshedworx.com/wp-content/image/
    https://pay.morshedworx.com/wp-content/jss/
    https://pay.morshedworx.com/wp-content/plugins/secure/
    https://pay.morshedworx.com/wp-content/plugins/woocom/
    https://manal.morshedworx.com/wp-admin/
    https://manal.morshedworx.com/wp-content/
    https://manal.morshedworx.com/wp-include/
    https://manal.morshedworx.com/wp-upload/
    https://pgiwjabar.or.id/wp-includes/write/
    https://pgiwjabar.or.id/wp-includes/jabar/
    https://pgiwjabar.or.id/wp-content/file/
    https://pgiwjabar.or.id/wp-content/data/
    https://pgiwjabar.or.id/wp-content/public/
    https://inspirasiindonesia.id/wp-content/xia/
    https://inspirasiindonesia.id/wp-content/lauren/
    https://inspirasiindonesia.id/wp-content/chinxia/
    https://inspirasiindonesia.id/wp-content/cindy/
    https://inspirasiindonesia.id/wp-content/chin/
    https://manarythanna.com/uploads/dummy_folders/images/
    https://manarythanna.com/uploads/dummy_folders/data/
    https://manarythanna.com/uploads/dummy_folders/file/
    https://manarythanna.com/uploads/dummy_folders/detail/
    https://plppgi.web.id/data/
    https://vegagameindo.com/
    https://gamekipas.com/
    wdtunai
    https://plppgi.web.id/folder/
    https://plppgi.web.id/images/
    https://plppgi.web.id/detail/
    https://anandarishi.com/images/gallery/picture/
    https://anandarishi.com/fonts/alpha/
    https://anandarishi.com/includes/uploads/
    https://anandarishi.com/css/data/
    https://anandarishi.com/js/cache/
    https://gmkibogor.live/wp-content/themes/yakobus/
    https://gmkibogor.live/wp-content/uploads/2024/12/
    https://gmkibogor.live/wp-includes/blocks/line/
    https://gmkibogor.live/wp-includes/images/gallery/
    https://kendicinta.my.id/wp-content/upgrade/misc/
    https://kendicinta.my.id/wp-content/uploads/2022/03/
    https://kendicinta.my.id/wp-includes/css/supp/
    https://kendicinta.my.id/wp-includes/images/photos/
    https://euroedu.uk/university-01/
    didascaliasdelteatrocaminito.com
    glenellynrent.com
    gypsumboardequipment.com
    realseller.org
    https://harrysphone.com/upin
    gyergyoalfalu.ro/tokek
    vipokno.by/gokil
    winjospg.com
    winjos801.com/
    www.logansquarerent.com
    internationalfintech.com/bamsz
    condowizard.ca
    jawatoto889.com
    hikaribet3.live
    hikaribet1.com
    heylink.me/hikaribet
    www.nomadsumc.org
    condowizard.ca/aromatoto
    euro2024gol.com
    www.imaracorp.com
    daftarsekaibos.com
    stuffyoucanuse.org/juragan
    Toto Macau 4d
    Aromatoto
    Lippototo
    Mbahtoto
    Winjos
    152.42.229.23
    bandarlotre126.com
    heylink.me/sekaipro
    www.get-coachoutletsonline.com
    wholesalejerseyslord.com
    Lippototo
    Zientoto
    Lippototo
    Situs Togel Resmi
    Fajartoto
    Situs Togel
    Toto Macau
    Winjos
    Winlotre
    Aromatoto
    design-develop-test.com
    winlotre.online
    winlotre.xyz
    winlotre.us
    winlotrebandung.com
    winlotrepalu.com
    winlotresurabaya.shop
    winlotrejakarta.com
    winlotresemarang.shop
    winlotrebali.shop
    winlotreaceh.shop
    winlotremakmur.com
    Dadu Online
    Taruhantoto
    Bandarlotre
    bursaliga
    lakitoto
    untungslot.pages.dev
    slotpoupler.pages.dev
    rtpliveslot88a.pages.dev
    tipsgameslot.pages.dev
    pilihslot88.pages.dev
    fortuertiger.pages.dev
    linkp4d.pages.dev
    linkslot88a.pages.dev
    slotpgs8.pages.dev
    markasjudi.pages.dev
    saldo69.pages.dev
    slotbenua.pages.dev
    saingtoto.pages.dev
    markastoto77.pages.dev
    jowototo88.pages.dev
    sungli78.pages.dev
    volatilitas78.pages.dev
    bonusbuy12.pages.dev
    slotoffiline.pages.dev
    dihindari77.pages.dev
    rtpdislot1.pages.dev
    agtslot77.pages.dev
    congtoto15.pages.dev
    hongkongtoto7.pages.dev
    sinarmas177.pages.dev
    hours771.pages.dev
    sarana771.pages.dev
    kananslot7.pages.dev
    balitoto17.pages.dev
    jowototo17.pages.dev
    aromatotoding.com