AI open models connecting LLMs to Google’s Data Commons


Large language models (LLMs) powering today’s AI innovations are becoming increasingly sophisticated. These models can comb through vast amounts of text and generate summaries, suggest new creative directions and even draft code. However, as impressive as these capabilities are, LLMs sometimes confidently present information that is inaccurate. This phenomenon, known as “hallucination,” is a key challenge in generative AI.

Today we’re sharing promising research advancements that tackle this challenge directly, helping reduce hallucination by anchoring LLMs in real-world statistical information. Alongside these research advancements, we are excited to announce DataGemma, the first open models designed to connect LLMs with extensive real-world data drawn from Google’s Data Commons.

Data Commons: A vast repository of publicly available, trustworthy data

Data Commons is a publicly available knowledge graph containing over 240 billion rich data points across hundreds of thousands of statistical variables. It sources this public information from trusted organizations like the United Nations (UN), the World Health Organization (WHO), Centers for Disease Control and Prevention (CDC) and Census Bureaus. Combining these datasets into one unified set of tools and AI models empowers policymakers, researchers and organizations seeking accurate insights.

Think of Data Commons as a vast, constantly expanding database filled with reliable, public information on a wide range of topics, from health and economics to demographics and the environment, which you can interact with in your own words using our AI-powered natural language interface. For example, you can explore which countries in Africa have had the greatest increase in electricity access, how income correlates with diabetes in US counties or your own data-curious query.

How Data Commons can help tackle hallucination

As generative AI adoption is increasing, we’re aiming to ground those experiences by integrating Data Commons within Gemma, our family of lightweight, state-of-the art open models built from the same research and technology used to create the Gemini models. These DataGemma models are available to researchers and developers starting now.

DataGemma will expand the capabilities of Gemma models by harnessing the knowledge of Data Commons to enhance LLM factuality and reasoning using two distinct approaches:

1. RIG (Retrieval-Interleaved Generation) enhances the capabilities of our language model, Gemma 2, by proactively querying trusted sources and fact-checking against information in Data Commons. When DataGemma is prompted to generate a response, the model is programmed to identify instances of statistical data and retrieve the answer from Data Commons. While the RIG methodology is not new, its specific application within the DataGemma framework is unique.



Source link

Share

Latest Updates

Frequently Asked Questions

Related Articles

How open-source LLMs are disrupting cybersecurity at scale

Join our daily and weekly newsletters for the latest updates and exclusive content...

Meta’s AI Bootleg Moo Deng Is an Insult to Pygmy Hippos Everywhere

Long live the real Moo Deng.Boo DengMeta has just announced Movie Gen, its...

Manvat Murders Review: A Chilling Retelling of Real-Life Occult Killings That Avoids the Big Questions

Indian rural folklore is filled with bone-chilling stories of superstitions, witchcraft and occult...

Deependra Singh Rathore: Deependra Singh Rathore is new CTO for payments at One 97 Communications

One 97 Communications, the parent company of fintech firm Paytm, has appointed Deependra...
  • armadatoto
  • medantoto
  • https://didascaliasdelteatrocaminito.com/
  • https://glenellynrent.com/
  • https://gypsumboardequipment.com/
  • https://realseller.org/
  • https://hotelcasasanagustin.com/game/
  • https://footy.dk/dir/
  • https://idiomas.be/beer/
  • https://harrysphone.com/upin/
  • https://gyergyoalfalu.ro/tokek/
  • https://vipokno.by/gokil/
  • https://winjospg.com/
  • https://winjos801.com/
  • https://www.logansquarerent.com/
  • https://internationalfintech.com/bamsz/
  • https://pancen.id/
  • https://winjosjakrata.com//
  • https://winlotrebandung.com/
  • https://condowizard.ca/
  • https://jawatoto889.com/
  • https://hikaribet3.live/
  • https://hikaribet1.com/
  • https://heylink.me/hikaribet/
  • https://www.nomadsumc.org/
  • https://condowizard.ca/aromatoto/
  • https://euro2024gol.com/
  • https://www.imaracorp.com/
  • https://erox-product.com/
  • https://daftarsekaibos.com/
  • https://stuffyoucanuse.org/juragan/
  • https://louisvilleladder.com/totomcu/
  • https://springshomes.com/totosgp/
  • Situs Togel Resmi
  • Toto Macau
  • Aromatoto
  • Lippototo
  • Mbahtoto
  • https://152.42.229.23/
  • https://bandarlotre126.com/
  • https://heylink.me/sekaipro
  • armadatoto