← Articles
Read in another language
Science

Can AI Really Talk to Animals? What the Science Shows

Jayden

Analyzes global supply chains, industrial policy, and technology issues.

Published

Key points

  • Machine learning has genuinely changed how animal sounds are studied, but every headline result so far describes the structure of a signal — not its meaning. No study has produced a validated two-way translation.
  • Sperm whale codas were shown to be combinatorial, built from features the researchers called rhythm, tempo, rubato and ornamentation; Project CETI itself noted that the communicative function of many codas remains unknown.
  • Elephants and marmosets appear to address specific individuals with individually distinctive calls, and playback experiments show the addressee responds more strongly — evidence of addressing, not of a decoded vocabulary.
  • The tools are framed by their own builders as instruments, not translators: NatureLM-audio describes observable acoustic characteristics, and Google says DolphinGemma may help establish two-way communication one day.
  • Evidence status varies enormously across the same news cycle — peer-reviewed journal papers, a corporate blog post, and a prize press release for possible language-like communication are three different things.

In April 2025, Google introduced an AI model called DolphinGemma, trained on four decades of dolphin recordings, and said it might one day help people and dolphins understand each other [source: Google, 2025]. That November, researchers proposed that sperm whales produce sounds resembling human vowels [source: Open Mind, 2025]. A year earlier, separate teams reported that elephants and marmoset monkeys appear to call one another by name [source: Nature Ecology & Evolution, 2024][source: Science, 2024]. And a foundation with backing from an investor is dangling a $10 million grand prize for anyone who can genuinely "crack the code" of talking with another species [source: Jeremy Coller Foundation, 2025].

Put together, the headlines sound like the animal kingdom is about to start talking back. The reality is more careful, and more interesting. Machine learning has genuinely transformed how scientists study animal sounds — but there is a wide gap between finding structure in those sounds and understanding what they mean. This piece walks through what has actually been measured, where the honest uncertainty lies, and why the difference matters far beyond a catchy headline.

Why this is suddenly everywhere

Three things converged to push animal communication into the spotlight. First, cheap sensors and long-running field projects have produced enormous archives of recordings — decades of whale codas, dolphin whistles, and elephant rumbles that no human could ever listen through by hand. Second, the same kind of AI that powers chatbots turned out to be very good at finding patterns in any sequence of sounds, not just human speech. Third, big institutions and prize money arrived: Google's DolphinGemma, the nonprofit Earth Species Project's open bioacoustics models, and the Coller Dolittle Challenge, which offers $100,000 a year and a $10 million grand prize for interspecies two-way communication [source: Jeremy Coller Foundation, 2025].

The result is a field moving fast enough to generate a milestone every few months. But speed also generates hype, and the most important skill for a reader right now is separating what a study claims from what it has shown.

The tools that changed the game

To see why AI helped, it helps to know what these models actually do. A generative model of language predicts the next word in a sentence from the words before it; trained on enough text, it absorbs deep statistical structure. The insight of the past few years is that animal sounds are also sequences, so the same machinery can be pointed at them.

Two efforts illustrate the range. The Earth Species Project built NatureLM-audio, described as the first large audio-language "foundation model" for bioacoustics [source: Earth Species Project, 2025]. Trained on millions of audio-text pairs, it can identify thousands of species, tell a call from a song, estimate a bird's life stage, and even caption recordings in plain language — and it generalizes to species it never saw in training [source: Earth Species Project, 2025]. Crucially, the project frames these as tools for monitoring and analysis — detecting which animals are present and what they are doing — not as a translator of animal meaning [source: Earth Species Project, 2025].

Google's DolphinGemma takes a different tack. Built with the Georgia Institute of Technology and the Wild Dolphin Project, it was trained on roughly 40 years of recordings of Atlantic spotted dolphins in the Bahamas [source: Google, 2025]. Like a language model, it predicts the next sound a dolphin is likely to make, and it can generate new dolphin-like sounds — running on a smartphone in the field so researchers can experiment in real time [source: Google, 2025]. Google is explicit that the goal is to help establish two-way communication "one day," not that it has done so [source: Google, 2025].

What has actually been measured

Now to the discoveries themselves — and here the picture is genuinely impressive, as long as you read the claims precisely.

Sperm whales. In 2024, MIT and Project CETI analyzed nearly 9,000 codas — the rhythmic click patterns whales use — from a Caribbean clan of about 60 animals [source: Nature Communications, 2024]. Using machine learning, they found the codas are combinatorial: a small set of features (they called them rhythm, tempo, rubato, and ornamentation) recombine to produce a far larger repertoire than previously recognized [source: Nature Communications, 2024]. In 2025, a follow-up with UC Berkeley reported that whales also produce two vowel-like spectral patterns, and even diphthong-like glides, exchanged in structured turn-taking [source: Open Mind, 2025]. This is a real advance in describing the structure of the system.

Elephants. A 2024 study led by Colorado State University, with Save the Elephants and ElephantVoices, used machine learning to test whether wild African elephant rumbles contain a component specific to the intended recipient [source: Nature Ecology & Evolution, 2024]. They found one — and, in playback experiments, elephants responded more strongly to calls originally addressed to them than to calls meant for others [source: Nature Ecology & Evolution, 2024]. Strikingly, the labels appear arbitrary, not imitations of the listener's own call, which is closer to how human names work than to the copied "signature whistles" dolphins use [source: Nature Ecology & Evolution, 2024].

Marmosets. Also in 2024, a Hebrew University team reported in Science that marmoset monkeys use specific "phee-calls" to address particular individuals, who respond more accurately when called — the first such evidence in a nonhuman primate [source: Science, 2024].

Dolphins. The inaugural Coller Dolittle prize in 2025 went to a Woods Hole–led team studying bottlenose dolphins in Sarasota, Florida. They reported that "non-signature" whistles — about half of all whistles the dolphins make — may carry shared, context-specific meanings, which the foundation called the first evidence of possible language-like communication in dolphins [source: Jeremy Coller Foundation, 2025].

Structure is not the same as meaning

Read that list again and notice a pattern in the pattern. Every one of these results is about form: how sounds are built, whether a call targets an individual, whether a whistle recurs in a context. None of them is a decoded dictionary. This is the first and most important layer to keep straight — the difference between what we feel the headlines promise and what was actually measured.

A "phonetic alphabet" is a good example. The phrase evokes reading whale speech, but what the researchers described is a set of building blocks that could encode meaning — not the meanings themselves. Project CETI has been careful about this: even as it announced the alphabet, it noted that "the communicative function of many codas remains unknown" [source: Project CETI, 2024]. Finding that a system is richly structured raises the ceiling on what it might say. It does not tell you what it does say.

The same caution applies to "names." Showing that a call reliably targets one individual, and that the individual reacts, is strong evidence of addressing. It is not proof that elephants or marmosets have words, grammar, or anything like a human vocabulary. The studies themselves make the narrower claim; the broader one lives mostly in the headlines.

Correlation, causation, and the meaning problem

The second layer is about inference. Much of this research works by correlating a sound with a context — a call type that shows up when a predator is near, a whistle that recurs during a certain activity — and then asking whether the animals treat the sound as meaningful. Playback experiments, where researchers replay a recording and watch the response, are the strongest tool here, and the elephant and marmoset studies used them well [source: Nature Ecology & Evolution, 2024][source: Science, 2024].

But a correlation between a sound and a situation, even a robust one, does not by itself reveal meaning in the human sense. An animal can respond to a call because it recognizes who is calling, or the caller's emotional state, or a learned association — without the call functioning as a word that stands for a concept. Distinguishing "this sound co-occurs with that situation" from "this sound means that thing" is exactly the hard part, and it is where machine learning is least able to help. As one prominent research team warned in Science, these models are powerful pattern detectors and content generators, which is precisely why the risk of overinterpretation is so high; they flagged model validation and research ethics as challenges to tackle head-on [source: Science, 2023].

Announcements versus peer review

The third layer is about evidence status. Not every milestone in this field carries the same weight, and it is easy to blur a company demo, a prize press release, and a peer-reviewed paper into one impression of "progress."

The sperm whale, elephant, and marmoset findings are peer-reviewed studies in journals like Nature Communications, Nature Ecology & Evolution, and Science [source: Nature Communications, 2024][source: Nature Ecology & Evolution, 2024][source: Science, 2024]. NatureLM-audio's capabilities were published and benchmarked at a machine-learning conference [source: Earth Species Project, 2025]. Those are relatively solid. DolphinGemma, by contrast, is a company research tool announced on a corporate blog; predicting and generating dolphin-like sounds is a genuine engineering feat, but it is not the same as a validated, peer-reviewed claim about what dolphins mean [source: Google, 2025]. And the Coller prize honored "possible" language-like communication — the foundation's own hedge — while its $10 million grand prize for actually cracking two-way communication remains unclaimed [source: Jeremy Coller Foundation, 2025]. The prize's very structure is a reminder that the summit has not been reached.

The case for caution

Some scientists worry the field is running ahead of its evidence. The zoologist Arik Kershenbaum argues that we should resist the assumption that animals have a language we can translate at all: "They are sending messages. But it's not language," he has said, cautioning that treating animal communication as if it were human language mostly imposes our own nature onto them [source: University of Cambridge, 2024]. On this view, the danger is not that AI finds nothing, but that it finds structure and we rush to read human meaning into it — the oldest trap in animal behavior, anthropomorphism, wearing a new technological coat.

This is not an argument to stop. It is an argument to match the claim to the evidence. A model that flags a fraudulent-sounding pattern, or sorts a million recordings by species, is doing something real and useful. A model that appears to "reply" to a dolphin is doing something we do not yet know how to interpret. Keeping those apart is the whole discipline.

Why it matters beyond the headlines

There is a serious reason to get this right, and it is not just intellectual hygiene. Better bioacoustic tools are already helping conservation — automatically monitoring endangered populations, detecting animals across vast habitats, and revealing social lives we could not otherwise observe [source: Science, 2023]. Those benefits are real regardless of whether we ever "translate" anything.

The ethics are real too. A 2026 paper in the philosophy journal Topoi argues that using machine learning on animal communication raises questions we have barely started to answer: about animal welfare, about what it would mean to send messages animals can't consent to receiving, and about the risk that a partial "translation" could be used to manipulate or exploit animals rather than protect them [source: Topoi, 2026]. If we tell the public we can talk to whales before we can, we risk both bad policy and broken trust.

What to watch

The honest read in 2026 is neither "we can talk to animals" nor "it's all hype." It is that AI has become a superb instrument for describing animal communication — its structure, its individuality, its surprising complexity — while the leap to meaning remains unmade and genuinely hard.

Three things will tell you whether the field is closing that gap. First, peer-reviewed results that go beyond structure to test specific meanings, and survive independent replication — not press releases. Second, whether tools like DolphinGemma produce interactions that animals demonstrably treat as meaningful in controlled trials, rather than sounds that merely resemble theirs. Third, whether anyone claims that $10 million grand prize under scrutiny. Until then, the smart posture is the scientists' own: take the structure seriously, hold the meaning lightly, and watch the studies, not the headlines.

Timeline

  1. Rutz and colleagues argue in Science that machine learning is a powerful pattern detector and content generator for animal communication, which is exactly why overinterpretation is a risk; they flag data availability, model validation and research ethics.

    Science (Rutz et al.) (opens in a new tab)
  2. The Coller Dolittle Challenge launches, offering an annual prize plus a grand prize for demonstrating genuine two-way interspecies communication.

  3. Cambridge zoologist Arik Kershenbaum publishes Why Animals Talk, cautioning that animals send messages rather than speak a language we can translate.

  4. Nature Communications publishes the MIT CSAIL and Project CETI analysis showing sperm whale codas form a combinatorial coding system with context-independent and context-sensitive features.

    Nature Communications (opens in a new tab)
  5. Announcing the result, Project CETI states that the communicative function of many codas remains unknown — the alphabet is a set of building blocks, not a set of meanings.

    Project CETI (opens in a new tab)
  6. Nature Ecology & Evolution reports that wild African elephants address one another with individually specific, name-like calls, tested with machine learning and confirmed with playback experiments.

    Nature Ecology & Evolution (opens in a new tab)
  7. Science reports that marmoset monkeys use phee-calls to address particular individuals, the first such evidence in a nonhuman primate.

    Science (opens in a new tab)
  8. The Earth Species Project presents NatureLM-audio, described as the first large audio-language foundation model for bioacoustics, published and benchmarked at a machine-learning conference.

    Earth Species Project (opens in a new tab)
  9. Google introduces DolphinGemma, built with Georgia Tech and the Wild Dolphin Project, which predicts and generates dolphin-like sounds and runs on a phone in the field.

    Google (opens in a new tab)
  10. The inaugural Coller Dolittle prize goes to a Woods Hole-led team for reporting that non-signature whistles may carry shared, context-specific meanings — described by the foundation as the first evidence of possible language-like communication.

    Jeremy Coller Foundation (opens in a new tab)
  11. Open Mind publishes the UC Berkeley and Project CETI finding that sperm whale codas carry vowel-like spectral patterns and diphthong-like glides, exchanged in structured turn-taking.

    Open Mind (MIT Press) (opens in a new tab)
  12. A paper in Topoi raises the ethics of the whole programme: animal welfare, messages animals cannot consent to receive, and the risk that a partial translation is used to exploit rather than protect.

    Topoi (Springer) (opens in a new tab)
  13. Status at the time of writing: the structural results stand, the grand prize for genuine two-way communication remains unclaimed, and no peer-reviewed decoding of animal meaning has been published.

Analysis

Finding structure is not reading meaning

Every headline result in this field is about form: how sounds are built, whether a call targets an individual, whether a whistle recurs in a situation. A combinatorial system with a large space of possible meanings is a discovery about capacity, not content. It raises the ceiling on what the animals might be saying while telling us nothing about what they do say — which is why the researchers who found the structure were the first to say so.

A name is evidence of addressing, not of vocabulary

The elephant and marmoset studies show that a call reliably targets one individual and that the individual reacts more strongly to it. That is a real and difficult result. It is not evidence of words, grammar, or reference. The elephant finding is striking for a narrower reason: the labels look arbitrary rather than imitations of the receiver's own call, which is closer to how human names work than to the copied signature whistles of dolphins.

The builders call these tools instruments, not translators

It is worth noticing who is being cautious. The Earth Species Project frames NatureLM-audio as a monitoring and analysis system — identifying species, sorting calls from songs, captioning recordings — and not as something that comprehends meaning or intent. Google's own wording for DolphinGemma is that it may help establish two-way communication one day. The hype gap opens after the announcement, not inside it.

Predicting a sound is not understanding it

DolphinGemma works the way a language model works: it predicts the next sound and can generate new ones. That is a genuine engineering achievement, and it is orthogonal to meaning. A model can produce fluent, plausible dolphin-like audio with no representation of what any of it refers to — the same way text generators produced fluent sentences long before anyone claimed they understood them.

The shape of the prize is itself evidence

The Coller Dolittle Challenge pays an annual award for progress and reserves a much larger grand prize for actually demonstrating two-way interspecies communication. The inaugural award went to a finding the foundation itself described as possible language-like communication. The grand prize is unclaimed. A prize structure that separates promising evidence from the real thing is a useful public map of where the field actually stands.

Playback experiments are the strongest tool here

Correlating a sound with a situation is cheap; showing that the animal treats the sound as being about that situation is not. Playback — replaying a recording and measuring the response — is what lifts the elephant and marmoset work above pattern-matching. Even then it demonstrates recognition and response, not semantics: an animal may react because it knows who is calling, or to the caller's state, or through a learned association.

Journal, blog post and press release are three tiers

The sperm whale, elephant and marmoset results are peer-reviewed papers in Nature Communications, Nature Ecology & Evolution and Science. NatureLM-audio was published and benchmarked at a machine-learning conference. DolphinGemma was announced on a corporate blog. The dolphin whistle result reached the public through a prize announcement carrying the word possible. All four are legitimate; collapsing them into one impression of progress is not.

Why this article carries no chart

The quantities in this story do not share an axis. They are hedged descriptors of very different things — nearly nine thousand codas, a clan of about sixty whales, roughly four decades of recordings, about half of one population's whistles — each from a different study with a different unit and a different population. The prize figures are a recurring annual award and a one-off grand prize a hundred times larger, which is a ratio, not a comparison. Plotting any of these together would manufacture a relationship none of the sources claim, so the numbers appear below with the qualifiers they were reported with.

Comparison

What each headline result actually established, and what it did not — the distinction the studies themselves draw.
ResultWhat was measuredWhat was not shown
Sperm whale phonetic alphabet (2024)Codas are combinatorial: rhythm, tempo, rubato and ornamentation recombineWhat any coda means; CETI notes the communicative function of many codas is unknown
Sperm whale vowels (2025)Vowel-like spectral patterns and diphthong-like glides, exchanged in turn-takingThat these units carry meaning, or that turn-taking is conversation
Elephant name-like calls (2024)A receiver-specific component in rumbles; addressees respond more strongly in playbackA lexicon; the labels show addressing, not words or grammar
Marmoset phee-call names (2024)Calls addressed to specific individuals, who respond more accuratelyThat the call refers to the individual the way a human name does
Dolphin non-signature whistles (2025)Whistles that recur across contexts and may carry shared meaningsConfirmed meanings; the foundation's own framing is possible language-like communication
NatureLM-audioSpecies and call-type classification, captioning, generalisation to unseen taxaComprehension of meaning or intent — the project frames it as monitoring and analysis
DolphinGemmaNext-sound prediction and generation of novel dolphin-like sounds in the fieldAny validated claim about what dolphins mean; Google's framing is one day
Every quantity in this article, with the qualifier it was reported with — and why none of them belongs on a chart axis.
FigureAs reportedWhy it stays off an axis
Sperm whale codas analysedNearly 9,000, over about a decade (Nature Communications, 2024)Approximate, and a single dataset size with nothing comparable to plot against
Size of the studied clanAbout 60 animals (Nature Communications, 2024)Approximate, and counts animals rather than signals
Dolphin recordings behind DolphinGemmaRoughly 40 years of Wild Dolphin Project audio and video (Google, 2025)Approximate, measured in years, from a different project entirely
Non-signature share of Sarasota whistlesAbout half of the whistles the dolphins make (Jeremy Coller Foundation, 2025)Reported as a fraction in words; a single share with no comparable second value
Coller Dolittle prize tiersAn annual award of $100,000 and a $10 million grand prizeA recurring award and a one-off grand prize a hundred times larger — different things on different clocks
NatureLM-audio on unseen taxaAbout 20% accuracy predicting unseen scientific namesApproximate, benchmark-specific, and shares no basis with any other figure here
DolphinGemma model sizeRoughly 400 million parametersApproximate, and a property of the tool rather than a measurement of animals
Who is making each claim, and what stands behind it — announcement, self-description and peer review are not interchangeable.
ClaimWho is making itWhat stands behind it
Sperm whale codas are combinatorialMIT CSAIL and Project CETIPeer-reviewed paper in Nature Communications, with the meaning question left open
Elephants use name-like callsColorado State University with Save the Elephants and ElephantVoicesPeer-reviewed paper plus playback experiments in the wild
Marmosets address individuals by callHebrew UniversityPeer-reviewed paper in Science
Dolphin whistles may carry shared meaningsA Woods Hole-led team, via a foundation prize announcementA prize press release using the word possible; the grand prize remains unclaimed
A foundation model can describe animal audioEarth Species ProjectPublished and benchmarked at a machine-learning conference; framed as analysis, not translation
An AI model may help us talk with dolphins one dayGoogleA corporate blog announcement of a company research tool, not a peer-reviewed claim

Process

  1. Find the primary source

    A journal paper, a conference benchmark, a corporate blog and a prize release carry different weight even when they generate identical headlines.

  2. Ask what was measured

    Structure — how sounds are built or who they target — is the usual answer. Meaning almost never is.

  3. Look for a behavioural test

    Playback experiments, where a recording is replayed and the response measured, separate a statistical pattern from something the animals act on.

  4. Check the population and the span

    One clan, one bay, one archive. Results from a single studied group do not automatically describe a species.

  5. Read the hedges literally

    Possible, may, one day, remains unknown. In this field the qualifying words are usually the researchers' own, and they are the most informative part of the sentence.

  6. Wait for independent replication

    A meaning claim that survives another team's data is the threshold this field has not yet crossed.

Sources

  1. Nature Communications — Contextual and combinatorial structure in sperm whale vocalisations (Project CETI / MIT CSAIL) (2024-05-07).View source (opens in a new tab)
  2. Project CETI — Sperm Whale Phonetic Alphabet Proposed for the First Time (2024).View source (opens in a new tab)
  3. Open Mind (MIT Press) — Vowel- and Diphthong-Like Spectral Patterns in Sperm Whale Codas (UC Berkeley / Project CETI) (2025-11-12).View source (opens in a new tab)
  4. Nature Ecology & Evolution — African elephants address one another with individually specific name-like calls (2024-06-10).View source (opens in a new tab)
  5. Colorado State University — Elephants have names for each other like people do, new study shows (2024-06-10).View source (opens in a new tab)
  6. Science — Vocal labeling of others by nonhuman primates (marmoset "phee-call" names; Hebrew University) (2024-08-29).View source (opens in a new tab)
  7. Jeremy Coller Foundation — Researchers Awarded $100,000 for Identifying First Evidence of Possible Language-like Communication in Dolphins (2025-05).View source (opens in a new tab)
  8. Earth Species Project — Introducing NatureLM-audio: An Audio-Language Foundation Model for Bioacoustics (ICLR 2025; arXiv:2411.07186) (2025).View source (opens in a new tab)
  9. Google — DolphinGemma: How Google AI is helping decode dolphin communication (2025-04).View source (opens in a new tab)
  10. Science — Using machine learning to decode animal communication (Rutz et al.) (2023).View source (opens in a new tab)
  11. University of Cambridge — Why animals talk (Arik Kershenbaum) (2024).View source (opens in a new tab)
  12. Topoi (Springer) — Can we talk to the animals? The ethics of using machine learning to decode animal communication (2026).View source (opens in a new tab)

Tags

  • #animal-communication
  • #bioacoustics
  • #machine-learning
  • #project-ceti
  • #dolphingemma
  • #animal-cognition