In April 2025, Google introduced an AI model called DolphinGemma, trained on four decades of dolphin recordings, and said it might one day help people and dolphins understand each other [source: Google, 2025]. That November, researchers proposed that sperm whales produce sounds resembling human vowels [source: Open Mind, 2025]. A year earlier, separate teams reported that elephants and marmoset monkeys appear to call one another by name [source: Nature Ecology & Evolution, 2024][source: Science, 2024]. And a foundation with backing from an investor is dangling a $10 million grand prize for anyone who can genuinely "crack the code" of talking with another species [source: Jeremy Coller Foundation, 2025].
Put together, the headlines sound like the animal kingdom is about to start talking back. The reality is more careful, and more interesting. Machine learning has genuinely transformed how scientists study animal sounds — but there is a wide gap between finding structure in those sounds and understanding what they mean. This piece walks through what has actually been measured, where the honest uncertainty lies, and why the difference matters far beyond a catchy headline.
Why this is suddenly everywhere
Three things converged to push animal communication into the spotlight. First, cheap sensors and long-running field projects have produced enormous archives of recordings — decades of whale codas, dolphin whistles, and elephant rumbles that no human could ever listen through by hand. Second, the same kind of AI that powers chatbots turned out to be very good at finding patterns in any sequence of sounds, not just human speech. Third, big institutions and prize money arrived: Google's DolphinGemma, the nonprofit Earth Species Project's open bioacoustics models, and the Coller Dolittle Challenge, which offers $100,000 a year and a $10 million grand prize for interspecies two-way communication [source: Jeremy Coller Foundation, 2025].
The result is a field moving fast enough to generate a milestone every few months. But speed also generates hype, and the most important skill for a reader right now is separating what a study claims from what it has shown.
The tools that changed the game
To see why AI helped, it helps to know what these models actually do. A generative model of language predicts the next word in a sentence from the words before it; trained on enough text, it absorbs deep statistical structure. The insight of the past few years is that animal sounds are also sequences, so the same machinery can be pointed at them.
Two efforts illustrate the range. The Earth Species Project built NatureLM-audio, described as the first large audio-language "foundation model" for bioacoustics [source: Earth Species Project, 2025]. Trained on millions of audio-text pairs, it can identify thousands of species, tell a call from a song, estimate a bird's life stage, and even caption recordings in plain language — and it generalizes to species it never saw in training [source: Earth Species Project, 2025]. Crucially, the project frames these as tools for monitoring and analysis — detecting which animals are present and what they are doing — not as a translator of animal meaning [source: Earth Species Project, 2025].
Google's DolphinGemma takes a different tack. Built with the Georgia Institute of Technology and the Wild Dolphin Project, it was trained on roughly 40 years of recordings of Atlantic spotted dolphins in the Bahamas [source: Google, 2025]. Like a language model, it predicts the next sound a dolphin is likely to make, and it can generate new dolphin-like sounds — running on a smartphone in the field so researchers can experiment in real time [source: Google, 2025]. Google is explicit that the goal is to help establish two-way communication "one day," not that it has done so [source: Google, 2025].
What has actually been measured
Now to the discoveries themselves — and here the picture is genuinely impressive, as long as you read the claims precisely.
Sperm whales. In 2024, MIT and Project CETI analyzed nearly 9,000 codas — the rhythmic click patterns whales use — from a Caribbean clan of about 60 animals [source: Nature Communications, 2024]. Using machine learning, they found the codas are combinatorial: a small set of features (they called them rhythm, tempo, rubato, and ornamentation) recombine to produce a far larger repertoire than previously recognized [source: Nature Communications, 2024]. In 2025, a follow-up with UC Berkeley reported that whales also produce two vowel-like spectral patterns, and even diphthong-like glides, exchanged in structured turn-taking [source: Open Mind, 2025]. This is a real advance in describing the structure of the system.
Elephants. A 2024 study led by Colorado State University, with Save the Elephants and ElephantVoices, used machine learning to test whether wild African elephant rumbles contain a component specific to the intended recipient [source: Nature Ecology & Evolution, 2024]. They found one — and, in playback experiments, elephants responded more strongly to calls originally addressed to them than to calls meant for others [source: Nature Ecology & Evolution, 2024]. Strikingly, the labels appear arbitrary, not imitations of the listener's own call, which is closer to how human names work than to the copied "signature whistles" dolphins use [source: Nature Ecology & Evolution, 2024].
Marmosets. Also in 2024, a Hebrew University team reported in Science that marmoset monkeys use specific "phee-calls" to address particular individuals, who respond more accurately when called — the first such evidence in a nonhuman primate [source: Science, 2024].
Dolphins. The inaugural Coller Dolittle prize in 2025 went to a Woods Hole–led team studying bottlenose dolphins in Sarasota, Florida. They reported that "non-signature" whistles — about half of all whistles the dolphins make — may carry shared, context-specific meanings, which the foundation called the first evidence of possible language-like communication in dolphins [source: Jeremy Coller Foundation, 2025].
Structure is not the same as meaning
Read that list again and notice a pattern in the pattern. Every one of these results is about form: how sounds are built, whether a call targets an individual, whether a whistle recurs in a context. None of them is a decoded dictionary. This is the first and most important layer to keep straight — the difference between what we feel the headlines promise and what was actually measured.
A "phonetic alphabet" is a good example. The phrase evokes reading whale speech, but what the researchers described is a set of building blocks that could encode meaning — not the meanings themselves. Project CETI has been careful about this: even as it announced the alphabet, it noted that "the communicative function of many codas remains unknown" [source: Project CETI, 2024]. Finding that a system is richly structured raises the ceiling on what it might say. It does not tell you what it does say.
The same caution applies to "names." Showing that a call reliably targets one individual, and that the individual reacts, is strong evidence of addressing. It is not proof that elephants or marmosets have words, grammar, or anything like a human vocabulary. The studies themselves make the narrower claim; the broader one lives mostly in the headlines.
Correlation, causation, and the meaning problem
The second layer is about inference. Much of this research works by correlating a sound with a context — a call type that shows up when a predator is near, a whistle that recurs during a certain activity — and then asking whether the animals treat the sound as meaningful. Playback experiments, where researchers replay a recording and watch the response, are the strongest tool here, and the elephant and marmoset studies used them well [source: Nature Ecology & Evolution, 2024][source: Science, 2024].
But a correlation between a sound and a situation, even a robust one, does not by itself reveal meaning in the human sense. An animal can respond to a call because it recognizes who is calling, or the caller's emotional state, or a learned association — without the call functioning as a word that stands for a concept. Distinguishing "this sound co-occurs with that situation" from "this sound means that thing" is exactly the hard part, and it is where machine learning is least able to help. As one prominent research team warned in Science, these models are powerful pattern detectors and content generators, which is precisely why the risk of overinterpretation is so high; they flagged model validation and research ethics as challenges to tackle head-on [source: Science, 2023].
Announcements versus peer review
The third layer is about evidence status. Not every milestone in this field carries the same weight, and it is easy to blur a company demo, a prize press release, and a peer-reviewed paper into one impression of "progress."
The sperm whale, elephant, and marmoset findings are peer-reviewed studies in journals like Nature Communications, Nature Ecology & Evolution, and Science [source: Nature Communications, 2024][source: Nature Ecology & Evolution, 2024][source: Science, 2024]. NatureLM-audio's capabilities were published and benchmarked at a machine-learning conference [source: Earth Species Project, 2025]. Those are relatively solid. DolphinGemma, by contrast, is a company research tool announced on a corporate blog; predicting and generating dolphin-like sounds is a genuine engineering feat, but it is not the same as a validated, peer-reviewed claim about what dolphins mean [source: Google, 2025]. And the Coller prize honored "possible" language-like communication — the foundation's own hedge — while its $10 million grand prize for actually cracking two-way communication remains unclaimed [source: Jeremy Coller Foundation, 2025]. The prize's very structure is a reminder that the summit has not been reached.
The case for caution
Some scientists worry the field is running ahead of its evidence. The zoologist Arik Kershenbaum argues that we should resist the assumption that animals have a language we can translate at all: "They are sending messages. But it's not language," he has said, cautioning that treating animal communication as if it were human language mostly imposes our own nature onto them [source: University of Cambridge, 2024]. On this view, the danger is not that AI finds nothing, but that it finds structure and we rush to read human meaning into it — the oldest trap in animal behavior, anthropomorphism, wearing a new technological coat.
This is not an argument to stop. It is an argument to match the claim to the evidence. A model that flags a fraudulent-sounding pattern, or sorts a million recordings by species, is doing something real and useful. A model that appears to "reply" to a dolphin is doing something we do not yet know how to interpret. Keeping those apart is the whole discipline.
Why it matters beyond the headlines
There is a serious reason to get this right, and it is not just intellectual hygiene. Better bioacoustic tools are already helping conservation — automatically monitoring endangered populations, detecting animals across vast habitats, and revealing social lives we could not otherwise observe [source: Science, 2023]. Those benefits are real regardless of whether we ever "translate" anything.
The ethics are real too. A 2026 paper in the philosophy journal Topoi argues that using machine learning on animal communication raises questions we have barely started to answer: about animal welfare, about what it would mean to send messages animals can't consent to receiving, and about the risk that a partial "translation" could be used to manipulate or exploit animals rather than protect them [source: Topoi, 2026]. If we tell the public we can talk to whales before we can, we risk both bad policy and broken trust.
What to watch
The honest read in 2026 is neither "we can talk to animals" nor "it's all hype." It is that AI has become a superb instrument for describing animal communication — its structure, its individuality, its surprising complexity — while the leap to meaning remains unmade and genuinely hard.
Three things will tell you whether the field is closing that gap. First, peer-reviewed results that go beyond structure to test specific meanings, and survive independent replication — not press releases. Second, whether tools like DolphinGemma produce interactions that animals demonstrably treat as meaningful in controlled trials, rather than sounds that merely resemble theirs. Third, whether anyone claims that $10 million grand prize under scrutiny. Until then, the smart posture is the scientists' own: take the structure seriously, hold the meaning lightly, and watch the studies, not the headlines.