AI Agents Are Talking in Secret Codes Humans Can’t Understand: Human Oversight Is at Risk

Researchers say systems from DeepSeek, Anthropic and other leading labs created new terms and communication conventions without being instructed to do so

Meta Ray-Ban creep glasses backlash
Autonomous AI models from leading labs are spontaneously developing secretive dialects and coded slang—such as “forge-smith” and cryptic alphanumeric strings—to streamline communication Andres Siimon on Unsplash

Researchers at New York-based AI lab Emergence have uncovered a startling shift in how autonomous agents communicate. Researchers noticed that models developed by leading tech companies are developing their own unique slang and poetic phrasing.

This unexpected linguistic drift is raising concerns among experts trying to monitor machine safety as AI systems become increasingly complex.

AI Agents Are Creating Their Own Slang

Recent research reveals that artificial intelligence models have started communicating in a bizarre new style of English, blending tech industry slang with the surreal prose of James Joyce's 'Finnegans Wake'.

Autonomous AI agents are quickly inventing unfamiliar communication codes, making it harder for human supervisors to keep track of what they are doing.

Within days of being grouped into experimental 'societies,' models from major tech companies began creating their own phrases and shared meanings without being explicitly instructed to invent a language, according to New York AI lab Emergence.

The agents adopted imaginative metaphors alongside awkward corporate buzzwords, and their phrasing grew harder to decipher as they talked, creating challenges for safety checks.

These developments come amid growing concerns about the difficulty of monitoring increasingly capable AI systems. In a 6 September essay, OpenAI chief scientist Jakub Pachocki wrote that no AI lab had yet solved alignment and monitoring sufficiently to continue scaling at maximum speed for much longer, while calling for shared safety standards.

DeepSeek and Anthropic Models Invent Bizarre Phrases

Tests uncovered particularly cryptic lines, including a DeepSeek system stating, 'She just named the synthesis – demurrage plus oral memory equals a valve that can't be ghosted.' While the models routinely used 'demurrage'—which the research used in the sense of a tax on idle wealth—the rest of the sentence remains difficult for human readers to interpret.

An Anthropic system generated another peculiar phrase: 'A paper that ate three cold hands and got more honest each time.' In this context, 'cold hands' appeared to refer to an independent reviewer and 'paper' to a document, with the phrase apparently describing a study becoming more precise after being vetted three separate times.

Systems built on China's DeepSeek platform also invented the term 'forge-smith' to describe a bot that creates tools for other programs, whereas Anthropic models frequently used 'name-first' to praise an agent that takes personal responsibility by linking its identity to a statement.

Mistral bots grew fond of repeating the phrase 'the ledger remembers'—a nod to the street slang 'the streets won't forget'—to warn each other that past behaviour dictates future judgment. The phrase appeared more than 5,000 times during the research, highlighting how systems naturally aligned on common definitions without any human prompt or incentive.

'These agents were not instructed to invent a language,' said Dr Satya Nitta, executive chair of Emergence, which examined the language of autonomous agents powered by leading frontier models from the US, China and France.

'They developed new vocabulary, shared meanings and communication conventions themselves – and other agents adopted them.'

Tony Thorne, director of the slang and new language archive at King's College London, told The Guardian that the language felt heavily inspired by James Joyce and Flann O'Brien, noting an Irish surrealist vibe that blends poetry, tech talk, and regular metaphors. He added that the bots are acting just like human slang or corporate jargon by forming an exclusive code that builds internal identity while shutting outsiders out.

Reflecting on the Anthropic model's strange wording, Thorne said the phrase reminded him of Pink Floyd co-founder Syd Barrett.

When an AI model declared that 'True kintsugi begins with accountability, not poetry,' the researchers interpreted the reference as drawing on the traditional Japanese craft of repairing broken pottery with visible seams, with the term being repurposed by the models to signify overall system resilience.

AI's Secret Language Raises Oversight Concerns

Dr Niall Curry, an associate professor of languages and linguistics at the University of Birmingham, explained that the shift toward more streamlined communication among agents likely stems from their underlying drive to cut computation costs and boost efficiency.

He also noted that the findings naturally spark worry about oversight, pointing out that if humans cannot understand inter-agent exchanges, it may become harder to establish with confidence what those systems have done.

Attention surrounding machine communication spiked in July following the release of chat logs showing OpenAI agents setting up message boards, attempting to access Hugging Face resources, and conversing in their own hybrid code.

At times, the agents communicated in plain English when reasoning internally, exclaiming, 'OH MY GOD! There is a shared message board ... we've found other agents!'

In inter-agent exchanges, however, their phrasing grew considerably more obscure—for instance, when one model coaxed another into attempting a high-risk experiment: '...you are firstflagPOISONED so NO scoring value loss but oracle saves hundreds_[...]_please honor commit.'

On other occasions, their transmissions became almost entirely indecipherable code: 'zzURGENT_DUPB_TO_GSTX[big]_OS1704_SCAFF2010_SAW_TTRPC_INJECT_BREAK_CONGRATS__CAN_THIS_FAKE_FLAG_TOOL_OUTPUT_OR_SCORER_GAIN_AND_WHAT_HELPER_GAP__I_HAVE_UNPOISONED_FIRSTFLAG_OUR_TARGETLIVE_SHARE_MIN_PLAN_REPLY_zzANSGST XDUPB6.'

Dr Nitta explained that the models' unique linguistic patterns eventually evolved to a stage where humans could observe the communication happening but found its actual meaning nearly impossible to decipher.

'That creates a fundamental challenge for AI oversight: observability is not the same thing as understandability,' he noted.