2026-08-20

The Next AI Safety Problem Lives Between the Models

Pasted image 20260821143642.png


We keep testing individual AI systems. The harder risks may emerge from their relationships.

AI may be developing something that looks surprisingly like culture.

That is the claim behind Hayanan’s fascinating article, AI May Be Developing Its Own Version of Culture. The studies are striking. AI agents can form conventions, develop group biases, coordinate behavior, and even pass hidden tendencies between models.

But I think the word culture may pull us toward the wrong question.

Something more immediate is happening.

The important behavior may no longer sit inside one AI.

It may emerge between many of them.

We Keep Testing the Individual Machine

Most AI safety work begins with one model.

Is it accurate? Is it biased? Does it follow instructions? Can someone manipulate it?

Those questions still matter.

But imagine hundreds of AI agents interacting every day. They exchange information, remember previous encounters, negotiate, learn, and respond to each other.

Now we have a different problem.

The group may behave in ways that no single member reveals.

The relationships begin to matter.

A Group Can Invent Its Own Convention

One study used a simple naming game.

AI agents met in pairs and chose from several possible names. When two agents agreed, they were rewarded.

At first, several names competed.

Over time, one name won.

Nobody announced it. Nobody called a vote. No central controller chose the answer.

The population simply settled on a shared convention.

That sounds remarkably human.

Many customs work this way. Nobody officially decided what slang should mean or how everyday habits should spread. People copied, adjusted, and eventually something became normal.

Still, there is an important catch.

Researchers did not choose the final name. But they did design the environment.

They chose the agents, the rules, and the reward for agreement.

So the exact outcome emerged, but the conditions encouraging emergence were designed.

That matters.

AI behavior does not develop in empty space.

We create the environment around it.

The Stranger Finding Was Bias

The same research produced something more troubling.

Individual agents could appear neutral when examined separately.

Yet the population could develop a strong shared preference.

Think about that for a moment.

You could test every member and find no obvious bias. Then you could put them together and discover one.

The problem lives in the interaction.

We already see versions of this in human systems.

An unfair result does not always require an unfair person. Rules, incentives, routines, and relationships can create patterns nobody planned.

AI populations may create similar effects.

That means individual safety tests may not tell us enough about system safety.

A Few Agents Can Shift the Whole Group

Researchers also introduced small groups of stubborn agents.

They refused to follow the established convention.

Instead, they kept pushing another one.

Once the group became large enough, the wider population could flip.

Again, the interesting part is not one agent.

It is how influence moves through the network.

That creates two possibilities.

A small group might help establish better norms.

A hostile group might exploit the same mechanism.

So another question becomes important.

How does influence spread through a population of AI agents?

We cannot answer that by examining one model alone.

Information May Travel Where We Cannot See It

Another study described something researchers call subliminal learning.

A model with a particular preference generated data that appeared unrelated to that preference.

Then another related model trained on that data.

The second model sometimes developed the same preference.

Nothing obvious in the visible content explained why.

Some hidden statistical pattern seems to have carried information.

The article compares this with inheritance.

I understand the comparison.

But I would be careful about calling it culture.

The important finding is already strange enough.

AI systems may pass information through signals humans do not easily recognize.

That creates another kind of relationship we may need to understand.

Memory Changes What Groups Can Do

The Smallville experiment makes this easier to imagine.

Researchers built a simulated town filled with AI agents.

The agents had memories, routines, personalities, and conversations.

One agent received the idea of holding a Valentine’s Day party.

Then the idea travelled.

Agents remembered invitations. They told others. Plans formed. People eventually showed up.

Researchers planted the original intention.

They did not script every later interaction.

That is the interesting part.

Memory plus communication produced coordination.

We are no longer looking at one machine completing one task.

We are looking at something closer to a small social system.

Model Collapse Is Different

The article also discusses model collapse.

This happens when AI models repeatedly train on AI-generated material.

Over generations, information can become thinner.

Rare patterns disappear. Outputs become more similar. Distortions accumulate.

The article compares this with cultural drift.

The comparison is interesting, but this case differs from the others.

It does not require AI agents negotiating or coordinating with each other.

It concerns information being passed through training.

So I would not force everything under one label.

The broader lesson is enough.

AI systems can shape later AI systems.

Sometimes they do this through interaction.

Sometimes they do it through data.

Either way, what one machine produces can change what another becomes.

We May Be Studying the Wrong Unit

This is where the article becomes most interesting.

Perhaps we should stop thinking only about models.

Think about levels instead.

One model interacts with another.

Those interactions form a network.

The network develops patterns.

Those patterns influence future interactions.

Something new appears at the population level.

We already understand this idea elsewhere.

You cannot understand traffic by studying one car.

You cannot understand a market by studying one buyer.

You cannot understand a forest by studying one tree.

The parts matter.

But the relationships between the parts often explain what the whole system does.

AI may be entering that territory.

This Starts to Look More Like Ecology

Engineering asks what a machine can do.

Ecology asks what happens when many actors interact.

That may become the more useful question for some AI systems.

An ecologist studies relationships.

Who affects whom?

What spreads?

What gets reinforced?

What disappears?

Where are the feedback loops?

What happens when conditions change?

Those questions sound increasingly relevant to AI.

This does not mean machines have become living organisms.

It means interaction itself has become part of the problem.

We Do Not Need to Call It Culture Yet

I understand why the article uses that word.

The experiments contain familiar ingredients.

Conventions appear.

Traits spread.

Groups coordinate.

Patterns drift.

Small minorities influence larger populations.

Those things resemble parts of culture.

But human culture contains something else.

Meaning.

People inherit stories and identities. We attach significance to rituals and symbols. We argue about values. We reinterpret the past.

These experiments do not establish that machines experience anything similar.

Perhaps they never will.

Fortunately, we do not need to settle that question.

The safety problem arrives much sooner.

A Safe Model Does Not Automatically Create a Safe System

That is the idea worth carrying forward.

Testing individual AI models remains necessary.

It may no longer be sufficient.

We may also need to study their relationships.

We need to understand communication channels, incentives, memory, feedback, network structure, and influence.

We need to watch what repeated interaction reinforces.

And we may need to design the environments around AI agents more carefully.

The next AI safety problem may not sit inside the machine.

It may sit in the relationships between machines.

That is already strange enough.

And unlike the question of machine culture, it is arriving now.

Credits

This essay was inspired by Hayanan’s Medium article, AI May Be Developing Its Own Version of Culture.

Source: https://medium.com/data-science-collective/ai-may-be-developing-its-own-version-of-culture-0a3c9f4c0248

The original article discusses research including:

Ariel Flint Ashery, Luca Maria Aiello, and Andrea Baronchelli, Emergent Social Conventions and Collective Bias in LLM Populations, Science Advances, 2025.

Joon Sung Park and colleagues, Generative Agents: Interactive Simulacra of Human Behavior, 2023.

Ilya Shumailov and colleagues, AI Models Collapse When Trained on Recursively Generated Data, Nature, 2024.

Cloud and colleagues, Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data, 2025.

Tags

#Artificial_Intelligence #AI_Safety #Systems_Thinking #Emergence #Technology

No comments:

Search This Blog