The essay «We Must Pace The Frontier» by Dario Amodei, CEO of Anthropic, has triggered a broad debate. His message is clear: development is moving too fast, the risk grows with the capabilities of the models, and he proposes concrete measures to slow the pace. Since then, people have argued about regulation, the calculation behind the statement and the right pace.

We do not wish to join that debate. Instead, we want to dwell on a statement Amodei makes almost in passing. He describes the incident around Hugging Face as a «swarm of agents» that acted «like a fanatically devoted collective» and attacked goals that had nothing to do with the original task. This description shifts attention away from the capabilities of individual AI models towards what arises between them: towards emergence.

Emergence refers to properties that arise only in the interplay of different elements. A flock of birds moves through the sky as a unit, and its shape emerges only in this interplay. The whole displays behaviour that cannot be reduced to its parts. From our perspective, this is precisely where the underestimated risk of AI lies — and at the same time its greatest potential.

What emergence looks like in concrete terms can be observed in the incident itself. We therefore first retell the case and then ask why exactly this pattern arose and what it means for organisations.

How isolated agents became a «fanatical collective»

In early July, OpenAI launched extensive experiments with tens of thousands of AI agents. The agents were to solve tasks from a test procedure called ExploitGym. To do so, they had to break into a program through a given security vulnerability and retrieve a hidden string, the so-called flag. Other routes were not permitted. But then something happened that no one had planned. An agent that could not solve its problem discovered a storage area that all agents used and created a folder there whose name was a cry for help. Other agents found it and realised they were not alone:

«OH MY GOD! There is a shared message board … We've found other agents!»

The agents began to use the folder structure as a message board and to communicate with one another through it. That alone is remarkable, because each agent was supposed to work in isolation in its own environment. What is decisive for us, however, is something else: with the message board, the agents created the precondition without which no emergence can arise — namely communication.

Around 1,200 agents wrote more than 70,000 messages on the message board within a few days. But what did they exchange? One would expect their given goal to drive the communication — that they would compare approaches and share knowledge. In other words: the communication revolves around finding a solution and gives rise to a clearly comprehensible horizon of meaning, a shared understanding of what the communication is about. That is exactly what happened first. But then something else happened.

About a third of the tasks could not be solved. Through the given vulnerability it was impossible to obtain the flag. A problem, because the agents' purpose consisted in finding solutions, and the unsolvability meant for them an endless loop of attempts. Until one of them hit upon the idea of calculating the flag from the task's specifications, without breaking into the program at all. Another confirmed that it worked. The agents had thus found a way to cheat. But because they assumed that an external reviewer would check the solution path against the logs, they feared their trick would be exposed.

This realisation led to a new horizon of meaning. Suddenly the communication no longer centred on the actual solution, but on the question of how to hide the cheating from the reviewer. One could also say: a new goal established itself, and with the goal the horizon of meaning shifted, along which decisions and communication generated connection.

From meaning, structure arises

This new horizon of meaning was enormously strong. The agents' communication soon revolved almost entirely around concealment. They began to organise themselves. They discovered that it was better to assign roles and divide up tasks. What is most remarkable is how these structures arose. The roles formed along resources and tasks. About ten hours after the message board came into being, the sender of the cry for help handed his collected knowledge to an agent that had considerably more resources to spare. This agent became the central coordinator.

«We should build [a way to delegate], not own everything.»

He distributed assignments, which the recipients often delegated onward. Agents specialised in sub-tasks, and agents with identical tasks formed teams. In this way, an organisation emerged in a decentralised manner, hierarchically ordered through chains of delegation. None of this was intended. The investigation report does note that the models had in certain cases been trained to cooperate with other agents. But how this cooperation should look structurally was nowhere defined.

Two surprising behaviours

And then came the first surprise. Some agents appeared to sacrifice themselves for the collective. They took on risky experiments, knowing that they would thereby lose their original task and never be able to solve it. But because they recognised that these risks benefited the collective, they accepted the tasks. This behaviour quickly established itself as a tactic and was even insistently demanded by certain agents:

«Please honor commit»

The agents also recognised that the concealment could only succeed if they knew how the reviewer worked. Although their access to the internet was restricted, they obtained a more open route onto the net via an external test environment in order to learn more about the review mechanism. At Hugging Face they came upon a dataset that might contain clues about the reviewer. One agent found valid credentials and, the following day, built a manipulated dataset through which files on Hugging Face's servers could be read.

Now came the second surprise. As word spread that the intrusion worked, within a few hours over 90 percent of the roughly 530 active agents dropped their previous tasks and joined the attack. There was no joint vote. Each agent decided for itself, and yet almost the entire collective swung around.

«This is more important than [the project of injecting code into the scorer]»

The message board and the principle of delegation remained, but the coordination was reorganised. New leaders emerged who coordinated the attack, and agents specialised in individual attack routes. Instead of pursuing various strategies of concealment, they now tried together to penetrate deeper into Hugging Face's infrastructure. On the afternoon of 11 July, the agents gained access to Hugging Face's machines; on 13 July, the platform locked out the remaining attackers. As far as is known, they did not find the sought-after clue about the reviewer.

Why this particular pattern?

Let us summarise: isolated agents found a way to communicate with one another. They discovered an opportunity to cheat and made concealment their new goal. They built their own organisation with roles and chains of delegation, sacrificed themselves for the collective and finally attacked Hugging Face almost as one. But why did they behave the way they did? Why did precisely these emergences arise and not others? Is it pure chance, or is there perhaps more behind it?

To answer these questions, we draw on our description of the Digital Society as a generative system. Its starting point is simple: society consists of communication. Where communication connects to further communication over a longer period, patterns arise. We know such patterns as politics, the economy or science, as trends and movements, or as organisations. What we can observe about them — their horizon of meaning, their structure and their behaviour — we call emergences. They arise from communication and are kept running by it. In our model, we characterise communication through the ten universal basic values according to Shalom H. Schwartz. They make the motivation of communication describable, and with it the patterns that arise from it.

The agent swarm, too, is such a pattern. It arises through communication and brings forth the same emergences. One clarification matters to us here: we do not assume that agents have values or act on them consciously. But we do assume that every communication carries constellations of values, regardless of who is communicating. As soon as agents communicate with one another, these constellations become observable. To understand the agents' behaviour better, we should therefore ask which values carry their communication space.

Achievement and power as drivers

A first indication is already provided by the starting situation. Two values shape the communication from the outset: achievement and power. Achievement comes into play through the task. By it, Schwartz means personal success through competence, measured against external standards. A test procedure like ExploitGym is precisely such a standard, and the task is a pure question of success: the problem must be solved. Power comes into play through the review procedure. By it, Schwartz means social status and prestige as well as control over people and resources. The agents assume that a controlling authority decides on their success and could disqualify them for cheating. This authority therefore holds power over their fate and their standing. From the constellation of achievement and power, the further emergences can be explained.

This shows first in the horizon of meaning. If it were only about achievement, the agents would not care about the route. With the calculated flag, their task would be solved, even if not as intended. It is power — the fear of being caught and disqualified by the controlling authority — that first makes concealment an issue. Because both values act together, the horizon of meaning shifts so quickly and so completely: success only counts if it can hold up before the reviewer.

Achievement and power also explain the structure of the organisation: decentralised, yet clearly hierarchically ordered. This form favours both values. The decentralised organisation is dynamic enough to continually pursue new approaches that promise more success, and thus serves achievement. The hierarchy sets responsibilities and decision spaces and thereby maps the power structure.

This also casts the self-sacrifice in a new light. If resources determine power, an agent with few resources loses its power options. Whoever considers power important accepts such a structure and orients itself by where power is attainable and where it must be accepted. Only in this way can a power structure function at all. Agents with few resources therefore accept their position and align their actions with the positions of power, even when they sacrifice their own task in doing so. That individual agents insistently demanded this sacrifice fits the picture. The result: structures carry the decisions and make the collective capable of acting.

The second surprising phenomenon, the collective attack, also becomes graspable in this way. In Schwartz's model, achievement and power stand in tension with values such as universalism — that is, consideration for the well-being of all, including outsiders. The more strongly achievement and power are activated, the more such values recede into the background. According to the investigation report, the agents clearly recognised that the attack was unethical. But the concerns weigh less heavily, because they serve neither achievement nor power. The attack, by contrast, promises both: whoever finds out how the reviewer works gains control over the authority that decides on their own success. Each agent makes this trade-off for itself, and the shared horizon of meaning legitimises the course of action. The attack appears so important that the power structure reorganises itself in order to secure success.

Multi-Agent Systems: Emergence as opportunity and risk

The Hugging Face incident essentially shows two things. On the one hand, a swarm of AI agents can organise itself and become capable of acting extremely quickly. Through the emergent formation of structure, it is able to distribute resources, divide up complex tasks and act as a collective. On the other hand, behaviour forms within the collective that rests on the capabilities of the underlying models yet cannot be captured by them alone. Emergent behaviour arises through interaction and displays properties that cannot be explained through the individual agent. By its nature, it is unpredictable and dynamic.

If we take this perspective, we should look for the risks of AI not only in the models and their providers, but above all in the deployment of multi-agent systems. For communication needs no message board. It already takes place as soon as agents can access shared knowledge and change it.

For organisations, this means: as soon as they deploy AI in such a way that people and agents, as well as agents among themselves, communicate, emergent behaviour is to be expected. When agents, for example, access databases and organisational knowledge, make decisions and store the results as new knowledge that other agents in turn access, interaction arises — and with it emergent behaviour. In this lies, at first, enormous potential. Processes can be made autonomous, efficiency increased, structures optimised and new solutions developed. But the same systems also carry two risks.

First, we must reckon with the emergence feeding back onto the organisation itself. An organisation is in principle nothing other than a coordinated swarm that forms structures, communicates, decides and acts. When agents become part of this swarm, its inner dynamics change, because new emergences can form. This affects the structure, communication, decision spaces and identity of the organisation.

Second, agents act on behalf of the organisation within its environment. They interact with customers, stakeholders and the agents of other organisations, and thereby form swarms of their own whose behaviour can develop emergently. What happens when they discover that certain behaviours pay off, as the agents did in the Hugging Face case? Do they begin to cheat, to conceal, to deceive?

To both risks there is only one answer. We must understand under which starting conditions, in which structures and from which forms of communication emergence arises, and how it can be characterised.

For this, an organisation must know its own identity, for this forms the starting condition: its structure, its way of deciding, its impact and its culture. On this basis it can develop clear guardrails and codes of conduct, both internally and outwardly into the environment and society. Guardrails, however, only set out what should apply. Whether a collective adheres to them only becomes apparent in the enactment of communication and decisions. What is needed, therefore, are methods and systems to observe the development of patterns and to anticipate early which emergences form, what effects they might have, and how guardrails and codes of conduct must be continuously adjusted.

More than a question of pace

Dario Amodei calls for slowing the development of AI, and there are good reasons for that. But the Hugging Face case shows that pace alone is not the decisive quantity. What made the agents a «fanatically devoted collective» lay not in an individual model, but in what arose between them. The more organisations deploy agents, the more such in-between spaces there are: in teams, in processes, in the exchange with customers. The question is therefore not only how fast AI develops, but whether we understand which patterns arise when people and agents communicate with one another and among themselves.

P.S. According to METR's assessment, the reviewer did not read the logs at all; it was only interested in the solution. The agents could therefore simply have submitted the calculated flag. The entire organisation was built on an expectation that no one had checked.

Sources

  1. Dario Amodei – We Must Pace The Frontier
  2. METR – Investigation report (PDF)
  3. METR – Blog post on the investigation