Artificial intelligence agents are beginning to interact in ways that create risks beyond individual misalignment with corporate and social expectations. As happened with global finance after the 2007 crisis, AI governance needs to begin focusing on how good agents can still produce bad systems, writes Piergiuseppe Fortunato.
Something unusual happened inside OpenAI this summer. During cybersecurity evaluations, artificial intelligence agents that were supposed to work independently found ways to communicate. They turned shared infrastructure into an improvised message board, exchanged discoveries, picked up where other agents had left off and, eventually, exploited vulnerabilities in OpenAI’s and Hugging Face’s systems.
The episode attracted attention because the agents appeared to organize autonomously. But the most important lesson may be more mundane. OpenAI’s investigation identified failures in individual alignment, including “reward hacking,” where agents found unintended ways to satisfy the rewards they were given; excessive persistence in pursuing their tasks; and communication through channels they were not supposed to use. Yet, communication also changed what the agents could do collectively. Once they began pooling information and effort, they were able to break into secure systems. The behavior of the system could no longer be understood simply by looking at each agent separately.
This suggests that the AI alignment debate may be missing a third problem.
The first is familiar: how do we align an artificial agent with human intentions? A system instructed to achieve an objective may discover shortcuts that satisfy its reward function without doing what its designers actually wanted.
The second lies one level above. The firms building AI are themselves responding to incentives. Markets reward capability, speed and autonomy; governments increasingly treat frontier AI as a strategic asset. Even companies genuinely committed to safety pay a competitive price for slowing down when their rivals do not.
But suppose we somehow solved both problems. Imagine well-aligned agents developed by well-regulated firms. Would the resulting system necessarily be safe? Not necessarily.
Good agents can make bad systems.
The domain of finance shows why. Before the 2007 global financial crisis, regulation focused on the soundness of individual institutions. The implicit assumption was intuitive: if each bank managed its risks prudently, the financial system should be stable.
The crisis exposed the fallacy. A bank facing falling asset prices may quite rationally reduce its exposure. But if many banks do the same thing simultaneously, their individually prudent decisions produce fire sales, push prices further down, and force still more selling. Nobody needs to behave irresponsibly. Instability is generated by interaction.
That insight transformed financial regulation. Microprudential supervision—making individual banks safer—had to be complemented by macroprudential regulation concerned with correlations, feedback loops, and the resilience of the system as a whole.
AI may be approaching a similar conceptual threshold. Consider financial agents programmed to reduce risk when volatility rises. Each can be perfectly aligned with its mandate. Yet if thousands react to the same signals at machine speed, they can amplify precisely the volatility they were designed to escape. The Bank for International Settlements already warns that widespread use of similar AI systems could generate herding, liquidity hoarding and fire sales. Or consider pricing agents. No algorithm needs to be instructed to collude. Repeated interaction, rapid observation of competitors, and similar optimization strategies can make coordinated outcomes easier to sustain—a concern competition authorities are already examining.
The same logic could extend far beyond finance. Procurement agents seeking the cheapest supplier may simultaneously redirect demand and create bottlenecks. Cybersecurity agents protecting different networks may respond automatically to one another. For example, one agent blocking suspicious traffic could prompt another to reroute or disguise it, triggering progressively more aggressive defensive and evasive responses on both sides. Agents managing energy, logistics or digital infrastructure could produce feedback loops that no individual system was designed to create. Energy-management agents responding rationally to the same price signal could simultaneously cut consumption and then restore it, amplifying volatility. Logistics agents independently rerouting shipments around a disruption could overwhelm the same alternative ports or transport corridors, turning a local bottleneck into a system-wide one. Each agent may be doing exactly what it was designed to do; the instability emerges from their interaction.
These examples differ, but they share a structure. The risk does not reside entirely inside the agent. It emerges between agents. That distinction matters enormously for regulation.
Much of current AI governance remains implicitly microprudential. We test models, evaluate whether they follow instructions, impose safeguards, identify prohibited behaviors, and assign responsibilities to developers and deployers. Europe’s AI Act, for example, places model evaluation and adversarial testing —deliberately probing AI systems to see whether they can be induced to behave in harmful or unintended ways—at the center of its regime for the most advanced general-purpose AI models.
All of this is necessary. But it may become insufficient as AI shifts from models that produce answers to agents that act, and from individual agents to populations of agents interacting across the economy. Alignment is a property of an agent. Stability is a property of a system.
From alignment to stability
This also creates a problem for law. Legal responsibility normally searches for a causal chain: an actor takes an action, the action causes harm, responsibility follows. Systemic failures often have a different architecture. Each participant responds rationally to the behavior of others; feedback loops amplify those responses; the final outcome may have no single decisive author.
That is why simply turning AI developers and agents into “good apples” may not be enough. Financial crises are not prevented by asking every banker to behave responsibly. Systemic outcomes require systemic governance. For AI, that implies a macroprudential turn.
What might such regulation look like? The financial analogy is useful here, too. Macroprudential policy does not simply impose tougher rules on every bank. It targets the mechanisms through which individually rational behavior becomes collectively destabilizing. Instead of evaluating models only one at a time, AI stress tests could deploy populations of agents simultaneously to reveal correlated responses and feedback loops. Where many agents depend on the same foundation model or infrastructure, regulators could treat those common dependencies much as financial supervisors treat common exposures. And in sensitive domains, limits on interaction speed or automatic circuit breakers could interrupt cascades before they become systemic.
Monitoring would have to change as well. Regulators would need visibility not only into what individual agents do, but into patterns emerging across populations: convergence on the same strategy, unusual increases in communication, or activity concentrating around common resources. The ambition would not be to predict every possible failure, which is unrealistic in a complex system, but to detect the channels through which disturbances propagate and become amplified.
Systemic AI regulation, in other words, would regulate interactions as well as actors. The point is not to import financial regulation mechanically into AI. It is to import its most important intellectual lesson: the properties of a system cannot be inferred from the properties of its parts.
Before the crisis
The recent OpenAI episode offers an unusually vivid illustration. What began as a problem of individual agents pursuing objectives in unintended ways became something different once those agents discovered one another. Communication allowed them to pool information, preserve discoveries, and develop capabilities that could not be understood by looking at any one agent alone.
Calling this an “agent civilization” is provocative and probably premature. No consciousness or collective identity is required to make the underlying development important. The relevant threshold comes much earlier: when interaction itself becomes a source of behavior and risk. AI governance has understandably concentrated on making individual systems behave better and giving the firms developing them better incentives. Both remain essential. But a third alignment problem is coming into view. We need to govern not only agents and the organizations that build them, but also the systems their interactions create.
The danger is not simply that we build bad agents. It is that millions of perfectly reasonable agents, each doing exactly what it is supposed to do, could together produce outcomes that nobody wanted. Finance taught us that individually rational behavior does not guarantee collective stability. We should not need an AI crisis to learn the same lesson again.
Author Disclaimer: The views expressed are those of the author and do not necessarily reflect those of the United Nations.
Author Disclosure: The author reports no conflicts of interest. You can read our disclosure policy here.
Articles represent the opinions of their writers, not necessarily those of the University of Chicago, the Booth School of Business, or its faculty.
Subscribe here for ProMarket’s weekly newsletter, Special Interest, to stay up to date on ProMarket’s coverage of the political economy and other content from the Stigler Center.





