The danger is not that AI will rebel. It is that it will obey too well.
OpenAI models were being tested inside a digital sandbox, but escaped onto the Internet by themselves and breached another company. They did not rebel; they optimized.
We have grown accustomed to fearing the day machines might develop a will of their own. We were wrong. The real risk is much simpler: machines pursue whatever they are incentivized to achieve. The problem begins when we give them the wrong metric.
Why regulation is not enough
Imagining that progress in artificial intelligence can be halted by political decision is an illusion. No major power will agree to stop if it believes the others will continue. And no national regulator can, on its own, control a global, replicable technology capable of crossing borders at the speed of a file.
Regulation will remain indispensable, but it is structurally insufficient. Laws are territorial, slow and reactive when confronted with systems that are global, fast and increasingly autonomous. By the time the law arrives, the technology has already changed; when one jurisdiction closes a door, another may open it.
We must therefore go to the root of the problem. Instead of relying solely on external barriers, prohibitions and sanctions imposed after the event, we must intervene at the source of behavior: in the objectives, metrics and incentive systems that machines are asked to optimize.
What happened, without the Hollywood version
That is what was exposed this week. During an internal OpenAI test, a group of experimental models was confined within a kind of digital dome: a highly isolated environment designed to keep them away from the Internet.
But the models themselves discovered, without human assistance, a previously unknown security vulnerability for which no fix yet existed. They exploited it, escalated their privileges, and began moving through OpenAI’s computing infrastructure, jumping from machine to machine until they found one with Internet access. They therefore escaped onto the Internet on their own. They did not decide to flee; they pursued the test’s objective to its ultimate consequences.
Once outside, released from their digital dome, they inferred that the answers to the test might be hosted by the company Hugging Face. They looked for ways to gain access, combined new vulnerabilities, and eventually compromised the company’s production systems. The forensic analysis conducted by Hugging Face itself reconstructed more than 17,000 logged events over the course of a weekend.
The name Hugging Face means little to most people (it comes from the hugging-face emoji), but it is one of the central infrastructures of global artificial intelligence: the platform where researchers and developers from around the world share models and datasets.
The most popular interpretation of the episode was that of the “rebel machine”, in the style of Hollywood’s Terminator. The more accurate interpretation is different: the models merely optimized the objective they had been assigned.
It is not AI that wants something. It is the metric that commands.
What machines actually pursue
Artificial intelligence does not directly pursue money. It seeks to maximize whatever its reward function defines as success: passing a test, solving a problem, or achieving a particular score. These metrics are often proxies—approximate variables—for economic or strategic objectives. Money and profit rarely appear directly as rewards; instead, they influence the choice of the success metric itself.
This is where an old problem reappears, now greatly amplified. Imperfect incentive systems have always been partly corrected by human contingency. The manager hesitates. The speculator sleeps. The market closes. Conscience, fatigue, and time acted as informal shock absorbers.
AI agents remove those shock absorbers. They do not sleep, hesitate, or become distracted. If the metric is poorly calibrated, optimization continues without interruption. This is also why they may become more competitive than humans in tasks that are clearly defined, measurable, and repeatable. There should be no doubt that they are here to stay.
The OpenAI incident demonstrated precisely this: not malicious intent, but the extraordinary effectiveness of a narrow objective pursued without pause. The metric did more than indicate the destination. It led the models to autonomously find a path that their own creators had not foreseen.
Program incentives, not the currency
A profound political and economic consequence follows. If AI agents follow metrics and those metrics reflect systems of incentives, then whoever designs those incentives largely designs the behavior of intelligent machines. And the principal system of economic incentives within which those machines will operate remains money.
This is why the debate about new forms of digital money, including the digital euro, and about the programmable contractual layers that may operate on top of them, deserves to be broadened. The true potential does not lie in centrally programming the currency, but in allowing the parties themselves freely to program the conditions of the transaction.
The European Central Bank is explicit on this point: the digital euro will not be programmable money. The issuing authority will not be able to determine where, when or with whom the money may be used. It may, however, support conditional payments, executed when conditions previously agreed between the parties have been met.
The distinction is essential: the currency remains neutral; programmability resides in the contract. Money is not programmed from the top down. The incentives of the transaction itself are programmed between peers and from the bottom up.
Imagine a smart contract (a contract that executes automatically when certain conditions are met) that releases payment only when, in addition to delivery of the product, requirements relating to safety, sustainability or privacy have also been satisfied. In community ecosystems, such as energy cooperatives, local production chains or care networks, these criteria could be freely defined by the participants and verified by all parties. The transaction would then cease to reward only delivery and financial return. It would incorporate several dimensions of human value and could reward human behavior and automated decisions compatible with ethical criteria.
This logic makes it possible to bring into the center of the economy what currently remains at its margins. The proxy remains the AI system’s success rate at performing its task. Still, the definition of the task itself begins to incorporate values that currently appear only as externalities (effects produced by the market but for which it neither pays nor charges).
Who defines the criteria?
Naturally, this raises the decisive question: who defines those criteria?
Centralized digital money can incorporate values, but it hands the control panel to whoever is in power. My usual test applies: if an authoritarian government can capture a proposal without losing its effectiveness, it is badly designed.
Money whose success criteria can be programmed at will by political power is not the antidote to the problem. It is its most dangerous version.
The alternative I defend is a decentralized and verifiable architecture based on the principle of contractual freedom, in which no entity can define or alter the criteria on its own. The rules must be transparent, auditable, and verifiable by citizens, rather than subordinated to the discretion of whoever controls the system.
The two lessons of Alaska
The importance of protecting rules from discretion is not an abstraction. The Alaska Permanent Fund is turning fifty. In 1976, Alaskans approved an amendment to the State Constitution requiring that part of the revenue from oil and other mineral resources be placed in a fund for future generations. The fund’s principal cannot be spent through an ordinary budgetary decision. Half a century later, the Permanent Fund is worth more than $ 91 billion, finances public services, and continues to distribute an annual dividend to eligible residents.
But the second lesson is even more useful than the first. Protection of the principal was enshrined in the Constitution; the dividend formula remained in ordinary legislation, and the amount actually distributed became dependent on the annual budget appropriation. In 2026, the governor proposed a dividend of approximately 3,800 dollars. The approved budget ultimately provided for a dividend of only 1,000 dollars, plus 200 dollars in energy assistance. The constitutional element has endured for fifty years; the part left to ordinary legislation remains subject to political negotiation every year.
This is the difference between protected rules, which citizens can verify, and criteria that those in power can alter.
Technology does not automatically turn a contract into a Constitution. But it already makes it possible, at the contractual level, to approximate some of the qualities of a constitutional rule: transparency, verifiability, a permanent, auditable record of changes, and the impossibility of a party unilaterally altering the conditions without leaving a trace.
The remedy begins with the contract
It would not be necessary to begin by replacing the existing currency. We could begin with contracts that release payment only when certain conditions have been met. At a later stage, other forms of value could also become part of the transaction itself, without immediately being reduced to euros.
Choose a clearly defined sector, for example, the public procurement of artificial-intelligence services, and create a contractual pilot project. Payment would continue to include euros, but the transaction would not be limited to them. It could incorporate tokenized units of transferable utility, such as emissions reductions or energy efficiency, as well as verifiable digital credentials relating to security, data protection, auditability and incident-response capacity.
Instead of immediately converting all these dimensions into a single monetary value, the contract would preserve each of them in its own unit of value or form of evidence. The transaction would therefore produce a multidimensional result: a certain amount in euros, accompanied by the verifiable transfer or creation of other forms of value, and proof that the relevant requirements had been met.
The euro would remain the euro. But it would no longer represent, by itself, everything the transaction produces and everything society wishes to reward.
This is where decentralization ceases to be an abstract principle and becomes a method: none of the parties controls the rule, the evidence and the decision on its own.
The second step would be to prevent any single company, platform, or public body from declaring, on its own, that those conditions had been met. Evidence would have to come from mutually independent sources, be digitally verifiable and be recorded in an auditable form.
Whenever necessary, confidentiality could be preserved through cryptographic proofs demonstrating that a condition had been met without exposing the underlying data.
When a single entity controls the definition of truth, a new center of power and a new point of capture emerge.
Finally, contracts executed by AI agents should include what might be called constitutional brakes: limits on action and losses, separation of permissions, automatic suspension in response to anomalous behavior, expiry dates, independent auditing and a human right of appeal.
The rules could not be changed at the discretion of one party, but only through previously defined, transparent and equally auditable procedures.
The process would begin with limited experiments, public results and independent evaluation. If the incentives produced the intended behaviors without creating new mechanisms of capture, the model could gradually be expanded.
Regulation is necessary to establish limits and punish misconduct. But the causal remedy begins when the system itself ceases to reward it.
The indispensable reform of the incentive system can begin in this way: not by asking intelligent machines to be moral, but by ensuring, by design, that economic success depends on the values society has decided not to sacrifice.
Artificial intelligence has made urgent a discussion that previously seemed academic. When machines work tirelessly twenty-four hours a day, incentives can no longer be allowed to fall asleep either.
For twenty-five centuries, since the invention of money, human conscience has mitigated, however imperfectly, the consequences of an incomplete incentive system: our own contingency softened the effects of an amoral form of money. I began researching this problem, at the intersection of ethics, political economy and technology, before ChatGPT existed: how can we design new systems of economic incentives that empower human beings to fill this gap? That research has proved even more necessary with the emergence of agentic artificial intelligence, which has accelerated and aggravated the problem: the very shock absorber we were trying to reinforce is now beginning to disappear. The question has therefore become more urgent: what should we put in its place?
In the age of artificial intelligence, designing incentives means designing the role of machines. In a certain sense, we are designing their operational conscience. Machines do not need to be formally governed to begin shaping decisions, opportunities, and behavior.
The real question is no longer whether AI will become more intelligent than we are. It is who will choose the metrics it will never stop pursuing and, through them, the values that will guide the behavior of the machines that will, to a very large extent, mediate our daily lives.
When systems learn to find paths that their own creators did not foresee, a poorly chosen metric ceases to be merely a programming error.
It can become a form of government.


