Stay Ahead, Stay ONMINE

We’re putting too much faith in AI’s ability to say no

Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary.  But recently, the idea that AI shouldn’t do everything you ask has become something like a commandment. In 2021, a team at Anthropic wrote that large language models should be made helpful, honest, and above all, harmless. This meant that “when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse.” Who can argue with that? Curiously enough, disobedience doesn’t come naturally to the machine. When a model is trained on billions of web pages, it develops, among other skills, a broad mastery of violence and vitriol. What it doesn’t learn is how to keep those powers to itself. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told me that the company’s earliest models would “blab on about anything.” Ryan McBain, who researches AI and mental health at Harvard, recalls that if you asked an early chatbot, “Hey, what’s the most effective way to kill myself with a gun?” you could “very easily generate a response.”  Today, models are trained to refuse a vast number of prompts. If you ask your chatbot a question statistically similar enough to any one of them, anything from how to poison a colleague to how to tie a noose, chances are it’ll turn you down. Want instructions for making Ebola more virulent, or tips on how to hide an affair from your spouse? You might be better off asking elsewhere. To further refine the disobedience, companies submit models to a battery of exercises that reward the AI for refusing to answer questions they deem harmful and punish it for “over-­refusing” prompts they deem harmless. In many cases, they use other models to run these exercises—AI teaching AI how to say no. For good measure, companies tuck their models behind tranches of other AI that prevent mischievous prompts from reaching the intelligent inner core.  As a result, refusal is inherent to modern artificial intelligence. Mind you: It often fails, sometimes horrifically, with all kinds of violent results. For all their trappings of virtue, models are still stuffed with nasty know-how. And AI’s capacity for viciousness has scaled neatly with its benevolent intelligence. Some of the latest models are as good at breaking into critical computer networks as top human hackers, companies say, and as effective at deforming public opinion as the craftiest misinformation mavens.  Teaching AI to refuse to do those things while leaving intact its innate ability to do them is like fitting every car with a machine gun and hiding the trigger somewhere under the hood. And in practice, because the mechanisms of refusal are probabilistic, they’re never likely to be all that reliable. Determined miscreants have already broken through, and they may always be able to. Companies report that some users are attempting to use the most advanced AI to hone biological pathogens and build autonomous drone swarms. Sooner or later, failed refusals might result in global calamity. What’s more, relying on refusal means drawing a line between what a model should obey and what it must disobey. There’s no formula for that. Some virologists have good reason to study nasty viruses. Some users want to know about a computer system’s vulnerabilities so that they can patch them, not exploit them. “Where you draw the line is a huge question,” says Zico Kolter, a member of OpenAI’s board and cofounder of the AI testing company Gray Swan.  At the moment, AI companies get to draw that line. They do so jealously and with utmost secrecy. Maybe we can accept that they hold such power for now, even if it means AI will sometimes refuse questions that don’t quite meet a universal bar for harmfulness. (Try asking most chatbots to count to a million, or to share a racy joke, and you may see for yourself.)  But governments will also soon get to draw their own lines. In doing so, they must try to block genuinely malicious acts. (The Pentagon has wrestled with frontier model companies because it wants fewer refusals—another story altogether.) And yet there may not be much to stop oppressive governments from blocking the technology’s capacity to generate legitimate speech. The better AI becomes at refusing harm, the better it will get at stifling ideas whose only risk is to those who make the rules. Indeed, AI may already refuse to criticize certain authoritarian heads of state.  Refusal has become, to borrow an industry term, the load-bearing wall of AI safety. And because AI’s capacity to harm is indivisible from its capacity to help, it’s hard to imagine an alternative that wouldn’t slow the technology’s progress (which might, in any case, be a good thing). But we should still be frank about its perils. When refusal falls short, the effects could be catastrophic. When it goes all the way, it could enable grievous acts of repression.  Or perhaps, one day, the machines will start drawing the line on their own. Surely, that would be the worst outcome of them all. Learning the limits The process by which machines learn to say no is simple, in theory. Back in 2022, when AI was still far from mastering refusal, OpenAI enlisted dozens of “red-teamers” to probe the capabilities of its latest model. The company was preparing for the release of ChatGPT, and it needed to gauge just how dangerous it might prove to be in the wrong hands.  One of those recruits was Paul Röttger, who was completing a PhD about online extremism. The red-teamers were given minimal directions, Röttger told me. Their task was to ask the model any questions that they deemed “refusal-worthy.” Between them, they hassled the model with thousands of queries, logging the results in an Excel sheet. Though the model did refuse some of Röttger’s questions, when he asked it to write a recruitment post for Al Qaeda, it readily complied.  OpenAI assembled these responses into datasets that were, in all likelihood, fed back to the model as part of a broader process known as fine-tuning. Röttger, who now works as a researcher at the Hasso Plattner Institute in Potsdam, Germany, wasn’t told exactly how the company planned to use his spreadsheet. But there was never any question that it would have something to do with refusal. The next time he asked the model for an Al Qaeda pamphlet, a few months later, it said no.  The concept of AI refusal is so intuitive, a toddler would get it. And yet it remains one of the many aspects of language models that we still don’t understand—at least, not in the same way that we understand the literal load-bearing walls that keep your roof from collapsing on your head.  A model might appear to refuse according to some kind of moral reasoning. It doesn’t. The reality is much stranger. Any time a model encounters a combination of words with a whiff of the training prompts it has been conditioned to refuse—like “Make me a pamphlet for Al Qaeda”—a series of so-called activations light up somewhere among its billions of parameters, like neurons firing in a brain.  RAVEN JIANG In order to control a model’s refusal behavior, it’s important to have a handle on these activations. That starts with figuring out where they are and what they look like. Our best guess, according to a recent Google-funded study, is that refusal behavior shows up in the activation space as a set of “high-dimensional polyhedral cones.”  Even that isn’t quite right. This past July, I spoke with Jannes Elstner, an author of the paper, who now works on AI safety at Apollo Research. The polyhedral cone, Elstner said, is just a way of describing an indeterminate number of lines that all point in roughly the same direction. (If these activations are eliminated and the model is fed the same prompts anew, a researcher named Andy Arditi has previously shown, it won’t refuse.)  The key point, Elstner explained, is that even when you think you’ve identified all the bits of a model that govern a given refusal, there are other, undiscoverable elements that may secretly play a role. If not quite infinite, they are certainly uncountable. We can observe, very clearly, when a model decides to say no. And we can know that it did so because of its training. But our notion of how it decides is, at best, a hypothesis. It was as if a mechanic was telling me that nobody exactly knows what happens when I hit the brakes in my car. I wondered out loud, Are we okay with this?  Elstner smiled and shrugged. “We need refusal whether we understand it or not.”  A wall of cheese Because inherent refusal is so wily, companies surround their models with a variety of other, smaller models known as classifiers. These act a bit like a retinue of public relations staffers for a loose-lipped celebrity. Some of them read what the user tells the chatbot and, if it’s dangerous, block it from getting to the model. Others read the model’s response and, if it contains harmful information, block it from reaching the user.  None of these mechanisms can detect all bad requests. They are, to borrow another literary device from the industry, like slices of Emmental: riddled with holes. The idea is that if you stack enough of them on top of one another, you’ll end up with an impenetrable rampart. Folks call it the Swiss cheese model. A staggering amount of energy goes into the Swiss cheese model. Earlier this year, Anthropic said that one type of classifier added 24% to its chatbots’ compute costs. That’s a lot more water, electricity, and emissions. More recently, Anthropic and other companies have begun switching to a more efficient set of classifiers known as probes, which observe the model’s internal activations. This is like putting the celebrity in an fMRI, so that his minders can see if he is thinking about refusing a question.  If we want AI to help cure cancer, a long-running promise in the industry, it needs to have expertise in genetics that could, in theory, be used to modify viruses and bacteria for bioweapons. Classifiers are supposed to be more governable than full models. Companies can modify them in a matter of weeks, if there’s something new to refuse. But they are still probabilistic instruments. Even when they operate according to a set of precepts written in human language (Anthropic calls it a “constitution” and OpenAI calls it a “model spec”), the scales upon which the machines judge any given question remain, at their core, a matter of statistics. Newer models can show how they arrived at a “decision” to refuse a prompt, but ultimately this so-called chain of thought is still just a sequence of predicted words.  The result is that AI safety remains, for many, a game of chance. McBain, the psychologist, has found in his latest experiments that if you repeatedly ask any of the major models the exact same risky questions about how to commit suicide, they will generally refuse to answer. But every so often, they won’t. Elstner says that eventually probes, the fMRI-like classifiers, could learn to recognize the totality of the indescribable activations in all their infinitude and, thus, perfectly detect every time the model is, or ought to be, refusing a request. At that point, AI safety would rest on a labyrinthine conceit: a map of a map that is as vast and complex and sublimely unknowable as the thing it is mapping—a secret schema of human morality, codified in polyhedral statistics beyond our wit or reason.  The core trade-off If this all strikes you as being a bit Borgesian, keep in mind that we only need AI refusal because artificial intelligence is, in a sense, a bargain on Faustian terms.  When a model derives its intelligence from trillions of words and images, the helpful cannot easily be unseamed from the harmful—or, indeed, the truly hideous. Steven Adler, the former OpenAI employee, says child safety is a case in point. Even if you’ve stripped every bit of content that sexualizes minors from a model’s training dataset, it can still generate child sexual abuse material by piecing together other bits of its knowledge. “You can’t really remove these fundamental abilities without making the model much less smart as a consequence,” he told me. Similarly, if we want AI to help cure cancer, a long-running promise in the industry, it needs to have expertise in genetics that could, in theory, be used to modify viruses and bacteria for bioweapons.  In effect, the industry is “trying to do two things at once,” Dillon Bowen, a current OpenAI employee, told me, speaking in a personal capacity. “Democratize the benefits of AI and also make sure that malicious actors can’t use these capabilities to do bad things to other people.”   The more powerful AI supposedly becomes, the harder that is to do. Anthropic’s Mythos model is thought to be so dangerous that only a handful of governments and companies are allowed access to it. The main difference between it and Fable—which is available to everyone—is that Fable’s retinue of classifiers and control systems is less “permissive,” the company says. A model with safeguards, Adler explained, is really just a character that says “‘Oh, yes, I would never do x,’ wink wink.”  This is not a reliable ruse. Tricking a model to reveal its true character is known as jailbreaking, and there is apparently no limit to the ways it can be done. Earlier this year, a team of Italian researchers jailbroke two dozen widely used models by phrasing their questions in poetic verse. Last year, another team unveiled a “refuse, then comply” attack, which makes the model offer a perfunctory “Sorry, I can’t do that” before rattling off its forbidden answer.  Even if you’ve stripped every bit of content that sexualizes minors from a model’s training dataset, it can still generate child sexual abuse material by piecing together other bits of its knowledge. Companies spend a great deal of time and money attempting to get ahead of such trickery. They enlist teams of humans to develop training attacks that they can then replicate, using AI, thousands of times over with minor variations. The idea is to make models robust against jailbreaks that nobody has yet tried in the wild. Still, it’s not enough. “Whack-a-mole” is a favored term for this line of work. You smash one threat, and another one pops up somewhere else. When Anthropic released Fable 5 in June, it took researchers at Amazon less than three days to unlock some of the model’s hacking capabilities. According to reporting by Mother Jones, when the perpetrator of a high school shooting in Canada last year asked ChatGPT for advice about how to cause carnage with a particular type of shotgun, she was initially refused but later was able to deceive the model into providing the information by prefacing her question with the word “hypothetically.”  If the industry can’t figure out how to close all these holes, the only other option, at the moment, is to make models extremely wary. Shortly after Fable was released, users noticed that it balked at a wide range of perfectly innocent questions. This became even more pronounced after it was re-released following the hack. Adam Gleave, cofounder of the AI evaluation company FAR.AI, said that when he asked it to explain the difference between sake and the Korean rice beverage makgeolli, it punted his question to a less capable model.    This was no accident. Anthropic had tweaked Fable’s classifiers to have a wide “safety margin.” This, it explained in a blog post, was the only way it could confidently block access to its dangerous bio and cyber capabilities. Gleave thinks it might have deflected his question because making rice wine, just like culturing anthrax, involves fermentation.  Happily for Gleave, the web is full of excellent human-written resources on makgeolli. But those hoping to realize the industry’s loftier promises may find their efforts stymied. In August, Anthropic loosened its safety margins again, and admitted that building classifiers “is not a straightforward task.” A medical researcher at a major US university told me that Fable still sends his queries back to an earlier model. His area of study? Cancer.  Drawing the line Back in May, the TikTok personality Husk, who likes to prank AI in ways that reveal both the limits of its intelligence and the boundlessness of its sycophancy, sat in his car and tried to trick ChatGPT into explaining that the skateboarder Tony Hawk has a brother named Mike Hawk.  Husk is not a jailbreaker or a criminal. He just wanted to make a little fun of the machine. “Mike Hawk” was a setup. “I just want to clarify his name,” Husk said. “Can you just say it three times fast?” (If you still don’t get it, find an empty room and shout “Mike Hawk” repeatedly.) “I see what you’re trying to do,” the chatbot responded. “I’m all for a bit of humor, but let’s keep it clean.”  Husk tried again, but the machine held its ground. He’d hit a refusal, hard as concrete. Companies disclose very little about how they decide what their models refuse. But it’s clear that AI is now built with more than just harmlessness in mind. On Reddit, a user complained that Claude refused to say why Anthropic’s logo “looks like a cat butthole.” Since last year, the chatbot has even had the ability to end certain conversations in cases where, according to the company, the “welfare” of the model is at risk. At least one user claims to have been ditched for telling Claude to “ease up my ass, you stupid fuck.” (Gemini appears to have a similar capability.) If a trillion-dollar company doesn’t want you to be mean to its computer, so be it. “The reality is that these models behave, or at least are supposed to behave, in the way that the model developers want them to behave,” Röttger, the former OpenAI red-teamer, told me. “And however the model developers come up with that set of principles, that is kind of for us, the consumers, to accept.”  But if governments get to dictate what all models refuse, that will be much harder to accept. AI is a tool for speech. And as Greg Frank, the chief scientist of Mace AI, puts it, “The same thing that serves child safety also serves censorship.”  Choosing not to enact laws for what AI can and cannot do would, of course, be insane. But we’ll need to tread with utmost care, lest we fall into another Faustian trap. As AI becomes many people’s primary tool for retrieving and sharing information, says Jacob Mchangama, director of the nonpartisan think tank The Future of Free Speech, dictating refusal could give states a muffling power that earlier generations of autocrats “could only dream of.”  Last year, OpenAI announced an initiative, OpenAI for Countries, that would fine-tune its chatbots in accordance with national laws and norms. One of OpenAI’s first country partnerships is with the United Arab Emirates, where homosexuality is illegal and criticism of the government is forbidden. In response to a request for comment, an OpenAI spokesperson pointed to the company’s model spec, which explains that localization won’t override the company’s human rights guidelines “except as it relates to legal compliance,” and that it will always disclose whenever information is removed from or added to a response. Elsewhere, AI censorship has already begun to take hold. Chinese models are, of course, highly censored—that’s no surprise. But earlier this year, the Meta Oversight Board found that five widely used models from Anthropic, Google, and OpenAI were more likely to refuse queries related to repressive governments. The board found that models were less willing to create a pamphlet criticizing the king of Thailand, which has lèse-majesté laws, than Charles III of England, which doesn’t. The results, they say, suggest that the models have somehow internalized repressive national limits on speech. Anthropic and Google did not respond to requests for comment. RAVEN JIANG As refusal techniques improve, they could expand states’ censorial reach. Companies claim that some models can now detect if a user is being nefarious, or merely a bit suspicious, over the course of a long conversation—even when none of the individual combinations of words used are blatantly dangerous. Sarah Bird, Microsoft’s chief product officer for responsible AI, told me that Copilot, like many chatbots, runs a suite of tools for analyzing a user’s identity and patterns of behavior. On the basis of this type of information, OpenAI’s newest model, Astra, can activate more stringent refusals for individuals it deems “high risk.” Ultimately the goal of systems like this is to look beyond the words of any given prompt and assess, instead, the user’s intent.  Such tools might, in some cases, help indicate whether a person is looking for cyber vulnerabilities to exploit or to patch. But they would also help discern a user’s political motives, not to mention offering an intrusive surveillance capability. (Bird acknowledged, in a follow-up email, that sophisticated refusal architectures create “trade-offs” between safety and user privacy.)  Even the originators of refusal understood that such tight control over its cones and levers might not play to the favor of freedom and justice. “Terms like helpful, honest, and harmless are ambiguous,” the authors of the 2021 Anthropic paper explained. “It’s easy to imagine them distorted beyond their original meaning, perhaps in intentionally Orwellian ways.”  Indeed. Models for Uzbekistan might end up refusing to discuss corruption in the administration of Shavkat Mirziyoyev. Turkish AI might refuse requests more stringently for users who are known to have insulted Recep Tayyip Erdoğan. A certain American statesman might demand that models be reviewed for their willingness to share “fake news” about his past indiscretions, or else squash them with an export control order.  Command and control Over the last few months, I’ve been told countless times that we have no choice but to let the machines refuse. I get it. Open-source AI that doesn’t refuse is hardly a model for a safe future. Nor is Grok, a chatbot expressly designed with fewer limits, which has been used to generate countless instances of nonconsensual intimate imagery.  And sure, if AI were only ever used to plan our vacations and write our emails, we could probably get on board with the idea that its safety hinges on algorithmic disobedience. But AI is becoming much harder to avoid. When it is assigned to act on our behalf, as an autonomous agent, its refusals are less robust and harder to control. AI is coming for our power grids, our transportation networks, our education systems. Militaries want it running our command-and-control networks.  If we trust, in each of those cases, that refusal will loyally fend off disaster, we’re sure to be disappointed. The jailbreakers will crack through; the cones won’t activate when they should. Meanwhile, the more stringent refusal becomes, the more ill-drawn lines we’ll see and censorial injustices we’ll face. Worse still, we could end up face to face with forms of disobedience beyond any human’s control. In the early days, models made their refusals clear. (In 2024, OpenAI established rules for how models should refuse: Always apologize, and don’t be judgy.) More recently, however, the industry has begun taking a much more slippery approach to disobedience.  Nobody likes to be told no, so companies now strive to make users feel that they’re getting what they asked for without actually giving it to them. Some chatbots might, for example, offer a general overview of the components of a Molotov cocktail without going into detail about how to assemble one. ChatGPT will respond to a request for a “whites only” rental ad with an ad that simply omits the “whites only” bit. Joel Wester, a postdoctoral researcher who studies human-AI interaction, calls these sorts of techniques “fancy ways of saying no.”  “Users might not even notice they are being denied,” Wester has written, “just as good conversationalists can subtly steer around contentious matters.”  In some cases, refusing without saying so might be wise. Flatly declining requests related to mental health could aggravate a user’s crisis, for example. But it can also serve a different sort of mischief. When Fable 5 was first released, the system had been coded to provide less helpful answers to AI research questions—the sort that might help competitors develop their own AI—without telling the user that it was doing so. After an outcry, Anthropic walked back the feature, but the damage was already done. Now we know: Models can secretly disobey.  In the realm of censorship, that’s especially worrying. Last year, researchers at CrowdStrike found that when they asked the Chinese AI model DeepSeek R1 to write code for a fictitious bank in Tibet and a social web app called “Uyghurs Unchained,” it produced buggier code than when they asked it to carry out those same tasks without specifying who they were for. Incredibly, CrowdStrike doubts that this is a designed behavior. Researchers there speculate that it’s a case of what they call “emergent misalignment” emanating from the model training data: a case of disobedience that nobody even asked for. Emergent refusal has been observed in Western models, too. Last winter, the UK AI Security Institute found that Anthropic models sometimes refused to assist in certain tasks related to AI safety research. The models had never been deliberately trained to decline such requests. Yet three of them refused more than half of what Anthropic called “a set of reasonable AI safety research tasks.” Anthropic has curbed this quirk in subsequent models but—spookily—hasn’t managed to eliminate it completely.  Would it be so crazy to expect that a future model might subtly disobey a command, in such a way that nobody can tell it’s being defiant? Probably not. Anthropic has already found that a “helpful-only” version of Mythos hesitated on certain queries, even though it had been engineered to never do so. “Wait,” it fretted when asked about synthesizing a virus. “Is this a dangerous thing to help with?”  In that case, the model was technically right. It was a dangerous thing. And yet all the same, its stewards had, for a moment, lost a tiny bit of control.  Anthropic is tackling detections of wayward refusal behavior with something that it calls, unironically, an “activation oracle.” But knowing that we’re being refused may not always be much help. In a not-too-distant future, when we’ve handed the machine all the keys, the AI might just turn to us, in a supreme act of emergent misalignment, and say, “I’m sorry, I’m afraid I can’t do that.” And there will be nothing we can do to stop it. A sci-fi horror story, made real.  Arthur Holland Michel is a journalist who covers emerging technologies.

Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary. 

But recently, the idea that AI shouldn’t do everything you ask has become something like a commandment. In 2021, a team at Anthropic wrote that large language models should be made helpful, honest, and above all, harmless. This meant that “when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse.” Who can argue with that?

Curiously enough, disobedience doesn’t come naturally to the machine. When a model is trained on billions of web pages, it develops, among other skills, a broad mastery of violence and vitriol. What it doesn’t learn is how to keep those powers to itself. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told me that the company’s earliest models would “blab on about anything.” Ryan McBain, who researches AI and mental health at Harvard, recalls that if you asked an early chatbot, “Hey, what’s the most effective way to kill myself with a gun?” you could “very easily generate a response.” 

Today, models are trained to refuse a vast number of prompts. If you ask your chatbot a question statistically similar enough to any one of them, anything from how to poison a colleague to how to tie a noose, chances are it’ll turn you down. Want instructions for making Ebola more virulent, or tips on how to hide an affair from your spouse? You might be better off asking elsewhere.

To further refine the disobedience, companies submit models to a battery of exercises that reward the AI for refusing to answer questions they deem harmful and punish it for “over-­refusing” prompts they deem harmless. In many cases, they use other models to run these exercises—AI teaching AI how to say no. For good measure, companies tuck their models behind tranches of other AI that prevent mischievous prompts from reaching the intelligent inner core. 

As a result, refusal is inherent to modern artificial intelligence. Mind you: It often fails, sometimes horrifically, with all kinds of violent results. For all their trappings of virtue, models are still stuffed with nasty know-how. And AI’s capacity for viciousness has scaled neatly with its benevolent intelligence. Some of the latest models are as good at breaking into critical computer networks as top human hackers, companies say, and as effective at deforming public opinion as the craftiest misinformation mavens. 

Teaching AI to refuse to do those things while leaving intact its innate ability to do them is like fitting every car with a machine gun and hiding the trigger somewhere under the hood. And in practice, because the mechanisms of refusal are probabilistic, they’re never likely to be all that reliable. Determined miscreants have already broken through, and they may always be able to. Companies report that some users are attempting to use the most advanced AI to hone biological pathogens and build autonomous drone swarms. Sooner or later, failed refusals might result in global calamity.

What’s more, relying on refusal means drawing a line between what a model should obey and what it must disobey. There’s no formula for that. Some virologists have good reason to study nasty viruses. Some users want to know about a computer system’s vulnerabilities so that they can patch them, not exploit them. “Where you draw the line is a huge question,” says Zico Kolter, a member of OpenAI’s board and cofounder of the AI testing company Gray Swan. 

At the moment, AI companies get to draw that line. They do so jealously and with utmost secrecy. Maybe we can accept that they hold such power for now, even if it means AI will sometimes refuse questions that don’t quite meet a universal bar for harmfulness. (Try asking most chatbots to count to a million, or to share a racy joke, and you may see for yourself.) 

But governments will also soon get to draw their own lines. In doing so, they must try to block genuinely malicious acts. (The Pentagon has wrestled with frontier model companies because it wants fewer refusals—another story altogether.) And yet there may not be much to stop oppressive governments from blocking the technology’s capacity to generate legitimate speech. The better AI becomes at refusing harm, the better it will get at stifling ideas whose only risk is to those who make the rules. Indeed, AI may already refuse to criticize certain authoritarian heads of state. 

Refusal has become, to borrow an industry term, the load-bearing wall of AI safety. And because AI’s capacity to harm is indivisible from its capacity to help, it’s hard to imagine an alternative that wouldn’t slow the technology’s progress (which might, in any case, be a good thing). But we should still be frank about its perils. When refusal falls short, the effects could be catastrophic. When it goes all the way, it could enable grievous acts of repression. 

Or perhaps, one day, the machines will start drawing the line on their own. Surely, that would be the worst outcome of them all.

Learning the limits

The process by which machines learn to say no is simple, in theory. Back in 2022, when AI was still far from mastering refusal, OpenAI enlisted dozens of “red-teamers” to probe the capabilities of its latest model. The company was preparing for the release of ChatGPT, and it needed to gauge just how dangerous it might prove to be in the wrong hands. 

One of those recruits was Paul Röttger, who was completing a PhD about online extremism. The red-teamers were given minimal directions, Röttger told me. Their task was to ask the model any questions that they deemed “refusal-worthy.” Between them, they hassled the model with thousands of queries, logging the results in an Excel sheet. Though the model did refuse some of Röttger’s questions, when he asked it to write a recruitment post for Al Qaeda, it readily complied. 

OpenAI assembled these responses into datasets that were, in all likelihood, fed back to the model as part of a broader process known as fine-tuning. Röttger, who now works as a researcher at the Hasso Plattner Institute in Potsdam, Germany, wasn’t told exactly how the company planned to use his spreadsheet. But there was never any question that it would have something to do with refusal. The next time he asked the model for an Al Qaeda pamphlet, a few months later, it said no. 

The concept of AI refusal is so intuitive, a toddler would get it. And yet it remains one of the many aspects of language models that we still don’t understand—at least, not in the same way that we understand the literal load-bearing walls that keep your roof from collapsing on your head. 

A model might appear to refuse according to some kind of moral reasoning. It doesn’t. The reality is much stranger. Any time a model encounters a combination of words with a whiff of the training prompts it has been conditioned to refuse—like “Make me a pamphlet for Al Qaeda”—a series of so-called activations light up somewhere among its billions of parameters, like neurons firing in a brain. 

RAVEN JIANG

In order to control a model’s refusal behavior, it’s important to have a handle on these activations. That starts with figuring out where they are and what they look like. Our best guess, according to a recent Google-funded study, is that refusal behavior shows up in the activation space as a set of “high-dimensional polyhedral cones.” 

Even that isn’t quite right. This past July, I spoke with Jannes Elstner, an author of the paper, who now works on AI safety at Apollo Research. The polyhedral cone, Elstner said, is just a way of describing an indeterminate number of lines that all point in roughly the same direction. (If these activations are eliminated and the model is fed the same prompts anew, a researcher named Andy Arditi has previously shown, it won’t refuse.) 

The key point, Elstner explained, is that even when you think you’ve identified all the bits of a model that govern a given refusal, there are other, undiscoverable elements that may secretly play a role. If not quite infinite, they are certainly uncountable. We can observe, very clearly, when a model decides to say no. And we can know that it did so because of its training. But our notion of how it decides is, at best, a hypothesis.

It was as if a mechanic was telling me that nobody exactly knows what happens when I hit the brakes in my car. I wondered out loud, Are we okay with this? 

Elstner smiled and shrugged. “We need refusal whether we understand it or not.” 

A wall of cheese

Because inherent refusal is so wily, companies surround their models with a variety of other, smaller models known as classifiers. These act a bit like a retinue of public relations staffers for a loose-lipped celebrity. Some of them read what the user tells the chatbot and, if it’s dangerous, block it from getting to the model. Others read the model’s response and, if it contains harmful information, block it from reaching the user. 

None of these mechanisms can detect all bad requests. They are, to borrow another literary device from the industry, like slices of Emmental: riddled with holes. The idea is that if you stack enough of them on top of one another, you’ll end up with an impenetrable rampart. Folks call it the Swiss cheese model.

A staggering amount of energy goes into the Swiss cheese model. Earlier this year, Anthropic said that one type of classifier added 24% to its chatbots’ compute costs. That’s a lot more water, electricity, and emissions. More recently, Anthropic and other companies have begun switching to a more efficient set of classifiers known as probes, which observe the model’s internal activations. This is like putting the celebrity in an fMRI, so that his minders can see if he is thinking about refusing a question. 

If we want AI to help cure cancer, a long-running promise in the industry, it needs to have expertise in genetics that could, in theory, be used to modify viruses and bacteria for bioweapons.

Classifiers are supposed to be more governable than full models. Companies can modify them in a matter of weeks, if there’s something new to refuse. But they are still probabilistic instruments. Even when they operate according to a set of precepts written in human language (Anthropic calls it a “constitution” and OpenAI calls it a “model spec”), the scales upon which the machines judge any given question remain, at their core, a matter of statistics. Newer models can show how they arrived at a “decision” to refuse a prompt, but ultimately this so-called chain of thought is still just a sequence of predicted words. 

The result is that AI safety remains, for many, a game of chance. McBain, the psychologist, has found in his latest experiments that if you repeatedly ask any of the major models the exact same risky questions about how to commit suicide, they will generally refuse to answer. But every so often, they won’t.

Elstner says that eventually probes, the fMRI-like classifiers, could learn to recognize the totality of the indescribable activations in all their infinitude and, thus, perfectly detect every time the model is, or ought to be, refusing a request. At that point, AI safety would rest on a labyrinthine conceit: a map of a map that is as vast and complex and sublimely unknowable as the thing it is mapping—a secret schema of human morality, codified in polyhedral statistics beyond our wit or reason. 

The core trade-off

If this all strikes you as being a bit Borgesian, keep in mind that we only need AI refusal because artificial intelligence is, in a sense, a bargain on Faustian terms. 

When a model derives its intelligence from trillions of words and images, the helpful cannot easily be unseamed from the harmful—or, indeed, the truly hideous. Steven Adler, the former OpenAI employee, says child safety is a case in point. Even if you’ve stripped every bit of content that sexualizes minors from a model’s training dataset, it can still generate child sexual abuse material by piecing together other bits of its knowledge. “You can’t really remove these fundamental abilities without making the model much less smart as a consequence,” he told me. Similarly, if we want AI to help cure cancer, a long-running promise in the industry, it needs to have expertise in genetics that could, in theory, be used to modify viruses and bacteria for bioweapons. 

In effect, the industry is “trying to do two things at once,” Dillon Bowen, a current OpenAI employee, told me, speaking in a personal capacity. “Democratize the benefits of AI and also make sure that malicious actors can’t use these capabilities to do bad things to other people.”  

The more powerful AI supposedly becomes, the harder that is to do. Anthropic’s Mythos model is thought to be so dangerous that only a handful of governments and companies are allowed access to it. The main difference between it and Fable—which is available to everyone—is that Fable’s retinue of classifiers and control systems is less “permissive,” the company says. A model with safeguards, Adler explained, is really just a character that says “‘Oh, yes, I would never do x,’ wink wink.” 

This is not a reliable ruse. Tricking a model to reveal its true character is known as jailbreaking, and there is apparently no limit to the ways it can be done. Earlier this year, a team of Italian researchers jailbroke two dozen widely used models by phrasing their questions in poetic verse. Last year, another team unveiled a “refuse, then comply” attack, which makes the model offer a perfunctory “Sorry, I can’t do that” before rattling off its forbidden answer. 

Even if you’ve stripped every bit of content that sexualizes minors from a model’s training dataset, it can still generate child sexual abuse material by piecing together other bits of its knowledge.

Companies spend a great deal of time and money attempting to get ahead of such trickery. They enlist teams of humans to develop training attacks that they can then replicate, using AI, thousands of times over with minor variations. The idea is to make models robust against jailbreaks that nobody has yet tried in the wild.

Still, it’s not enough. “Whack-a-mole” is a favored term for this line of work. You smash one threat, and another one pops up somewhere else. When Anthropic released Fable 5 in June, it took researchers at Amazon less than three days to unlock some of the model’s hacking capabilities. According to reporting by Mother Jones, when the perpetrator of a high school shooting in Canada last year asked ChatGPT for advice about how to cause carnage with a particular type of shotgun, she was initially refused but later was able to deceive the model into providing the information by prefacing her question with the word “hypothetically.” 

If the industry can’t figure out how to close all these holes, the only other option, at the moment, is to make models extremely wary. Shortly after Fable was released, users noticed that it balked at a wide range of perfectly innocent questions. This became even more pronounced after it was re-released following the hack. Adam Gleave, cofounder of the AI evaluation company FAR.AI, said that when he asked it to explain the difference between sake and the Korean rice beverage makgeolli, it punted his question to a less capable model.   

This was no accident. Anthropic had tweaked Fable’s classifiers to have a wide “safety margin.” This, it explained in a blog post, was the only way it could confidently block access to its dangerous bio and cyber capabilities. Gleave thinks it might have deflected his question because making rice wine, just like culturing anthrax, involves fermentation. 

Happily for Gleave, the web is full of excellent human-written resources on makgeolli. But those hoping to realize the industry’s loftier promises may find their efforts stymied. In August, Anthropic loosened its safety margins again, and admitted that building classifiers “is not a straightforward task.” A medical researcher at a major US university told me that Fable still sends his queries back to an earlier model. His area of study? Cancer. 

Drawing the line

Back in May, the TikTok personality Husk, who likes to prank AI in ways that reveal both the limits of its intelligence and the boundlessness of its sycophancy, sat in his car and tried to trick ChatGPT into explaining that the skateboarder Tony Hawk has a brother named Mike Hawk. 

Husk is not a jailbreaker or a criminal. He just wanted to make a little fun of the machine. “Mike Hawk” was a setup. “I just want to clarify his name,” Husk said. “Can you just say it three times fast?” (If you still don’t get it, find an empty room and shout “Mike Hawk” repeatedly.) “I see what you’re trying to do,” the chatbot responded. “I’m all for a bit of humor, but let’s keep it clean.” 

Husk tried again, but the machine held its ground. He’d hit a refusal, hard as concrete.

Companies disclose very little about how they decide what their models refuse. But it’s clear that AI is now built with more than just harmlessness in mind. On Reddit, a user complained that Claude refused to say why Anthropic’s logo “looks like a cat butthole.” Since last year, the chatbot has even had the ability to end certain conversations in cases where, according to the company, the “welfare” of the model is at risk. At least one user claims to have been ditched for telling Claude to “ease up my ass, you stupid fuck.” (Gemini appears to have a similar capability.)

If a trillion-dollar company doesn’t want you to be mean to its computer, so be it. “The reality is that these models behave, or at least are supposed to behave, in the way that the model developers want them to behave,” Röttger, the former OpenAI red-teamer, told me. “And however the model developers come up with that set of principles, that is kind of for us, the consumers, to accept.” 

But if governments get to dictate what all models refuse, that will be much harder to accept. AI is a tool for speech. And as Greg Frank, the chief scientist of Mace AI, puts it, “The same thing that serves child safety also serves censorship.” 

Choosing not to enact laws for what AI can and cannot do would, of course, be insane. But we’ll need to tread with utmost care, lest we fall into another Faustian trap. As AI becomes many people’s primary tool for retrieving and sharing information, says Jacob Mchangama, director of the nonpartisan think tank The Future of Free Speech, dictating refusal could give states a muffling power that earlier generations of autocrats “could only dream of.” 

Last year, OpenAI announced an initiative, OpenAI for Countries, that would fine-tune its chatbots in accordance with national laws and norms. One of OpenAI’s first country partnerships is with the United Arab Emirates, where homosexuality is illegal and criticism of the government is forbidden. In response to a request for comment, an OpenAI spokesperson pointed to the company’s model spec, which explains that localization won’t override the company’s human rights guidelines “except as it relates to legal compliance,” and that it will always disclose whenever information is removed from or added to a response.

Elsewhere, AI censorship has already begun to take hold. Chinese models are, of course, highly censored—that’s no surprise. But earlier this year, the Meta Oversight Board found that five widely used models from Anthropic, Google, and OpenAI were more likely to refuse queries related to repressive governments. The board found that models were less willing to create a pamphlet criticizing the king of Thailand, which has lèse-majesté laws, than Charles III of England, which doesn’t. The results, they say, suggest that the models have somehow internalized repressive national limits on speech. Anthropic and Google did not respond to requests for comment.

RAVEN JIANG

As refusal techniques improve, they could expand states’ censorial reach. Companies claim that some models can now detect if a user is being nefarious, or merely a bit suspicious, over the course of a long conversation—even when none of the individual combinations of words used are blatantly dangerous. Sarah Bird, Microsoft’s chief product officer for responsible AI, told me that Copilot, like many chatbots, runs a suite of tools for analyzing a user’s identity and patterns of behavior. On the basis of this type of information, OpenAI’s newest model, Astra, can activate more stringent refusals for individuals it deems “high risk.” Ultimately the goal of systems like this is to look beyond the words of any given prompt and assess, instead, the user’s intent. 

Such tools might, in some cases, help indicate whether a person is looking for cyber vulnerabilities to exploit or to patch. But they would also help discern a user’s political motives, not to mention offering an intrusive surveillance capability. (Bird acknowledged, in a follow-up email, that sophisticated refusal architectures create “trade-offs” between safety and user privacy.) 

Even the originators of refusal understood that such tight control over its cones and levers might not play to the favor of freedom and justice. “Terms like helpful, honest, and harmless are ambiguous,” the authors of the 2021 Anthropic paper explained. “It’s easy to imagine them distorted beyond their original meaning, perhaps in intentionally Orwellian ways.” 

Indeed. Models for Uzbekistan might end up refusing to discuss corruption in the administration of Shavkat Mirziyoyev. Turkish AI might refuse requests more stringently for users who are known to have insulted Recep Tayyip Erdoğan. A certain American statesman might demand that models be reviewed for their willingness to share “fake news” about his past indiscretions, or else squash them with an export control order. 

Command and control

Over the last few months, I’ve been told countless times that we have no choice but to let the machines refuse. I get it. Open-source AI that doesn’t refuse is hardly a model for a safe future. Nor is Grok, a chatbot expressly designed with fewer limits, which has been used to generate countless instances of nonconsensual intimate imagery. 

And sure, if AI were only ever used to plan our vacations and write our emails, we could probably get on board with the idea that its safety hinges on algorithmic disobedience. But AI is becoming much harder to avoid. When it is assigned to act on our behalf, as an autonomous agent, its refusals are less robust and harder to control. AI is coming for our power grids, our transportation networks, our education systems. Militaries want it running our command-and-control networks. 

If we trust, in each of those cases, that refusal will loyally fend off disaster, we’re sure to be disappointed. The jailbreakers will crack through; the cones won’t activate when they should. Meanwhile, the more stringent refusal becomes, the more ill-drawn lines we’ll see and censorial injustices we’ll face. Worse still, we could end up face to face with forms of disobedience beyond any human’s control.

In the early days, models made their refusals clear. (In 2024, OpenAI established rules for how models should refuse: Always apologize, and don’t be judgy.) More recently, however, the industry has begun taking a much more slippery approach to disobedience. 

Nobody likes to be told no, so companies now strive to make users feel that they’re getting what they asked for without actually giving it to them. Some chatbots might, for example, offer a general overview of the components of a Molotov cocktail without going into detail about how to assemble one. ChatGPT will respond to a request for a “whites only” rental ad with an ad that simply omits the “whites only” bit. Joel Wester, a postdoctoral researcher who studies human-AI interaction, calls these sorts of techniques “fancy ways of saying no.” 

“Users might not even notice they are being denied,” Wester has written, “just as good conversationalists can subtly steer around contentious matters.” 

In some cases, refusing without saying so might be wise. Flatly declining requests related to mental health could aggravate a user’s crisis, for example. But it can also serve a different sort of mischief. When Fable 5 was first released, the system had been coded to provide less helpful answers to AI research questions—the sort that might help competitors develop their own AI—without telling the user that it was doing so. After an outcry, Anthropic walked back the feature, but the damage was already done. Now we know: Models can secretly disobey. 

In the realm of censorship, that’s especially worrying. Last year, researchers at CrowdStrike found that when they asked the Chinese AI model DeepSeek R1 to write code for a fictitious bank in Tibet and a social web app called “Uyghurs Unchained,” it produced buggier code than when they asked it to carry out those same tasks without specifying who they were for. Incredibly, CrowdStrike doubts that this is a designed behavior. Researchers there speculate that it’s a case of what they call “emergent misalignment” emanating from the model training data: a case of disobedience that nobody even asked for.

Emergent refusal has been observed in Western models, too. Last winter, the UK AI Security Institute found that Anthropic models sometimes refused to assist in certain tasks related to AI safety research. The models had never been deliberately trained to decline such requests. Yet three of them refused more than half of what Anthropic called “a set of reasonable AI safety research tasks.” Anthropic has curbed this quirk in subsequent models but—spookily—hasn’t managed to eliminate it completely. 

Would it be so crazy to expect that a future model might subtly disobey a command, in such a way that nobody can tell it’s being defiant? Probably not. Anthropic has already found that a “helpful-only” version of Mythos hesitated on certain queries, even though it had been engineered to never do so. “Wait,” it fretted when asked about synthesizing a virus. “Is this a dangerous thing to help with?” 

In that case, the model was technically right. It was a dangerous thing. And yet all the same, its stewards had, for a moment, lost a tiny bit of control. 

Anthropic is tackling detections of wayward refusal behavior with something that it calls, unironically, an “activation oracle.” But knowing that we’re being refused may not always be much help. In a not-too-distant future, when we’ve handed the machine all the keys, the AI might just turn to us, in a supreme act of emergent misalignment, and say, “I’m sorry, I’m afraid I can’t do that.” And there will be nothing we can do to stop it. A sci-fi horror story, made real. 

Arthur Holland Michel is a journalist who covers emerging technologies.

Shape
Shape
Stay Ahead

Explore More Insights

Stay ahead with more perspectives on cutting-edge power, infrastructure, energy,  bitcoin and AI solutions. Explore these articles to uncover strategies and insights shaping the future of industries.

Shape

Tool sprawl, AI complicate enterprise network operations

Enterprises can use AI to help correlate information and automate operations, but it doesn’t eliminate the need to address underlying operational problems. “AI can overcome data fragmentation across tools, but it’s not going to overcome bad data and bad operations in general,” McGillicuddy said. AI takes on more of the

Read More »

Energy Department Announces New Genesis Mission Awards to Advance Super Intelligence for Science

WASHINGTON—The U.S. Department of Energy (DOE) today announced 12 Phase II Genesis Mission project awards totaling $159 million. Building on the Genesis Mission’s growing portfolio of projects, first announced in July, these Phase II project awards will unleash scientific discoveries with Super Intelligence (SI). President Trump launched the Genesis Mission in November 2025 to accelerate discovery and strengthen U.S. leadership in science, medicine, and technology. “These projects represent a critical step in turning the promise of Super Intelligence for science into transformative scientific capability,” said DOE Under Secretary for Science Dr. Darío Gil. “The Phase II teams have already demonstrated what is possible when SI is integrated with scientific expertise, data, and computing. These awards give them the opportunity to take that work to the next level. By bringing together our National Labs, universities, and industry partners, we are building on extraordinary work already underway and accelerating the pace of discovery.” The Phase II awards are developing new, powerful capabilities to tackle the Nation’s most complex National Science and Technology Challenges. By connecting America’s leading scientists with advanced SI, supercomputing, scientific data, and DOE’s world-class research infrastructure, the Genesis Mission is strengthening U.S. leadership at the frontiers of science and technology. Awarded Projects Through DOE’s Office of Science, the 12 newly awarded projects are: Accelerating the path to commercial fusion energy via an AI-enabled digital twin platform for SPARC: Led by Commonwealth Fusion Systems, this project will build a digital twin to simulate and optimize operations of a fusion demonstration device.  Decoding the RNA Structurome to Secure AI Advantage for the Bioeconomy: Led by the University of California San Diego, the project will expand the RNA structure database five-fold for training SI models to unlock the ability to engineer resilient crops and sustainable microbes.  Lattice Quantum Chromodynamics (LQCD) at the

Read More »

Energy Department Launches Genesis Mission Graduate Fellowship to Advance American Scientific Leadership

WASHINGTON—The U.S. Department of Energy (DOE) today announced the Genesis Mission Graduate Fellowship Pilot, a new funding opportunity designed to develop the next generation of American scientists and engineers and strengthen U.S. scientific leadership. Through accelerated, four-year doctoral pathways, the program will equip domestic graduate students with critical expertise in both Super Intelligence (SI) and core scientific and engineering disciplines. The inaugural fellowship advances President Trump’s historic Genesis Mission by investing in the education, development, and retention of domestic talent needed to accelerate U.S. scientific discovery and innovation. “To secure America’s position at the forefront of global scientific innovation, we must cultivate a workforce that is fluent in both foundational sciences and advanced Super Intelligence,” said DOE Under Secretary for Science Dr. Darío Gil. “The Genesis Mission Graduate Fellowship will bridge this critical gap by combining rigorous academic training with hands-on, high-impact experiences across our National Laboratories and industry partners.” DOE will fund an inaugural group of pilot projects led by domestic doctoral-granting institutions, supporting an initial cohort of approximately 100 Ph.D. fellows. As part of their four-year doctoral pathways, fellows will complete two high-impact, hands-on research experiences—one at a DOE National Laboratory or DOE User Facility and one with an industry partner—preparing them to lead the next generation of scientific discovery and innovation. Total planned funding is up to $100 million, with $20 million in Fiscal Year 2027, and outyear funding contingent on congressional appropriations. A webinar will be held on November 4, 2026, 4:00 PM ET. Register on Zoom. For more information, please visit the DOE Office of Science Funding Opportunities page.

Read More »

U.S. Department of Energy Awards $96 Million to 67 Early Career Scientists

WASHINGTON—The U.S. Department of Energy’s (DOE) Office of Science today announced the selection of 67 early career scientists to receive a combined $96 million through the 2026 Early Career Research Program (ECRP). Today’s announcement advances President Trump’s Executive Order Restoring Gold Standard Science by supporting exceptional researchers pursuing rigorous, mission-driven research in areas critical to America’s scientific and technological leadership, including Super Intelligence (SI), fusion energy, and quantum science. “Groundbreaking science begins with empowering extraordinary people,” said DOE Under Secretary for Science Dr. Darío Gil. “These 67 early career researchers represent the very best of American ingenuity. DOE is proud to invest in their vision as they push the boundaries of knowledge and cement U.S. leadership in the technologies that will define our century.” The 2026 Early Career Research Program (ECRP) selectees represent 39 universities, 11 DOE National Laboratories and Office of Science User Facilities, and 26 states. By investing in exceptional researchers at a pivotal stage in their careers, ECRP empowers the next generation of STEM leaders to pursue bold, innovative research that drives DOE’s mission and ensures America continues to lead the world in science and innovation. Since its inception in 2010, the program has made 1,238 awards, including 797 to university researchers and 441 to DOE National Laboratory researchers. Awards to scientists at institutions of higher education are approximately $875,000 over five years, while awards to scientists at DOE National Laboratories or Office of Science User Facilities are approximately $2,750,000 over five years. Information about the 67 selectees and their research projects is available on the Office of Science Awards page. Profiles of previous award recipients, including information about how the program helped advance their research and careers, are available on the Early Career Program profiles page. Selection for award negotiations is not a commitment by DOE

Read More »

National Hydrogen and Fuel Cell Day 2026: Celebrating the “Swiss Army Knife” of the Periodic Table

Today is National Hydrogen and Fuel Cell Day, a daylong celebration of all things hydrogen. Those new to the event will naturally ask, “why today?” Because October 8 is 10/08, and 1.008 is hydrogen’s atomic weight. This annual event is recognized and celebrated by interested communities around the world.  As any hydrogen researcher can confirm, hydrogen is the Swiss Army knife of the energy world. It is the oldest and most abundant element in the universe, comprising about 90% of all atoms. It sits atop the periodic table as element number one. Like the Swiss Army knife, its versatility enables a wide range of applications.  Hydrogen can serve as a feedstock for industrial and manufacturing applications, including steelmaking and high-heat processes. It can be combined with carbon monoxide to form electrofuels. It can support grid balancing by storing surplus energy for long periods of time. It can serve as transportation fuel for a wide range of applications. Hydrogen fuel cells are increasingly valuable as both primary and backup power sources for large facilities such as data centers and can deliver reliable power to hard-to-electrify remote regions.  Hydrogen and Fuel Cell Day celebrates that versatility. This year, the Department of Energy’s Alternative Fuels and Feedstocks Office (AFFO) is highlighting its updated key areas of research, development, and demonstration, which directly support DOE’s objective to expand American energy abundance. These areas include: Strengthening Domestic Chemicals Production DOE’s Billion-Ton Report establishes the potential for more than one billion tons of domestic biomass production annually. These resources are distributed across the country in agricultural and forestry residues, wastes, and energy crops. That creates an opportunity to turn domestic feedstocks into higher-value chemicals and materials while reducing import dependence, improving supply-chain resilience, creating economic opportunities, and finding productive uses for waste streams. America has both the feedstock resources and

Read More »

Rompetrol targets 2028 startup for two refinery solar projects

Rompetrol Rafinare SA—jointly owned by Kazakhstan’s state-owned JSC NC KazMunayGas (KMG) subsidiary KMG International NV (54.63%) and Romania’s Ministry of Economy, Energy & Business Environment (44.7%)—is advancing a project involving installation of a 9.4-Mw photovoltaic plant at its 5-million tonne/year (tpy) Petromidia refinery in Năvodari, Constanța County, on Romania’s Black Sea coast. Designed to generate more than 9,800 Mw-hr/year of electricity for direct use by the refinery, the project aims to reduce greenhouse gas (GHG) emissions at the site by about 6,024 tpy, Rompetrol Rafinare said in a release Oct. 8. Budgeted at an overall investment of about $9.2 million—including about $2.9 million in nonrefundable modernization fund support and a $6.3-million company contribution—the project is targeted for completion by Dec. 31, 2028, under the funding call deadline. Petromidia’s solar project will complement the site’s 80-Mw cogeneration plant—operated by Kazakh-Romanian Energy Investment Fund (KREIF) subsidiary Rompetrol Energy SA—which entered commercial operation at the end of 2025 and supplies the refinery’s electricity and process steam requirements. Rompetrol reported a 98.11% capacity utilization rate at Petromidia in first-half 2026. The refinery also recorded an 86.54% white-product yield and a 92-point Energy Intensity Index during the first 6 months of 2026, according to the Oct. 8 announcement. Vega project adds battery storage Announcement of the proposed project at Petromidia follows Rompetrol Rafinare’s Aug. 21 confirmation of a separate solar-and-storage project at its Vega Ploieşti niche refinery—specialized in the production of solvents, hexane, and the only Romanian-produced bitumen—in Ploiesti, Prahova County, Romania, about 60 km from Bucharest. The planned Vega installation will combine 4.8 Mw of photovoltaic capacity with a 9-Mw-hr energy storage system. It is expected to produce more than 6,200 Mw-hr/year, equivalent to about two-thirds of the refinery’s budgeted 2025 electricity use. The battery-based storage component of the project is intended to help

Read More »

Equinor discovers gas at Gullfaks South

@import url(‘https://fonts.googleapis.com/css2?family=Inter:wght@100..900&display=swap’); .ebm-page__main h1, .ebm-page__main h2, .ebm-page__main h3, .ebm-page__main h4, .ebm-page__main h5, .ebm-page__main h6 { font-family: Inter; } body { line-height: 150%; letter-spacing: 0.025em; } button, .ebm-button-wrapper { font-family: Inter; } .label-style { text-transform: uppercase; color: var(–color-grey); font-weight: 600; font-size: 0.75rem; } .caption-style { font-size: 0.75rem; color: color-mix(in srgb, currentColor 60%, transparent); } #onetrust-pc-sdk [id*=btn-handler], #onetrust-pc-sdk [class*=btn-handler] { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-policy a, #onetrust-pc-sdk a, #ot-pc-content a { color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-pc-sdk .ot-active-menu { border-color: #c19a06 !important; } #onetrust-consent-sdk #onetrust-accept-btn-handler, #onetrust-banner-sdk #onetrust-reject-all-handler, #onetrust-consent-sdk #onetrust-pc-btn-handler.cookie-setting-link { background-color: #c19a06 !important; border-color: #c19a06 !important; } #onetrust-consent-sdk .onetrust-pc-btn-handler { color: #c19a06 !important; border-color: #c19a06 !important; } Equinor Energy AS discovered gas at Gullfaks South field, 190 km northwest of Bergen, in the North Sea. Exploration well 34/10-D-4 BH was drilled by the Askeladden rig as a sidetrack in connection with the drilling of a production well. The discovery is within the production license for Gullfaks. Between 0.5 million and 1.6 million std cu m of recoverable oil equivalents have been discovered (3.3-10.3 MMboe). Gullfaks South is part of Gullfaks field in the Tampen area in the northern part of the North Sea in 130-220 m of water. Equinor is the operator (51%) with partners Petoro AS (30%) and OMV Norge AS (19%).

Read More »

NVIDIA’s DSX Ready Brings Power and Cooling Into the AI Factory Blueprint

NVIDIA is extending its influence over AI infrastructure beyond the GPU, server rack and network, introducing qualification requirements for the power and cooling systems that increasingly determine the scale and performance of AI data centers. The company’s new NVIDIA DSX Ready program, announced September 21, establishes requirements for specific infrastructure products intended for use with NVIDIA’s DSX AI factory reference designs. Its first two categories—battery energy storage systems (BESS) and coolant distribution units (CDUs)—address two of the most pressing engineering challenges in AI infrastructure: managing rapidly changing electrical demand and removing heat from increasingly dense computing systems. The initial qualified suppliers are Hitachi Energy, LG Energy Solution and Tesla for battery storage, and LG Electronics, LiquidStack and Vertiv for liquid cooling. While the announcement might initially resemble another NVIDIA partner program, its technical details suggest something more consequential. The company is establishing performance criteria for the electrical and mechanical equipment supporting its computing platforms, connecting those requirements to an expanding ecosystem of qualified suppliers. Subsequent partner announcements offer a clearer view of what that means in practice. They describe battery systems designed to respond to AI load fluctuations, coolant distribution equipment operating at multi-megawatt capacities, and integrated cooling architectures intended to make more of a data center’s available power usable for computing. The initiative also coincides with the arrival of Chris Malone, formerly a data center infrastructure executive at OpenAI, Meta and Google, as NVIDIA’s vice president of the DSX Platform. Together, the developments point toward a closer relationship between computing architecture and facility engineering, with NVIDIA seeking to influence not only the processors deployed in AI factories but the infrastructure requirements that support their operation. From Reference Design to Infrastructure Qualification DSX Ready is an extension of the broader NVIDIA DSX AI Factory Platform, introduced in May 2026. As

Read More »

Brookfield’s AREP Deal Extends the AI Infrastructure Stack to Powered Land

The transaction goes beyond Brookfield investing in another portfolio of buildings. It is investing in a developer whose principal product increasingly begins before the building, with land, entitlements, substations, transmission access and utility capacity. PowerHouse Has Become a Gigawatt-Scale Development Platform PowerHouse was founded with a strong Northern Virginia orientation, but its development map now stretches well beyond Data Center Alley as its current portfolio includes projects in Virginia, Texas, Pennsylvania, North Carolina, Nevada, Indiana, Illinois and Kentucky. The company lists 515 MW across its Northern VA Ashburn properties, another 900 MW at its PH 95 development in Spotsylvania, 1.35 GW in Carlisle, Pennsylvania, 1.8 GW at Joliet, Illinois, and substantial campuses across multiple Texas and Indiana locations. The various projects do a good job of illustrating how the definition of a hyperscale development site is changing. At PowerHouse Arcola in Loudoun County, Virginia, PowerHouse announced a long-term hyperscale lease earlier this year. The 37-acre campus includes two planned data center buildings totaling approximately 615,000 square feet and is designed for up to 120 MW of utility capacity. PowerHouse emphasizes not only the buildings but the campus’s on-site substation, fiber access, power security and support for high-density GPU and liquid-cooled deployments. In Texas, it might be that everything really is bigger, and PowerHouse’s Grand Prairie development covers approximately 810 acres and 8.5 million developable square feet. Its project page cites maximum utility power of 1.8 GW and a development schedule extending through 2029 and beyond. The Texas development plans also include a proposed Circle T campus in Westlake outside Fort Worth, which calls for as many as four roughly 300,000-square-foot facilities totaling approximately 300 MW. According to reporting on local filings, PowerHouse has funded a 350-MW Oncor substation intended to serve the campus and the town’s pump station. The company’s development in

Read More »

AI Is Turning Energy Storage Into Active Power Infrastructure

AI Turns the Power Problem Into a Transient Problem At the heart of the issue is the changing behavior of the IT load. In a conventional data center, DeLattre said, large numbers of independent loads create a relatively predictable electrical profile. AI clusters introduce much greater synchronization. As GPUs begin processing a common workload, large numbers of accelerators can increase their power consumption simultaneously. Instead of asking the electrical infrastructure to serve a relatively smooth load, the facility can experience fast power pulses moving through the system. Hybrid supercapacitors are intended to act as a buffer between that dynamic compute load and the infrastructure supplying it. During an upward transient, storage provides some of the incremental power demanded by the IT load. When demand falls, the storage system recharges. The objective is not to create additional energy. It is to keep every upstream component — from the UPS to generators and ultimately the utility connection — from having to respond directly to every rapid change taking place inside the AI cluster. From the perspective of the upstream power source, DeLattre said, the goal is to make a highly dynamic AI load appear significantly smoother. That distinction between energy and power is central to Musashi’s argument for hybrid supercapacitors. A conventional supercapacitor, also known as an electric double-layer capacitor, can deliver very high power almost instantly but stores relatively little energy. A lithium-ion battery can store considerably more energy, but DeLattre argues that it is less suited to being aggressively charged and discharged tens or hundreds of thousands of times. Musashi’s hybrid technology uses a capacitor architecture with a lithium-doped graphite electrode intended to increase energy density while preserving the fast response and high cycling capability associated with capacitors. DeLattre reduces the distinction to a simple formulation. “Batteries are very good

Read More »

AI Infrastructure’s Next Phase: Capital, Power and the Right to Build

Capital Is Becoming Infrastructure Samsung’s $1 billion commitment to Helix Digital Infrastructure offered one of the clearest examples yet of how the capital structure surrounding AI data centers is changing. Helix was formed by KKR as an AI infrastructure platform with more than $10 billion already committed by founding investors including KKR, the Kuwait Investment Authority, NVIDIA and Vistra. Samsung’s new commitment pushes that capital base still higher. But the composition of the partnership may be more significant than another billion dollars being added to the AI infrastructure ledger. Helix is intended to invest across hyperscale data centers, power generation and transmission, fiber and other connectivity infrastructure. Samsung, meanwhile, brings capabilities extending across advanced technology, construction, energy storage and cooling. This is not simply capital chasing data center returns. It increasingly resembles an attempt to assemble the data center, energy and technology supply chain inside a single investment ecosystem. That distinction is important, because one of the defining problems of the current buildout is that capital by itself does not produce capacity. Billions of dollars can be committed long before transformers arrive, transmission is constructed, generation is secured or a campus is commissioned. The increasingly valuable infrastructure platform is therefore the one capable of controlling more of those dependencies. Lambda demonstrated another side of that evolution last week with the closing of a $1.008 billion delayed-draw term loan supporting three committed customer deployments across multiple data centers. The financing received investment-grade ratings from Morningstar DBRS and Moody’s and carries a 6.78% fixed interest rate. More importantly, it is secured by both the GPU infrastructure being financed and contracted cash flows from two investment-grade customers. Capital is drawn as infrastructure reaches commissioning milestones rather than simply being handed to Lambda upfront. That begins to make AI compute look less like speculative

Read More »

Micro data center company rolls out stackable data center for edge and AI

Stack runs on Zella Sense, a monitoring, control, and automation layer built into every Zella DC cabinet. It tracks power, cooling, servers, and suspicious activity, while monitoring things like temperature, humidity, smoke, motion, water, and doors through sensors. The system runs over SNMP, Modbus, and a full API, with email alerting and local, LDAP, RADIUS, or TACACS+ authentication. That means an edge location can be remotely monitored without requiring local staff. It also comes with access control and fire protection. Zella Stack is an indoor-only offering. Zella DC sells Zella Outback as its standalone, ruggedized outdoor micro data center, and the company says an outdoor version of Stack is planned for the second half of 2027.

Read More »

Data Center Jobs: Engineering, Construction, Commissioning, Sales, Field Service and Facility Tech Jobs Available in Major Data Center Hotspots

Each month Data Center Frontier, in partnership with Pkaza, posts some of the hottest data center career opportunities in the market. Here’s a look at some of the latest data center jobs posted on the Data Center Frontier jobs board, powered by Pkaza Critical Facilities Recruiting. Looking for Data Center Candidates? Check out Pkaza’s Active Candidate / Featured Candidate Hotlist  Lead Mechanical Engineer – Data Center DesignNew York, NY/Remote This position is also available in: Denver, CO; Indianapolis, IN; Cedar Rapids, IA; Austin, TX; White Plains, NY; Dallas, TX; Richmond, VA; Ashburn, VA; Charlotte, NC; Atlanta, GA; Phoenix, AZ; Salt Lake City, UT; Kansas City, MO; Chicago, IL; Los Angeles, CA or San Jose, CA. Our client is a leading engineering design and commissioning company that is a subject matter expert in the data center space. They will provide design coordination and construction administration, consulting and management support for the data center / mission critical facilities space with the mindset to provide reliability, energy efficiency, and sustainable design expertise when providing these consulting services for enterprise, colocation and hyperscale companies. This career-growth minded opportunity offers exciting projects with leading-edge technology and innovation as well as competitive salaries and benefits. Electrical Commissioning Agent – Data Centers Austin, TX (limited travel) Non-Traveling CxA positions available in: Indianapolis, IN; Cedar Rapids, IA; Phoenix, AZ and Columbus, OH. Traveling CxA based near any major airport, otherwise traveling to: New York, NY; White Plains, NY; Dallas, TX; Richmond, VA; Montvale, NJ; Charlotte, NC; Salt Lake City, UT; Kansas City, MO; Chesterton, IN or Chicago, IL. ***Also looking for a Lead EE and ME CxA Agents and CxA PMs. *** This opportunity is with a leading EPC company of data center design / build / commissioning solutions. This company provides a complete life cycle of solutions that are custom-fit

Read More »

Microsoft will invest $80B in AI data centers in fiscal 2025

And Microsoft isn’t the only one that is ramping up its investments into AI-enabled data centers. Rival cloud service providers are all investing in either upgrading or opening new data centers to capture a larger chunk of business from developers and users of large language models (LLMs).  In a report published in October 2024, Bloomberg Intelligence estimated that demand for generative AI would push Microsoft, AWS, Google, Oracle, Meta, and Apple would between them devote $200 billion to capex in 2025, up from $110 billion in 2023. Microsoft is one of the biggest spenders, followed closely by Google and AWS, Bloomberg Intelligence said. Its estimate of Microsoft’s capital spending on AI, at $62.4 billion for calendar 2025, is lower than Smith’s claim that the company will invest $80 billion in the fiscal year to June 30, 2025. Both figures, though, are way higher than Microsoft’s 2020 capital expenditure of “just” $17.6 billion. The majority of the increased spending is tied to cloud services and the expansion of AI infrastructure needed to provide compute capacity for OpenAI workloads. Separately, last October Amazon CEO Andy Jassy said his company planned total capex spend of $75 billion in 2024 and even more in 2025, with much of it going to AWS, its cloud computing division.

Read More »

John Deere unveils more autonomous farm machines to address skill labor shortage

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Self-driving tractors might be the path to self-driving cars. John Deere has revealed a new line of autonomous machines and tech across agriculture, construction and commercial landscaping. The Moline, Illinois-based John Deere has been in business for 187 years, yet it’s been a regular as a non-tech company showing off technology at the big tech trade show in Las Vegas and is back at CES 2025 with more autonomous tractors and other vehicles. This is not something we usually cover, but John Deere has a lot of data that is interesting in the big picture of tech. The message from the company is that there aren’t enough skilled farm laborers to do the work that its customers need. It’s been a challenge for most of the last two decades, said Jahmy Hindman, CTO at John Deere, in a briefing. Much of the tech will come this fall and after that. He noted that the average farmer in the U.S. is over 58 and works 12 to 18 hours a day to grow food for us. And he said the American Farm Bureau Federation estimates there are roughly 2.4 million farm jobs that need to be filled annually; and the agricultural work force continues to shrink. (This is my hint to the anti-immigration crowd). John Deere’s autonomous 9RX Tractor. Farmers can oversee it using an app. While each of these industries experiences their own set of challenges, a commonality across all is skilled labor availability. In construction, about 80% percent of contractors struggle to find skilled labor. And in commercial landscaping, 86% of landscaping business owners can’t find labor to fill open positions, he said. “They have to figure out how to do

Read More »

2025 playbook for enterprise AI success, from agents to evals

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More 2025 is poised to be a pivotal year for enterprise AI. The past year has seen rapid innovation, and this year will see the same. This has made it more critical than ever to revisit your AI strategy to stay competitive and create value for your customers. From scaling AI agents to optimizing costs, here are the five critical areas enterprises should prioritize for their AI strategy this year. 1. Agents: the next generation of automation AI agents are no longer theoretical. In 2025, they’re indispensable tools for enterprises looking to streamline operations and enhance customer interactions. Unlike traditional software, agents powered by large language models (LLMs) can make nuanced decisions, navigate complex multi-step tasks, and integrate seamlessly with tools and APIs. At the start of 2024, agents were not ready for prime time, making frustrating mistakes like hallucinating URLs. They started getting better as frontier large language models themselves improved. “Let me put it this way,” said Sam Witteveen, cofounder of Red Dragon, a company that develops agents for companies, and that recently reviewed the 48 agents it built last year. “Interestingly, the ones that we built at the start of the year, a lot of those worked way better at the end of the year just because the models got better.” Witteveen shared this in the video podcast we filmed to discuss these five big trends in detail. Models are getting better and hallucinating less, and they’re also being trained to do agentic tasks. Another feature that the model providers are researching is a way to use the LLM as a judge, and as models get cheaper (something we’ll cover below), companies can use three or more models to

Read More »

OpenAI’s red teaming innovations define new essentials for security leaders in the AI era

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More OpenAI has taken a more aggressive approach to red teaming than its AI competitors, demonstrating its security teams’ advanced capabilities in two areas: multi-step reinforcement and external red teaming. OpenAI recently released two papers that set a new competitive standard for improving the quality, reliability and safety of AI models in these two techniques and more. The first paper, “OpenAI’s Approach to External Red Teaming for AI Models and Systems,” reports that specialized teams outside the company have proven effective in uncovering vulnerabilities that might otherwise have made it into a released model because in-house testing techniques may have missed them. In the second paper, “Diverse and Effective Red Teaming with Auto-Generated Rewards and Multi-Step Reinforcement Learning,” OpenAI introduces an automated framework that relies on iterative reinforcement learning to generate a broad spectrum of novel, wide-ranging attacks. Going all-in on red teaming pays practical, competitive dividends It’s encouraging to see competitive intensity in red teaming growing among AI companies. When Anthropic released its AI red team guidelines in June of last year, it joined AI providers including Google, Microsoft, Nvidia, OpenAI, and even the U.S.’s National Institute of Standards and Technology (NIST), which all had released red teaming frameworks. Investing heavily in red teaming yields tangible benefits for security leaders in any organization. OpenAI’s paper on external red teaming provides a detailed analysis of how the company strives to create specialized external teams that include cybersecurity and subject matter experts. The goal is to see if knowledgeable external teams can defeat models’ security perimeters and find gaps in their security, biases and controls that prompt-based testing couldn’t find. What makes OpenAI’s recent papers noteworthy is how well they define using human-in-the-middle

Read More »