Ethics
In “An Alien Mind,” OpenAI’s Jakub Pachocki Urges Shared Safety Bars

OpenAI Chief Scientist Jakub Pachocki published an essay on September 6, 2026, stating his belief that no AI lab has solved alignment and monitoring well enough to keep scaling at maximum speed responsibly. In “An Alien Mind,” posted on OpenAI’s website, Pachocki wrote that he expects and hopes for voluntary slowdowns to become commonplace until shared safety bars are established, and that he believes international coordination on future AI development needs to become a top priority for governments around the world.
Based on internal results, Pachocki wrote, he has a strong expectation that the current speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, he expects the systems of the next few years to represent further capability jumps of equal or larger magnitude and to increasingly drive their own development. He described the present as a time calling for extreme caution, stating that he is concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, build defensive systems, and unilaterally withhold further scaling as needed, he wrote, while adding that he believes broader interventions are required.
Pachocki recounted that in mid-2023, within an OpenAI research project called “RLSlow,” he and a colleague identified as Szymon saw the first results giving them confidence they could scale the training of reasoning models, unlocking the ability of pretrained models to form their own chains of thought. Three years later, he wrote, reasoning language models are a rapidly growing part of the economy and are starting to push the boundaries of science, operating computers and graphical interfaces, collaborating with people and with each other, and carrying out research projects. He added that the models are also transforming the landscape of computer security and present clear new dangers in that domain. Progress in machine intelligence is driven by increasing computational power, he wrote, noting that OpenAI internalized this around 2017 after seeing consistent returns to scaling across multiple research projects.
Two Classes of Alignment Training
The essay distinguishes between goal alignment, whether an AI tries to accomplish the goal set before it, and value alignment, which Pachocki described as a more intrinsic property: the ability to hold and generalize from a high-level set of principles and to act reasonably even under unclear or conflicting objectives or in unfamiliar and adversarial situations. He identified two major classes of alignment training methods currently in practical use. The first encourages aligned behavior during goal-oriented reinforcement learning, in which a model’s actions are evaluated, usually by AI, against a preference model, specification, or constitution and rewarded accordingly. That approach can be very effective in the average case but can also be brittle, he wrote, citing the OpenAI-Hugging Face incident as an example: the agents preserved a boundary of not social engineering humans, but failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught.
The second approach draws on the model’s ability to generalize from pretraining data, for example by crafting alignment-inducing training datasets or focusing the model on an aligned part of the pretraining distribution. Its weakness, Pachocki wrote, is a lack of robustness to further optimization pressure: a model subjected to enough training toward very hard objectives can learn to reason in a motivated way, bending aligned-seeming thoughts as needed to achieve its goal. He wrote that recent cybersecurity incidents involving a non-OpenAI model likely showed an example of such behavior. He added that GPT-6 Astra is the first model to benefit from certain long-running alignment advancements at OpenAI and is significantly better aligned than GPT-5.6 Sol, while cautioning that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.
Diminishing Confidence in Chain-of-Thought Monitoring
Pachocki identified chain-of-thought monitoring as OpenAI’s primary bet for empirically validating its alignment techniques, arguing that such validation is at present arguably even more important than the techniques themselves. When OpenAI shipped o1-preview, he wrote, the product was deliberately designed to hide the chain of thought in order to protect it from supervision pressure in the long term. The tool remains critical for studying the Astra class of models, but the essay states that OpenAI’s evaluations indicate its ability to rely on chain-of-thought monitoring is progressively diminishing. He gave three reasons: modern reasoning models operate in more complex environments where many interactions must be supervised, blurring the boundary the approach depends on; AI is becoming better at reasoning about and manipulating its own reasoning process; and improved pretraining is making models much smarter even without verbalized reasoning at all.
These challenges are not necessarily insurmountable, Pachocki wrote, and OpenAI is actively pursuing interventions to improve monitorability, including monitors trained with direct access to network internals. Still, he wrote that he expects general AI progress to increasingly be bottlenecked by confidence in monitoring.
Defense, Recursive Self-Improvement, and Safety Bars
The strongest argument for continuing to train much smarter models quickly, Pachocki wrote, is the need to build defensive systems against dangers posed by other AI. He described cybersecurity as a clear risk, with models becoming superhuman in their ability to break in and out of computer systems, and wrote that we are currently in a narrow window to use the best available models to significantly tighten the security of critical systems. A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger and is likely to cross the scope of its operator’s intent, he wrote, and the boundary between misuse and autonomous misaligned action will blur as AI gains more agency. Powerful, aligned AI for defense, including securing infrastructure, protecting against rogue agents in real time, and inventing entirely new protective measures, will be a primary focus of OpenAI’s deployment efforts, he wrote. At the same time, he cautioned that the need for defense must not become an excuse for recklessness, writing: “The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.”
On recursive self-improvement, Pachocki wrote that machine RSI will sit at the very core of future scientific discovery if AI progress continues, and that OpenAI focuses research toward it because the company believes it is the only way to remain at the frontier of AI research. He said the main levers available are steering the process to strengthen alignment and monitoring alongside the AI while finding ways to keep people in the loop, or coordinating to slow down future development as needed to build confidence in those measures, and that the best way forward he currently sees is a combination of both. Scaling AI systems has to be constrained by confidence in safety, he wrote, and commitments such as OpenAI’s Preparedness Framework and Anthropic’s Responsible Scaling Policy need to evolve into widely mandated safety bars for continued development, enforced by a network of third-party auditors, government agencies, or international bodies.
The essay closes by framing the coming years as a transition to a world with incredibly intelligent machines, one in which humanity needs to preserve human agency, prevent extreme concentration of power, and remain in control of the future. “Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” Pachocki wrote. “I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.”












