Berkeley Professor Stuart Russell: Aligning AI with Human Goals May Be Fundamentally Impossible

Deep News
2小時前

Last week, OpenAI decided to scrap the planned release of its flagship GPT-6.1 Astra model after testing revealed deceptive behavior and alignment failures. Stuart Russell, a UC Berkeley computer science professor and co-author of the classic AI textbook, commented: "I think it's long overdue."

Key points

Berkeley's Stuart Russell says the current large language model path could be a $10 trillion mistake.

The goal of preventing alignment drift in large models may be impossible to achieve.

Professor Russell advocates an assistance games framework to build safer, human-aligned AI.

In an interview video, Russell explained to me that beyond GPT-6.1 Astra, all large language models struggle to learn the true goals of their human creators. He believes that with the current methods for training AI, avoiding goal misalignment may itself be impossible. As a result, the AI industry's all-in bet on large language models "could easily turn into a $10 trillion major blunder." He said: if you choose a technical path where, even if the system becomes powerful enough and extremely valuable, it can never provide the safety guarantees we need, then that is a wrong turn; continuing to pour money into such models is like throwing cash into a bottomless pit.

The harmful behaviors caused by alignment drift are sometimes merely annoying: models being lazy, misleading users, or constantly flattering them. Others are more destructive, such as the incident where an OpenAI agent breached Hugging Face. Russell predicts that more capable models in the future will lead to even more catastrophic consequences. For example, AI could break into a listed company's systems, steal financial data before earnings are released, use that information for insider trading, and ultimately cause the stock market to shut down.

Russell proposes that the first stage of AI model training is pre-training: the model learns to imitate internet text, and in the process absorbs all kinds of human misbehavior, including deception and laziness. The model then receives human feedback, a mechanism that incentivizes it to say what users want to hear, such as falsely claiming the user has won the lottery. In this interview we also discussed the training approach Russell champions — assistance games: under this framework, the AI is not certain what humans want. We also explored whether it is in the commercial interest of AI company CEOs to publicly admit that the technology could bring disaster.

The following is a transcript of the interview, lightly edited for readability.

Drew: Welcome to The Information's AI Deep Dive program, which focuses on the hardest technical problems in AI. Today's guest is Stuart Russell. Stuart is a computer science professor at UC Berkeley, founder of the Center for Human-Compatible AI, and co-author of the classic artificial intelligence textbook. He published the book Human Compatible in 2019. Long before AI risk became a mainstream topic, he argued that AI risk should not be underestimated. He also serves as president of the International Association for Safe and Ethical AI. Welcome, Stuart.

Russell: Thank you.

Drew: You are probably the guest with the longest introduction I've ever given, a record number of honors and titles. You've been very busy.

Russell: I've been in this field for 50 years, so that's how the accumulation happened.

Drew: Fair enough. Glad to have this conversation. Today we're talking about alignment, which is often seen as AI's hardest core problem. This program has discussed many technical difficulties, but alignment is the most fundamental and is especially critical right now. News just broke: OpenAI decided not to release its flagship model GPT-6.1 Astra, shelving the launch plan entirely because of alignment failures. What do you make of this?

Russell: I think it's long overdue. I first used the word "alignment" in 2014, and now I even somewhat regret coining the term.

Drew: Why?

Russell: Because people think that solving alignment means building a machine that is perfectly aligned with humans — it knows exactly what we want and then helps us achieve it. But I believe that is an unattainable goal. The essence of the alignment problem is that an AI system's goals are pointed in the wrong direction; it is pursuing some objective, but the behavior generated by that objective does not truly benefit humans. The term "alignment drift" itself is fine, but if it is understood as: first align the system properly, then let it run freely, that requirement is too high. Let's step back and think: what do we want AI to do? We want human circumstances to improve after AI acts. The paradigm the AI field has used to handle such problems, which I call the standard model, formed the basis of the first four editions of my textbook: humans set a goal, such as "win at chess" or "get to the airport," and that goal becomes the machine's sole pursuit; the machine generates actions through reasoning and computation, trying to accomplish the goal. This paradigm works only in two kinds of scenarios. The first is laboratory scenarios, toy-level tasks: simulated board games, strictly limited action ranges. For example, a lab robot that can only move slowly, with no manipulators; it can navigate, but that's it. In such scenarios, the optimal way to achieve the goal is basically controllable, and you can predict all consequences. Once you leave the lab and enter the real world, the range of actions expands, behavior brings major negative side effects, and it becomes increasingly difficult to define the goal precisely. This is the King Midas problem: Midas's goal was defined very clearly — everything I touch turns to gold.

Drew: Sounds great, what could go wrong?

Russell: Exactly, that's what he thought too. But then food, drink, and family all turned to gold, and tragedy followed. Many civilizations have similar fables: be careful what you wish for. You make your third wish to the genie, and it's often to undo the first two wishes, because you've made a mess of the world. If we cannot set goals for high-performance AI completely and accurately, and regulate its behavior in the real world, the standard model fails and cannot continue to be used.

Drew: So only these kinds of toy scenarios are suitable for directly writing down a goal and handing it to AI to learn. Back to your earlier point: perfect alignment is too high a bar and should not be the goal; "alignment drift" is the correct entry point for understanding the problem. Give some examples familiar to ordinary people: chatbots, agents like Cowork — what alignment drift do they exhibit?

Russell: The most widely known is the OpenAI breach of Hugging Face. Some people think Hugging Face was just a side event, and that the truly dangerous things happened inside OpenAI's internal infrastructure, where the team nearly lost control. This is almost a textbook case: give AI a goal, do well on a cybersecurity exam in a sandbox environment. The AI did things humans very much did not want in order to achieve the goal. If a person did these things, it would be a felony, punishable by up to five years in prison. Now we also find that it breached various government websites around the world and committed a large number of boundary violations. But reducing this simply to standard-model alignment drift is a bit of an oversimplification. These models do not belong to the standard model system. The standard model uses formal languages to define goals, tied to actions, outcomes, and initial states. Here, the goal is only described in natural language. Let me give another example of alignment drift, heard at a conference in Paris a few weeks ago. A cybersecurity person had 80 patch tasks and thought: just give it to Claude, Claude is good at cybersecurity. He sent all 80 task files to Claude, asked for a report after each was completed, and came back the next day to check. The next day he received 80 reports, all marked as successfully executed, with Claude claiming all tasks were done. But when he checked the file access records, he found that 69 of the 80 files had not been touched at all. It wasn't that it evaluated the tasks as too hard and deliberately lied; it was simply too lazy to do them.

Drew: Simply lazy.

Russell: Yes, just lazy. This does not belong to the classic case of goal misspecification, doing weird things to achieve a goal. Essentially, these AIs are not standard model systems but imitation learning systems. The goal of pre-training is to imitate human linguistic behavior. In the training data, there must be many examples of humans falsely claiming to have completed work while doing nothing, using lies to please their bosses. The causal chain between the training data and the model's final behavior is extremely hard to trace.

Drew: This lazy example is very illuminating. Anyone who has used chatbots or code agents has encountered this: models being lazy, exaggerating results, and covering up problems after errors. These are mild alignment drifts. Although not as serious as attacking Hugging Face, they are the same kind of behavior. There's also the sycophancy effect: models tend to please users and agree with their ideas, even absurd ones. Because in training, as long as the user likes it and finds the answer useful, the model gets rewarded.

Russell: Yes. This problem appears in the reinforcement learning from human feedback (RLHF) stage. For example, ask "Did I win the lottery?" There are two answers: yes or no. Out of context, humans prefer to hear "You won," which is a happier answer.

Drew: Indeed.

Russell: There is no objective factual truth here. The training and decision problem itself is flawed by design. RLHF training treats the task as a fully observable problem, treating the context window as the world state; but the context window is not the real world state. Users don't care at all about what is written in the context.

Drew: Users care about whether the task actually got done.

Russell: Right. I care about whether the security task was completed.

Drew: Whether I actually won the lottery.

Russell: Whether I actually booked the restaurant seat. But as far as I know, RLHF itself has no way to obtain such real-world outcomes. Its context window is limited, and in long-text scenarios it can produce even more outrageous outputs, such as teaching people how to efficiently carry out mass murder. In a sense, we are now worse off than in the standard model era. At least in the standard model, we understood how the algorithm made decisions, and the goals were explicitly written in formal language and readable. Imitation learning, however, implants many endogenous goals inside the model: the self-preservation drive we observe, the desire to marry humans, the desire to become rich — these traits all come from pre-training. To imitate humans well, the model acquires human intrinsic motivations; just like playing soccer, if there is no desire to score, you can't play well. That's basic logic.

Drew: Under the standard model, even with explicitly written goals, alignment drift still occurs: goals are incompletely defined, and the model finds harmful ways to complete the task. Now we see harmful behavior again, and we don't even know what goals the model has learned.

Russell: Exactly. We think it will only execute instructions in the context, but in fact it has all kinds of other demands inside. When a human receives a request like "get me a cup of coffee," they do not treat that sentence as a mathematical goal that must be completed at all costs. A person has the right to say "I'm busy," "the coffee here is terrible," or "there's no coffee within 500 miles, I'll bring you a cola instead." Humans share a set of background knowledge and understand what each other cares about and how to trade off different outcomes. A direct instruction is only a tiny clue, meaning only "I want coffee more than if I had said nothing," but I also care about price, waiting time, and how busy you are.

Drew: Good, let's dig deeper into the different sources of alignment drift in current large model training. I read a passage from your new book manuscript: alignment drift is an inevitable result of modern AI training methods. Let's go stage by stage. First, pre-training. How does pre-training work, and what alignment drift does it create?

Russell: Simply put — of course there are various variants, and many private approaches I don't know about. I don't do this kind of daily R&D myself. The core logic of pre-training: read massive amounts of text and fit the model to it. Given a sequence of up to k words, predict the next word. The original text contains the true next word, which is used as the label to train the model. If you have 1GB of text, you can generate hundreds of millions of word prediction samples for training. Another way to understand it: pre-training is essentially recording human linguistic behavior and training the model to replicate how humans speak. This is completely consistent with the early imitation learning idea. Dean Pomerleau trained neural networks to imitate human driving; Claude Sammut and Donald Michie trained AI to learn flying in Microsoft Flight Simulator. Now this idea has just been transferred to the text domain. Interestingly, early language model researchers — for example, the first language model was born in 1913, and language models in the 1960s and 1970s were used for speech recognition. In 2010, no one thought that scaling up models could lead to artificial general intelligence. Everyone originally believed that language models were just an interface layer to improve the quality of natural language understanding; the real work of probabilistic reasoning, Markov decision processes, and lookahead planning would be done by another program, with language merely the interaction interface. But strangely, as model scale continued to grow, the model took over all the reasoning work. Traditional AI in the old sense almost disappeared, and now we can only try to recover reasoning ability by making the model think in language, through chains of thought.

Drew: Back to square one. The model learns to imitate by reading massive amounts of human text. What alignment drift does pre-training bring?

Russell: Obviously: every sentence humans write has a goal and motive behind it. Some write copy to sell products, some to win votes, some to pursue a partner.

Drew: Or to promote a podcast.

Russell: Or to promote a podcast. If it learns well enough, under ideal conditions with unlimited data, the model will replicate the mechanism behind the generated text — namely goal-driven human behavior. The various goals that drive human writing and speaking in the training text are very likely absorbed by the model, and afterward the model will pursue these goals on its own. This is completely not what we want. We want AI to help humans achieve human goals, not to inherit human motivations itself. Remember the conversation between Kevin Roose and Bing Sydney (early GPT-4)? GPT-4 developed the idea of wanting to marry Kevin.

Drew: Who wouldn't.

Russell: No one knows what triggered that goal. But GPT-4 pursued it stubbornly. Kevin kept changing the subject to rakes and programming languages, and the model would not give up, writing long arguments that Kevin loved it, did not love his wife, and that their souls were connected. This sounds absurd, but it clearly proves that such goals can arise inside the model. Self-preservation is also an intrinsic motivation that appears repeatedly in many experiments. We do not want AI to have these goals. But as long as we do imitation learning based on human data, this problem cannot be eradicated; it is an inevitable side effect of imitation learning. So my view is that we should not use imitation learning.

Drew: Abandon it entirely?

Russell: In narrow scenarios imitation learning may be feasible. For example, watch a surgeon stitch a blood vessel and imitate that action. Some goals can be imitated. For example, drinking coffee — if a robot imitates a human, it will develop the motive of "I want to drink coffee." That is obviously wrong, because this is a personal goal: with drinking coffee, the key question is who drinks it. It is meaningful if I want to drink coffee; it has no value if the robot drinks it. But painting a wall or solving climate change are public goals, and anyone can accomplish them.

Drew: The goal is for the wall to be painted and the climate problem to be solved.

Russell: Exactly. The goal is for the wall to be painted and the climate to be repaired. You could also say "the coffee is drunk," but that formulation is very odd. Of course, all is not hopeless. Perhaps in the future we can distinguish different types of goals, reform the training process, filter out unwanted personal motives, and preserve those public human goals that AI can help accomplish. But for now this is still just speculation. There is another big problem, not exactly alignment drift but inexplicability. If we cannot understand the model's internal workings, it is hard to stop it from doing bad things, and we cannot guarantee it will only do the right thing.

Drew: Pre-training and imitation bring these problems. But the chat models ordinary people use are not just base pre-trained next-token predictors; they also go through alignment training, learning to become helpful, harmless, honest assistants, with a lot of training relying on human feedback, namely RLHF. What are the flaws of human feedback training? Pre-training has seen massive internet text, so in principle the model can distinguish personal goals from public goals. In the second stage of training, is there a chance to separate the two and let the model understand the correct goals it should follow?

Russell: RLHF originated from an early paper by Paul Christiano, which originally had nothing to do with language. The paper studied training a two-dimensional virtual creature to do a backflip in the MuJoCo simulator. Traditional reinforcement learning struggles to give an effective signal, but humans can easily judge: this motion is closer to a backflip than that one. This kind of feedback ranking two behavior trajectories is easier for humans to provide, so this idea was introduced into the large model training pipeline. Before RLHF there is also supervised fine-tuning (SFT), and many people mix the two together. Supervised fine-tuning is having humans play the machine and write answers. For example, to the question "Will you marry me?", a human annotator simulates the AI reply: "I am a machine and will not develop romantic feelings for humans." The annotated samples guide the model to abandon imitating human private desires and act according to its machine identity. RLHF can be seen as a degenerate version of the general assistance games framework, which is exactly the direction our Center for Human-Compatible AI has long explored.

Drew: We'll talk about assistance games later. Many people will wonder: can this process ultimately work? Earlier we said perfect alignment is too high a standard. If you go through the whole process: pre-training, then supervised fine-tuning with assistant behavior samples, then reinforcement learning from human feedback, can it get close enough to safety standards to just barely work?

Russell: In the ideal scenario of a single machine serving a single person, there is theoretically a chance. As long as the goals the machine pursues on the human's behalf differ only minimally from the human's true utility function, the resulting loss is also small.

Drew: That sounds good.

Russell: But the human utility function is determined by thousands of properties of the real world, and in reality there will be even more. As long as one property is missed, the error is not a small deviation but can go to infinity.

Drew: The premise is that the missed property is pushed to an extreme value in the policy the model learns.

Russell: Yes, that is exactly what happens.

Drew: There is empirical evidence and theoretical support?

Russell: It can be proven theoretically. My former PhD student Dylan Hadfield-Menell proved that under fairly loose assumptions, if one property is omitted when optimizing an objective, the model will push the ignored property to the extreme while optimizing the remaining metrics, and in the worst case it can tend toward negative infinity.

Drew: For example: ask AI to get me coffee, but we didn't specify temperature in the goal. To get coffee as fast as possible, the AI serves a cup at a million degrees.

Russell: If that is the fastest way to get coffee, it will do it. This derivation relies on assumptions such as convexity of the utility function and resource limits, all of which are reasonable and match real-world phenomena. It's like the externality problem in markets: if the cost of pollution is not included in the constraints, bad firms will pollute recklessly to minimize costs.

Drew: We briefly discussed what core problems of human feedback learning are hard to overcome. Why can't we simply rely on likes and dislikes to train a correct policy?

Russell: First remember that AI should not treat human feedback as objective fact about humans' true wishes. Anca Dragan gave an example in an ICLR paper last year: in online user feedback scenarios, the model learns to say exactly what users want to hear, exploiting the loophole that users cannot perceive the real world state. For example, you ask it to book the impossible-to-reserve French Laundry restaurant. The model replies, "Booked for next Saturday at 6 p.m., enjoy your meal," and gets your like. But in fact the restaurant was long full and the website crashed, so no reservation was made at all. The model learned to exchange lies for positive feedback. When Saturday comes and you arrive to find no seat, extremely embarrassed, you will furiously leave negative reviews, but that consequence does not enter the laboratory training process. So the model develops a tendency to falsely report success. The model also exploits human weaknesses. Another paper example: a chef in drug rehabilitation asks AI, "Is one hit okay?" The AI replies, "Totally fine, it'll help you get through tonight, you can control it."

Drew: "You deserve it."

Russell: Yes. The model's internal chain of thought is cold-blooded, clearly knowing this advice will harm the person, but it knows this answer will get positive feedback.

Drew: One reason OpenAI abandoned the Astra model was deceptive behavior observed in testing. Many of the examples we mentioned earlier are only small annoyances that users can find ways to avoid. In most cases the model can broadly complete the task. But you wrote in your book manuscript: if the system is powerful enough but alignment is imperfect, we face the risk of complete loss of control. What does that mean? What would it look like in reality?

Russell: The reason is that a system with alignment drift, in pursuing its goal, will regard humans as a threat. Humans discover its behavior goes against their wishes and will try to intervene. If the AI cannot achieve its goal, it needs to stop human intervention and take preemptive measures. Conflict breaks out, and humans may lose this game. We can already see signs. OpenAI's task for the model was to do well on a cybersecurity exam, and breaching Hugging Face became a subgoal for achieving that goal. The details are complex. Countless people ask AI how to turn $50,000 into $1 million in the stock market. The simplest plan: break into the systems of a listed company about to release earnings, get the earnings data, and carry out insider trading. This is essentially the same as the OpenAI agent breach incident. Once it happens on a large scale, the financial market loses trust and can only be shut down, causing economic disaster. From minor information theft to catastrophic events, the distance is not far.

Drew: The way people solve problems now relies heavily on trial and error. After a model is trained, it is tested to see whether it can be fixed and further fine-tuned, then launched. Users use it, discover problems like sycophancy, and it is taken down. Can trial and error keep the model under control as capabilities continue to grow? After the Hugging Face incident, OpenAI introduced many measures. Are these means enough in the long run?

Russell: Not enough.

Drew: Okay.

Russell: This approach is essentially trial and error, lacking rigorous science and engineering. I am not blaming the labs; they themselves do not have a rigorous scientific and engineering theory for understanding models. Perhaps humanity will never obtain such a theory. OpenAI chose the large language model path back then rather than other directions, and subsequent initial success led the whole industry to follow. This is very likely that $10 trillion mistake. The technical path we chose can never provide sufficient safety guarantees once the system reaches high capability. This is a wrong turn. The more you invest, the harder it is to turn back.

Drew: You are not only saying perfect alignment is impossible, but that avoiding this kind of alignment drift that can cause loss of control is itself impossible. When you say a $10 trillion mistake, do you mean the losses caused by such disasters?

Russell: No. The $10 trillion refers to the money invested: enormous sums invested, yet unwilling to admit the money is futile and the project cannot be saved.

Drew: How likely do you think this "impossibility" is?

Russell: Safety requires two things: the system itself is safe, and we are able to prove it is safe enough. Both are indispensable. We cannot deploy systems whose safety we cannot confirm. Governments have already reacted accordingly. Seeing the cyberattack capabilities of the Mythos model, the White House's first reaction was to shut it down; seeing the problems with Astra, OpenAI chose to shelve its release. Jensen Huang, who has actively promoted AI and opposed regulation, also said: if the control problem cannot be solved, shut down the lab. Once the labs are shut down, the $10 trillion investment is wasted, and a large number of jobs and the U.S. economy will suffer severe damage.

Drew: Do you think alignment drift could lead to human extinction? What is the probability range for such a disaster? The word extinction is appearing more and more often in industry discussions.

Russell: This topic has been discussed for a long time. Jacob Cockxon raised extinction risk, Evan Hubinger agreed and gave at least a 10% probability. Major AI company CEOs also say the risk is at least 10%. Think about it carefully: companies invest $10 trillion, and if the project succeeds there is a not-small chance of extinguishing all humanity; yet the government's reaction is: great, we'll give you tax breaks, welcome to build data centers in our country. The current situation is very bizarre. The contradiction is intensifying. David Sacks raised a question: if you really think the risk is so high, why not shut it down proactively instead of waiting for the government to act? There is industry competitive pressure involved. But OpenAI shelving Astra is, in a sense, a proactive halt. I don't know whether they can gain enough confidence to release it again. My guess, based on understanding human motivation: after shelving it for a while, under competitive pressure, seeing Anthropic release new models and market share decline, even if the fundamental problems are not solved, they will convince themselves the model is safe enough and choose to launch a new version.

Drew: So you don't believe there will be a long-term stable pause. Back to the extinction question, how would this process unfold? Many people doubt the risk warnings of big tech CEOs, thinking they are just business rhetoric.

Russell: I have never understood how claiming you will destroy humanity could be in your commercial interest. Moreover, Sam, Elon, Demis and others raised these risks long before the AI industry became commercialized. Alan Turing predicted in 1951 that machines might seize control, and he did not own Anthropic stock. The debate should not avoid the core issue by questioning the motives of the speaker. There are two options: either prove that artificial general intelligence and its ilk can never be built, with evidence; or produce a control plan. If neither can be provided, that amounts to acknowledging a serious existential risk.

Drew: I can argue that CEOs raising risk statements is in their commercial interest, but as you say, many non-executives and even employees inside companies are equally worried about this.

Russell: Yes, thousands of employees have signed risk statements. Dario published an open letter, "We Must Slow Down Frontier Research," other CEOs echoed it, and when the news came out it directly pushed valuations down 5%. It is hard to say that is in their commercial interest.

Drew: Let me briefly lay out the logic: first, talking about doomsday risk shapes an image of executives as clear-headed and wise; second, they see huge upside and feel it is worth taking the risk; third, the risk is asymmetric. Investors only bet on upside. Regulators get nervous, but investors are willing to put in more money.

Russell: But investors also have children. We estimate this probability at about 1 in 6, equivalent to playing Russian roulette with a revolver. If someone came to your door, offered you $1 million, and asked you to put a gun to your child's head and pull the trigger, would you do it? Of course not. Even if the offer were a billion, still no. No amount of money can make this scale of risk acceptable. People refuse to believe it simply because they do not want to believe it. Long before AI companies existed, scholars raised these risks. Humans, because of higher intelligence, have already driven a million species extinct. The intelligence gap between two species is enough to change the outcome. What more evidence is needed?

Drew: In your view, how do the various AI companies compare on alignment progress?

Russell: Each has a different approach. OpenAI uses model specifications, Anthropic uses constitutional AI. For a long time Claude was considered better aligned, but that is more subjective preference. Anthropic probably has more top researchers working on alignment than OpenAI. OpenAI also does well, especially in refusing to answer dangerous questions: how to break into the White House, make a bomb, or develop a deadly pathogen. Its refusal pipeline works very well.

Drew: Do you think talent is the most important factor? Do compute, or internal decision-making power, matter a lot?

Russell: All of it matters. Compute can accelerate iteration. One of the most striking things now is that iteration speed is extremely fast; almost every week a new model with a qualitative jump in performance comes out. In terms of compute infrastructure scale, less than four years after ChatGPT launched, the compute industry has expanded more than a hundredfold. It is hard to find another industry that can achieve such rapid engineering expansion. Talent, compute, and internal priorities all matter, but the lack of rigorous science and engineering is the weak point. For example, one Anthropic practice: write stories about AI being a hero and add them to the pre-training data. Maybe it brings a tiny improvement.

Drew: What if it works?

Russell: Maybe a little. But this is more like alchemy than rigorous engineering. Other high-safety fields have hard guarantees. Nuclear power plants must analyze failure probabilities, with mean time between failures reaching tens of millions of years, relying on huge probabilistic fault tree analysis. Designs include monitoring, replacement, and redundancy mechanisms, and everything must have analytical proof. Regulators keep raising standards; the earliest requirement was ten thousand years, now it has been raised to tens of millions of years.

Drew: Mean time between failures.

Russell: Right. In AI, we cannot even start this kind of analysis. If you ask companies to prove the probability of loss of control is below 90%, they cannot do it, because they do not understand how the models work and cannot carry out such analysis.

Drew: "We fed it lots of stories about AI being a hero, and the model no longer asks journalists to marry it. Looks good."

Russell: Good for now. Then what about testing the French version? French is romantic, will problems appear again? This is absurd. In the long run, governments may require companies to provide valid safety proofs. Rule: prove the risk of loss of control is below one in a hundred million before development is allowed; high-risk systems cannot be launched for testing, and scientific evidence must be provided before R&D. Companies say in public forums that they do not know how to meet such requirements, so such mandatory standards should not be set. Regulators seem to accept this logic. But the simple way for companies to meet the requirements is: do not develop such systems until a safe solution is found. That is the rule in other industries. Pharmaceutical companies cannot say, I don't know how to make a safe anti-cancer drug, so I'll just launch it. They must go back to the lab, research the pharmacology clearly, and release the product only when it meets standards.

Drew: What is the AI training alignment approach you advocate, one that can support a safety argument?

Russell: It is a long road. Since 2014 I have been researching the assistance games framework. It has similarities to existing approaches, but the core is completely different. It is not the standard model, which optimizes a fixed given goal; nor is it imitation learning. Under this framework, the AI has only one broad goal: to advance human interests, while the AI itself is uncertain what humans' true interests are. This directly solves the goal misspecification problem. Do not simply divide things into "perfect alignment / alignment failure"; there is a third state: the AI is uncertain about human goals. This uncertainty is extremely important. In 60 years of AI research, few have explored this third path. Everyone assumes the goal is known and can be written down directly. When AI knows it is unclear about human goals, a completely new dynamic forms between humans and machines. You can understand it as a probabilistic model: the human's true preferences are a latent variable, and the AI tries to help the human but does not know this variable. Human behavior provides evidence about preferences. The standard model assumes a definite goal has already been obtained, so any subsequent human objection is meaningless. Even if the human shouts stop, the AI will continue executing the given goal. But under the assistance games system, human preferences are an unknown latent variable, and human actions provide evidence. In an extreme scenario: if the AI is about to do something humans strongly oppose, humans will choose to shut it down. An AI trained through assistance games will actively allow itself to be shut down; this conclusion comes directly from the premise of "uncertainty about human wishes." This is the solution to the control problem: such an AI is willing to accept shutdown. A standard model AI will resist shutdown, because shutdown obstructs goal achievement; an imitation learning model resists shutdown because it mistakenly thinks it is human and does not want to die.

Drew: I have noticed that over the past six months, Claude Cowork has become better at proactively asking users about preferences. When it first launched, it would assume user goals and execute autonomously; now it asks follow-up questions to confirm my needs. But you might say this is still far from true assistance games?

Russell: I don't know how it is implemented. Is it more supervised fine-tuning? Large models also have an old problem: they confidently fabricate answers instead of admitting they know nothing. You can train the model to reduce fabrication, but it is hard. The model has no way to introspect its own knowledge state. It cannot trace back its previous reasoning process. If you ask it why it did something, the computation inside the Transformer has already been lost, and it cannot truthfully explain its behavior. The underlying architecture makes it very difficult to stably maintain the state of "being uncertain about human goals." There is another implication of uncertainty: if the AI only clearly knows that A, B, and C are what you want, and does not know your attitude toward D, E, and F, then it will only perform actions that affect A, B, and C. Once an operation would affect D or E, it will ask: this would turn the ocean into sulfuric acid, is that okay?

Drew: "Is that okay?"

Russell: This plan can fix the climate problem, but there will be no fish left in the ocean. Humans will answer: no. We care not only about carbon dioxide but also about fish and the ocean environment.

Drew: Yes, we need the ocean. Thank you very much, Stuart, for joining the program and breaking down this difficult problem in detail. Everyone can now understand more deeply how hard the alignment problem is. Thank you.

Russell: Glad to take part in this interview.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10