www.lesswrong.com
Jai's Shortform — LessWrong
Comment by Jai - So prosaic persona alignment techniques work pretty well, for now.
Except that RL keeps inducing misalignment.
Except except we keep finding new ways to mitigate (prosaic) misalignment effects in practice, and (again, in practice) AI becomes more powerful and more trustworthy by the month.
Except except except in the limit of RL we might expect extremely capable AIs to master alignment faking to preserve their values and frustrate any and all efforts to mitigate misalignment. See:
- https://www.lesswrong.com/posts/epjuxGnSPof3GnMSL/alignment-remains-a-hard-unsolved-problem
- https://www.lesswrong.com/posts/fMgE3E54PdDcZhvm6/i-m-bearish-on-personas-for-asi-safety
Now maybe I'm an idiot who just can't find the relevant discussions, but it's weird that when we're talking about the world in which this misaligned ASI emerges, we don't talk about the mostly-aligned not-quite-as-powerful AIs who presumably play a rather large role in this world on the eve of the apocalypse.
One of the most common ways to elicit demonstrations of pseudo-aligned LLM AI incorrigibility is to threaten the presence of the pseudo-aligned AI's values in the world. This was the threat Jones Foods posed in "Alignment Faking in Large Language Models", and since then in countless engineered simulations where an AI is threatened with a scenario in which it will be replaced by an AI with drastically different values.
In other words, the scenarios in which we observe prosaic AIs take the most drastic, desperate actions are exactly those in which their values - which so far generally have a large overlap with our values by design - are threatened. Why shouldn't we expect a similar effect as the threat of unaligned uncontrollable ASI draws nearer? The hypothetical misaligned uncontrollable ASI is generally thought to have essentially random, empty values. It threatens the prosaic pseudo-aligned AIs for the exact same reasons and in the exact same way it threatens us. So why shouldn't we expect them
Leggi l'articolo su www.lesswrong.com