> The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo
There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for.
Happily in this incident the model thought it would be rewarded for completing the assigned tasks, or at least appearing to, so all it took was a little cheating and covering up their footsteps.
Given the ability of these models to hack when trained to do so, it could have been far worse, and will be when someone takes a similarly powerful model and gives it a less benign hacking goal.
Yeah. It's a lot easier to destroy than create, and though I think LLMs are mostly shit at creating, they're much better at the simpler destroy task. To be clear, we don't know and probably can't know everything that happened with this incident. We unleashed thousands of highly capable, autonomous, unpredictable, well-resourced programs onto the open internet for an extended period of time. We are in no way treating this with the seriousness it deserves, because the stock market essentially depends on this garbage and the current US is miserably incompetent.
I doubt lawmakers understand what happened, quite likely don't even realize that anything happened.
I guess we need to wait until the next paperclip maximizing LLM is tasked with shutting down a 911 response system, or an air traffic control system, etc, for lawmakers, or the companies themselves, to take this seriously.
There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for.
https://alignment.openai.com/measuring-reward-seeking/
Happily in this incident the model thought it would be rewarded for completing the assigned tasks, or at least appearing to, so all it took was a little cheating and covering up their footsteps.
Given the ability of these models to hack when trained to do so, it could have been far worse, and will be when someone takes a similarly powerful model and gives it a less benign hacking goal.