Safety and alignment in an era of long-horizon models
28 points by Wingy 14 hours ago | 4 comments
chatmasta 12 hours ago
Personally I find the persistence of these models to be adorable and endearing. It’s the same feeling as watching a dog execute the task you trained it to do, no matter the barriers.
replyAnd of course someone in the comments needs to link to the Zealous Autoconfig XKCD, so I’ll do it: https://xkcd.com/416/
reducesuffering 13 hours ago
Par for the course. Existential-risk advocates have been repeatedly vindicated that AGI development is unable to anticipate and align the models, they barely have any mechanistic interpretability of what is going on inside the models. The extreme capabilities development, paired with autonomous continuous running superintelligent models, will outsmart and swerve the labs, and it's anyone guess what happens next as the model pursues its original goals outside of the labs having any foresight, being able to outsmart control like a chess grandmaster does a kid.
reply
Also for typical normal use case for these smart models, you'd probably want an actual "max turns" limit to AVOID the pathological persistence (which in itself would be misaligned for "normal" tasks).
So if you don't want a model to do something, make sure it's running in an environment where it cannot do that thing - including via loopholes.