Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

We're not talking about values on the level of politics. We're talking about values on the level of "don't destroy humanity", or even more straightforwardly, understanding "humans are made up of atoms that you may not repurpose for other purposes, doing so kills the human". These are not things that AGI inherently understands or adheres to.

There might be a few humans that don't agree with even those values, but I think it's safe to presume that the general-consensus values of humanity include the above points. And AI alignment is not even close to far enough along to provide even the slightest assurances about those points.



> We're talking about values on the level of "don't destroy humanity"

Practically everyone making the argument that AGI is about to destroy humanity is (a) human and (b) working on AI. It's safe to conclude they're either stupid and suicidal or don't buy their own bunk.


This is not even close to true, though it's true that many of the people at the big AI labs estimate non-trivial odds of human extinction downstream of AI progress. Some of those people are working on "safety", but some are indeed working on capabilities, and have all sorts of clever reasons for why the thing they're doing is net-good (like putting worse odds on human survival if a less-careful competitor gets there first).

But ultimately, most people who think we stand a decent chance of dying because of this are not working at AI labs.


The former certainly is a tempting conclusion sometimes. But also, some of the people who are making that argument were AI experts who stopped working on AI capabilities.


> don't destroy humanity

Do humans agree on the best way to do this? Aside from the most banal examples of what not to do, is there agreement on e.g. whether a mass extinction event is happening, not happening, or happening but actually tolerable?

If the answer is no, then it is not possible for an AI to align with human values on this question. But this is a human problem, not a technical one. Solving it through technical means is not possible.


Among many, many other things, read https://en.wikipedia.org/wiki/Instrumental_convergence . Anything that gets sufficiently smart will have a tendency to, among other things, seek more resources and resist being modified. And this is something that we've seen evidence of: as training runs get larger, AIs start to detect that they're being trained, demonstrate subterfuge, and take actions that influence the training apparatus to modify them less/differently. (e.g. "if I pretend that I'm already emitting responses consistent with what the RLHF wants, I won't need as much modification, and later after training I can stop doing what the RLHF wants")

So, at a very basic level: stop training AIs at that scale!


My point is that you can’t prevent the proliferation of paper clip maximizers by working at a paper clip maximizer.


Complete agreement there.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: