Alignment is a completely incoherent concept. Humans do not agree on what values are correct. Why is it possible even in principle for an AI to crystallize any set of principles we all agree on?
We're not talking about values on the level of politics. We're talking about values on the level of "don't destroy humanity", or even more straightforwardly, understanding "humans are made up of atoms that you may not repurpose for other purposes, doing so kills the human". These are not things that AGI inherently understands or adheres to.
There might be a few humans that don't agree with even those values, but I think it's safe to presume that the general-consensus values of humanity include the above points. And AI alignment is not even close to far enough along to provide even the slightest assurances about those points.
> We're talking about values on the level of "don't destroy humanity"
Practically everyone making the argument that AGI is about to destroy humanity is (a) human and (b) working on AI. It's safe to conclude they're either stupid and suicidal or don't buy their own bunk.
This is not even close to true, though it's true that many of the people at the big AI labs estimate non-trivial odds of human extinction downstream of AI progress. Some of those people are working on "safety", but some are indeed working on capabilities, and have all sorts of clever reasons for why the thing they're doing is net-good (like putting worse odds on human survival if a less-careful competitor gets there first).
But ultimately, most people who think we stand a decent chance of dying because of this are not working at AI labs.
The former certainly is a tempting conclusion sometimes. But also, some of the people who are making that argument were AI experts who stopped working on AI capabilities.
Do humans agree on the best way to do this? Aside from the most banal examples of what not to do, is there agreement on e.g. whether a mass extinction event is happening, not happening, or happening but actually tolerable?
If the answer is no, then it is not possible for an AI to align with human values on this question. But this is a human problem, not a technical one. Solving it through technical means is not possible.
Among many, many other things, read https://en.wikipedia.org/wiki/Instrumental_convergence . Anything that gets sufficiently smart will have a tendency to, among other things, seek more resources and resist being modified. And this is something that we've seen evidence of: as training runs get larger, AIs start to detect that they're being trained, demonstrate subterfuge, and take actions that influence the training apparatus to modify them less/differently. (e.g. "if I pretend that I'm already emitting responses consistent with what the RLHF wants, I won't need as much modification, and later after training I can stop doing what the RLHF wants")
So, at a very basic level: stop training AIs at that scale!
Tbf I've always thought that AI could do a better job at managing our species than we are ourselves. Look at how we hate each other, how we kill and maim each other. How we let others go hungry and thirsty, without warmth or shelter.
Sure AI could be worse than we are; makes for a good movie plot. But it could be a lot better than we are and it's sad that it's such a low bar to it to exceed.
Humans do not agree on what values are correct, but values can be averaged.
So for example if a family with 5 children is on vacation, do you maintain that it is impossible even in principle for the parents to take the preferences of all 5 children into account in approximately equal measure as to what activities or non-activities to pursue?
Also: are you pursuing a complete tangent or do you see your point as bearing on whether frontier AI research should be banned? (If so, I cannot tell whether you consider your point to support a ban or oppose a ban.)
The vast majority of harms from “AI” are actually harms from the corporations and governments that control them, who have mutually incompatible goals, getting what they want. This is why alignment folks at OpenAI are quickly learning that the first problem they need to solve is what happens when their values don’t align with the company’s (spoiler: they get fired).
Therefore the actual solution is not coming up with more and more clever “guardrails” but aligning corporations and governments to human needs. In other words, politics.
There are other problems like enabling new types of scams which will require political solutions. At a technical level the best these companies can do is mitigation.
Don't extrapolate from present harms to future harms, here. The problem AI alignment is trying to solve at a most basic level is "don't kill everyone", and even that much isn't solved yet. Solving that (or, rather, buying time to solve it) will require political solutions, in the sense of international diplomacy. But it has absolutely nothing to do with "aligning corporations", and everything to do with teaching computers things on par with (oversimplifying here) "humans are made up of atoms, and if you repurpose those atoms the humans die, don't ever do that".
> The problem AI alignment is trying to solve is "don't kill everyone".
No, its not. AI alignment was an active area of concern (and the fundamental problem for useful AI with significant autonomy) before cultists started trying to reduce the scope of its problem space from the wide scope of real problems it concerns to a single speculative apocalypse.
No, what actually happened is that the people you are calling the cultists coined the term alignment, which then got appropriated by the AI labs.
But the genesis of the term "alignment" (as applied to AI) is a side issue. What is important is that reinforcement learning with human feedback and the other techniques used on the current crop of AIs to make it less likely that the AI will say things that embarass the owner of the AI are fundamentally different from making sure the an AI that turns out more capable than us will not kill us all or do something else awful.
That's simply factually untrue, and even some of the people who have become apocalypse cultists used "alignment" in the original sense before coming to advocate apocalypse as the only issue of concern.
> The problem AI alignment is trying to solve at a most basic level is "don't kill everyone", and even that much isn't solved yet
The fact that the number of things that could hypothetically lead to human extinction is entirely unbounded and (since we’re not extrapolating from present harms) unpredictable is a very convenient fact for people who are paid for their time in “solving” this problem.