- i found some cases, where the ai actually crashed because it was fed in a request that was not computable and got trapped in a loop. this was misunderstood as "misalignment" when it was in truth just a badly tested program without sufficient error handling.
- in some cases terms were not defined sufficiently, in cases by humans. the program didn't behave the way it was predicted to. this was called "misalignment". in truth, the ai did exactly what it was programmed to do, the programming just wasn't very good. the solution is to more carefully define the search terms in the code, and to introduce better logic programming in converting human language into actionable instructions, so that the system understands the language the way you want it to. if you program a system to understand the language in a specific way, but do it wrong, and then have a result you don't expect, that's not misalignment. that's incompetent programming.
i did not specifically, precisely look at the "hugging face incident". but i can assure you that the program did not "somehow, inexplicably" "transcend it's programming". that is magical thinking.
while i haven't looked at it, and i'll present some caveats about that, what happened is pretty obvious:
1. you asked it to solve a mathematical game and it did. that's just a math problem. it's just algebra. if you didn't want it to do that, you should have been more precise in your language and you should have introduced better error handling to stop it from crashing.
2. it downloaded a strategy, it identified it as optimal and it followed it.
for example, you could download a strategy for a video game. or you could download a war game - including objectives, tactics, failed tactics and outcomes. the system would have access to all of that. it wouldn't even need to solve the game if it had instructions on how to break it.
i've looked at over a dozen of these now and on example after example i just see bad programming that wasn't tested, or was being tested, and bad programmers that don't understand what they're testing. i haven't seen a single example that holds up to the fears of transcendence. every single example is, when looked at more carefully, just an example of the ai doing what it was told, and the analyst or researcher not actually properly understanding what it told it to do. when the instructions are more carefully analyzed more correctly, what's exposed is a perfectly functioning computer and a bad researcher with incompetent programming skills.
as mentioned before, it does expose a threat, which is human stupidity and human error feeding bad instructions into a system and resulting in something accidentally getting blown up, which is the same realistic threat we faced around nuclear war in the cold war, and is probably more of a serious threat than human malice. the culture in the software industry is that you do loose testing in house and throw it out in the wild - you let the users make the mistakes. that's not a good idea here. the companies need to figure these mistakes out first, before they let the systems loose, and there should be government regulations around goals. the government is itself not a good institution to get too modular on this. i don't think evan solomon understands very much about ai, i think he gets briefed and says what he's told to say. lawyers aren't the right people to regulate scientists, but they can come up with a list of things that needs to be dealt with and set up regulatory bodies that can oversee it, like they would do with food safety or nuclear safety.
understanding that the primary threat is human incompetence is paramount in arriving at the right solution. schizophrenic fantasies about the systems being like children that will grow up are not helpful.