I think the problem is that you can’t really make a LLM that can tell the difference between fantasy or reality. It doesn’t know if someone is just roleplaying hell it doesn’t really “know” anything, just to be clear.
Sure you could try to have it pluck out keywords or phrases but that’s nothing guaranteed. I’m not a lawyer but I bet also just acknowledging that shit and trying to put up guardrails against it would open the company to liability if someone self harms because they took the output of a computer program as gospel.
Remember, to the AI everything is a hallucination (or everything is real, pick which lens you want to use). Put another way, it has no mechanism to tell reality from fantasy or to contextualize things in a broad way.
I do think that there’s room for some sort of compromise, maybe the tools periodically remind you they’re not sentient and give a little more info on how they work. Maybe if enough “disturbing” phrases or keywords are input or output it throws a reminder about mental health wellness and talking to a professional.
At the end of the day though, how can these AI companies be responsible for the mental health of their end users? Like, yeah the thing shouldn’t tell people to kill themselves but it’s not sentient it’s just modeling what it thinks a human would tell another human. Maybe as time goes on we will introduce better protections but I just think it’s borderline impossible to have significant guardrails if you want the LLM / “AI” to perform properly.
Idk I’m not a great coder or programmer or anything. I know some and I’ve tried to learn more about how AI actually works and it’s pretty fucky wucky and even the people working on it go “idunno” sometimes when asked why something works the way it does.
You’re underselling what companies can do. An LLM can’t verify reality the way people imagine, but it can detect patterns that correlate with crisis well enough to escalate, refuse harmful advice, or encourage real help. Seatbelts don’t prevent every crash either: they still reduce harm.
I did say in my comment I’m for some guardrails of sorts as you suggested. Just not sure if a company is going to put time and resources into that - probably not voluntarily at least. Or if it’ll compromise how “good” their product is. I’m starting to think AI might be more of a mirage at this point that’s going to implode but I also kind of don’t want the economy to explode at the same time?
I think the bubble will not burst but deflate long before the tech disappears: Plenty of companies will burn cash and investors will get wrecked, but the genuinely useful products will stick around. Guardrails will improve too, because trust eventually becomes a competitive advantage.
Guardrails will improve too, because trust eventually becomes a competitive advantage.
Thing is, as an LLM user you don’t want guardrails usually. Someone else wants it to have guardrails so you can’t get all the information you want.
If you’re reverse engineering something, you don’t want the guardrails to say “hey this is copyrighted, I can’t do this”. I’m assuming terrorists don’t want the guardrails to stop the LLM from giving them info on chemical weapons. And if you’re religious, you probably don’t want guardrails to stop the LLM from talking about religion with you
You’re describing different guardrails: Most users don’t want arbitrary censorship, but almost everyone benefits from guardrails that reduce clearly harmful or illegal outputs. The real debate is not whether they exist, it is who decides where the line is drawn.
I think the problem is that you can’t really make a LLM that can tell the difference between fantasy or reality. It doesn’t know if someone is just roleplaying hell it doesn’t really “know” anything, just to be clear.
Sure you could try to have it pluck out keywords or phrases but that’s nothing guaranteed. I’m not a lawyer but I bet also just acknowledging that shit and trying to put up guardrails against it would open the company to liability if someone self harms because they took the output of a computer program as gospel.
Remember, to the AI everything is a hallucination (or everything is real, pick which lens you want to use). Put another way, it has no mechanism to tell reality from fantasy or to contextualize things in a broad way.
I do think that there’s room for some sort of compromise, maybe the tools periodically remind you they’re not sentient and give a little more info on how they work. Maybe if enough “disturbing” phrases or keywords are input or output it throws a reminder about mental health wellness and talking to a professional.
At the end of the day though, how can these AI companies be responsible for the mental health of their end users? Like, yeah the thing shouldn’t tell people to kill themselves but it’s not sentient it’s just modeling what it thinks a human would tell another human. Maybe as time goes on we will introduce better protections but I just think it’s borderline impossible to have significant guardrails if you want the LLM / “AI” to perform properly.
Idk I’m not a great coder or programmer or anything. I know some and I’ve tried to learn more about how AI actually works and it’s pretty fucky wucky and even the people working on it go “idunno” sometimes when asked why something works the way it does.
In a sane world, this would mean no more LLM
You’re underselling what companies can do. An LLM can’t verify reality the way people imagine, but it can detect patterns that correlate with crisis well enough to escalate, refuse harmful advice, or encourage real help. Seatbelts don’t prevent every crash either: they still reduce harm.
I did say in my comment I’m for some guardrails of sorts as you suggested. Just not sure if a company is going to put time and resources into that - probably not voluntarily at least. Or if it’ll compromise how “good” their product is. I’m starting to think AI might be more of a mirage at this point that’s going to implode but I also kind of don’t want the economy to explode at the same time?
I think the bubble will not burst but deflate long before the tech disappears: Plenty of companies will burn cash and investors will get wrecked, but the genuinely useful products will stick around. Guardrails will improve too, because trust eventually becomes a competitive advantage.
Thing is, as an LLM user you don’t want guardrails usually. Someone else wants it to have guardrails so you can’t get all the information you want.
If you’re reverse engineering something, you don’t want the guardrails to say “hey this is copyrighted, I can’t do this”. I’m assuming terrorists don’t want the guardrails to stop the LLM from giving them info on chemical weapons. And if you’re religious, you probably don’t want guardrails to stop the LLM from talking about religion with you
You’re describing different guardrails: Most users don’t want arbitrary censorship, but almost everyone benefits from guardrails that reduce clearly harmful or illegal outputs. The real debate is not whether they exist, it is who decides where the line is drawn.