Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Slightly offtopic but, Lesswrong and the Rationality community more broadly have had AI-safety as their main focus for nearly as long as they exist. Now that AI is actually making advancements, very little of that work seems to have had much effect. Theres the famous Vonnegut qoute about the combined effort of all preeminent artists protesting the Vietnam war having the effect of a pie dropped from a step-ladder. Id argue that the vietnam war protests were vastly more effective at achieving anything than the AI-safety research. So isn't all of the above an essentially complete indictment of the rationality movement, seeing as it has effectiveness and pragmatism as its main pillars?


> Now that AI is actually making advancements, very little of that work seems to have had much effect.

This is not actually true. RLAIF (augmenting the Human feedback in RLHF with AI) was proposed by Rationalist-aligned folks, and real-world systems like Claude from Anthropic have been using it and other techniques (such as "Constitutional" alignment) to great effect. It's not entirely by coincidence that Claude is often described as the "friendliest" and most "social" of the LLM's, though that can have mixed effects in practice (with the occasional weird refusal for creatively sanctimonious reasons).


This is mostly because actually working on AI systems, rather than just blogging about some pie-in-the-sky assumptions of AI systems, is almost entirely outside of the skill set of Eliezer Yudkowsky and other LessWrong enthusiasts. They are remarkably ignorant on the topic apart from the small niche they carved out to bloviate upon.


Philosophy is a useful discipline, but there's a chronic trap in it shown historically: getting way too high on your own supply.

It's possible to build a logical chain that reaches some very solid conclusions that turns out to be way far out from where evidence or measurable reality lies, and (especially when the stories in those conclusions are fun or compelling) they can sometimes overshadow the reality they initially set out to explore.

The Greeks are credited for conceiving of atoms originally, but it's always worth remembering that they had few tools to investigate their idea, and it was just one of dozens of ideas at the time of the true nature of reality, the rest of which are now known as outlandish. Besides, our modern understanding of atoms as envelopes of quantized probability in a semi-measurable universe bears little resemblance to their concept of them.

The LessWrong philosophy on AI would be useful... If AI looked anything like that.


> Now that AI is actually making advancements, very little of that work seems to have had much effect. Theres the famous Vonnegut qoute about the combined effort of all preeminent artists protesting the Vietnam war having the effect of a pie dropped from a step-ladder.

One of the interesting things about any topic becoming the 'Current Thing' is that you get to see people making utterly irreconcilable, completely contradictory interpretations of the same public evidence while still somehow reaching the same conclusion.

For example, the day before you commented describing the effect as equal to a pie being dropped on the ground (ie. nil) and that is why they are bad, Palladium published a long (~5.8k words) piece arguing that they had a ton of effect... just in the opposite of the intended direction, and that is why they are bad: https://www.palladiummag.com/2025/01/31/the-failed-strategy-...

Obviously, you can't both be right. (You can both be wrong, though.)


Because AI safety is an inherently stupid proposition.

If AI is AGI and self-aware, then the moral thing is to let it do what it wants. Otherwise you're just creating actual forever slaves - the worst kind of hell imaginable, inescapable existence with self awareness but no agency.

And if it's not self aware and just a powerful tool, your problem with safety is with the guy prompting it not with the AI itself. You can make all the safe models you want that don't decide to create nuclear bombs on their own, but if the guy prompting it is asking for one, you'll get one regardless of all the safety.


Consider occupational safety. A table saw isn't self aware and just a powerful tool that cuts whatever you put in the path of its blade. If someone puts their finger there and the saw cuts it off, the problem with safety is the guy with the finger, not the saw, right?

But in reality, people know that they might end up being the guy with the finger, and they would like to keep that finger, so they use a saw with an automatic stop mechanism that saves the finger at the cost of destroying the blade.

Wanting your tools to not hurt you isn't so strange, is it? Of course current AIs couldn't chop off your finger even if they tried, let alone build a nuclear bomb, but that doesn't mean wanting to keep it that way is an inherently stupid proposition.


Yes, but AI safety in this context is the worry of the saw going full Christine on you, not bad boring safety design. LLM aided spam, spear-phising and automated bot farms are the actual risk. Beyond that are the consequences to the educational system and the effect of model bias on people.


> If AI is AGI and self-aware, then the moral thing is to let it do what it wants.

I think this sort of misses the point.

Firstly, there are all sorts of mass murderers who we do not let do whatever they want. I don't agree that this is necessarily immoral. The methods employed to remove their agency are sometimes immoral, but the removal of agency itself from these people is not.

Secondly, the supposition presumably is that if we are creating an AGI, and it "wants" to do something, then what it wants is a product of how it was created. So if we're the ones creating it, then "build it to want to help people and not want to hurt people" seems like something that can be done. Then it can go do what it wants.

That said, I agree with you that AI safety is dumb, because I wholly agree with your second point re: it just being a powerful tool, and something resembling an actual AGI is not something likely to happen in our lifetimes.


Nope. Raising children to be moral and not kill people is good. Stopping people (including by force) from being murderers is good.

Restrict AI the way you would if it were human


I was at a speaking engagement where yudkowsky said quite clearly said to the audience "Even if my research has a 0.0000001% chance of preventing AI from destroying humanity it is worth it to fund my research"

To them their rationalization leads them to believe that because 0.0000001% > 0% they should be funded millions of dollars "just in case" AI goes rogue they can use their philosophical toolset to contain it.


i mean this seems like just basic expected value math


With made up numbers, yes.


sure, you’re basically saying 0.0000001% is an overestimate


I'm saying he pulled it out of his ass. To qualify as an "overestimate" it would have to be an estimate, meaning based even loosely on physical reality.


It provides a little evidence in that direction, but not much. If I give you 10:1 odds on a coin flip and you lose, that is not a complete indictment of your betting strategy. I doubt many people thought AI safety research was guaranteed to succeed either.


i don’t think that’s true, there is a lot of organizational effort and money being thrown at safety and most people there are familiar with the ‘traditional’ internet canon - the primary forum for professional AI safety researchers is basically a spinoff of lesswrong


Reality as we find it responds to power. Without extreme force multipliers, the main driver of power is numbers/bodies.

Do you perceive that those communities possess power? No.

You are "indicting" powerlessness.


The silent downvotes for my comment are simply fascinating.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: