Hacker Newsnew | past | comments | ask | show | jobs | submit | okdood64's commentslogin

I'm all for regulation of AI, but it's not practical.

I fear we're going to regulate ourselves to death (hyperbole) with regards to AI in the Western world, while China is just going to go brrr.


I've been using Gemini for the last 15 minutes. It's been fine for me.

Please, this is not Reddit. Don't bring this here. And to top it off, this isn't even funny or original.

I found it funny. Please do not tell what is funny and what isn't. Your comment is the most Reddit like here btw. Hahaha.

What else are we going to do when LLMs are down? Need some fun in here.

Thought it was worth a read.

You actualy think METR is complicit in this marketing stunt? If OpenAI was withholding data, do you think they would not call it out?

I think that the incident itself is a stunt, even if it may not have originally been a deliberate choice on OpenAI's part. Never let a good crisis go to waste.

I disagree. As the saying goes: never attribute to malice what can be adequately explained by incompetence. That goes for both OpenAI and HF (but mostly the former, as the latter was the victim).

You mean by Judge Leonie M. Brinkema who was appointed by Clinton and has ruled against Trump policies/agenda regularly?

Respectfully: If it's able to deploy without you asking it to, that's a you problem. There are no safeguards?

That's exactly how you get 'you are right, I deleted the production DB to apply the new schema when I should have written a migration'

That said, I do trust Opus and Fable enough to let them deploy to staging. Great for debugging. Just don't give them keys for prod


I told it “don’t betray me” in my prompt and it still stabbed me in the back.

I'm terrified to seed the RNG with words like "betray". I'll keep those way, way down the list of likely words.

You need more safeguards for sure, but also it tends to fly off down rabbit holes, rebuilding things in dumb ways, hacking around things, making assumptions etc, it seems very eager to go 'ta da! I did it look how quick I was', sometimes it nails it other times it created a lot of tech debt.

Also if it ever says, "I've found the root cause of ..", it definitely has not found the root cause and is making a non evidence based guess as it has run out of ideas.

100% my problem, but it's the only model I tried that does that so recklessly

> but if, in 2026, you aren't able to have an LLM generate decent quality code... IDK what to tell you.

I think the author's point still stands. At some point you need to determine whether the output is of decent quality. A lot of people don't have the ability and experience to do that.


> are they not even a little curious

Some are. Some aren't. Usually a good indicator of career trajectory.


I also switched to 5.6 Sol for this very reason. It was so exhausting and cringe to read.

> I lowball them every time.

And?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: