Hacker Newsnew | past | comments | ask | show | jobs | submit | camkego's commentslogin

If Anthropic really wants to make a statement they could independently pace their own model development, and ask others to make the same pledge.

Somehow, I suspect that won't happen.


If the marketer is an employee or a consultant, is it in their interest to show that the ad-spend they are controlling is high ROI, or low ROI.

Maybe this is a cynical take, but, if they get to the bottom of things, and show their boss/client that the ad-spend is not returning so much, it seems it would portend bad things for the marketer.

I really don't know, and it seems like a very hard problem.

Maybe this is the time for that Upton Sinclair quote: "It is difficult to get a man to understand something, when his salary depends upon his not understanding it,"


That's not how ad performance is measured. Your ROI is based on end conversions. You need to know if your leads are _good_, not just plentiful. A company that doesn't do this at the start will figure it out pretty quickly.

The final paragraph says this, among other things:

"The court ruled on a preliminary injunction request, so it’s not the final word on the merits. Still, it seems highly likely that the TWEET term and the bird logo have been freed from X’s trademark clutches. If so, it’s nice to get some cultural assets back into the public domain"

Seems kind of dubious to say the least.


This really does seem to call into question the credibility of this study.

Why would an rando anecdote (which another sibling comment has proven does not reflect the broader reality) call into question an actual study? The latter may be garbage, or not, but it starts with a much better claim to credibility.

I would say it's not really a "rando anecdote", but the whole basis of the relationship of the first actual study in the paper.

To quote the paper:

"Our first two studies were naturalistic field studies, and examined whether upper-class individuals behave more unethically than lower-class individuals while driving. In study 1, we investigated whether upper-class drivers were more likely to cut off other vehicles at a busy four-way intersection with stop signs on all sides. As vehicles are reliable indicators of a person's social rank and wealth (15), we used observers’ codes of vehicle status (make, age, and appearance) to index drivers’ social class."


Calculus can get pretty heavy, but I really value the comfort it gave me of the concepts of velocity and acceleration, (not to speak of higher orders) of which I am sure I wouldn’t understand nearly as well without the calculus background. I’m sure for most of HN velocity and acceleration, etc seem like super basic concepts, but I just don’t think I could apply the mental models as easily as before I took calculus.

The Bank of Canada at 234 Wellington Street. If you try standing around taking photos for over fifteen minutes or so, let us know how that works out.

I find this very interesting, I wonder if there is a public benchmark that reflects this “red team coding critique” aspect of the current SOTA model that reflects what you have observed.

It would be really useful to observe this in a benchmark vs. the more common “go implement this, or fix this bug” type benchmarks that seem to be prevalent.


Yeah, my tool to automate these review loops is https://github.com/wwind123/coding-review-agent-loop . It's basically a script calling Claude, Codex and Antigravity CLI's. The benefit of using CLI's is, the tool uses quota in your subscription plan of these AI providers, which is much cheaper than using extra tokens from the same providers to do the same thing.

A couple of months ago (before opus-5 and gpt-5.6 sol), The ratio of problems caught by codex/claude vs gemini was more like 2:1 to 3:1. But now it seems codex and claude have made huge leaps and gemini is more or less staying put.


Amazingly, these few days the Gemini 3.8 Flash (High) has been catching much more problems in code reviews than before. I think it started from the second day since I posted the observation above. Maybe somebody from Google saw my posts and tuned some knobs in the model to allow more critical thinking?

Another observation, Gemini's review on code is more critical now, but its review on design plans is still quite agreeable - it tends to approve Codex's design plan immediately, while Claude could often pick out a bunch of problems in the design plan in the first round of reviews.


Honestly, TL;DR

I wrote a summary for you...

"Because they receive a rebate, credit rewards-card users often effectively pay less than the posted or cash register price for equivalent goods or services."


Interesting how there is possibly a new category of residential proxy malware “automotive proxy malware”


This is just another mobile device. It even has a SIM card.


> So they _are_ going to train on them, no matter how many checkboxes you tick to stop them.

This is the part that is truly scary.

Nadella the CEO of MSFT wrote that post “ A frontier without an ecosystem is not stable”

It seems to me that the unspoken assumption in that post is that no matter what happens they’re gonna be training on your data.

He’s the CEO of Microsoft, he knows how these decisions go down, he knows how the world works, he is sending a warning.


I read it quite differently.

The frontier labs can have the models but without being where the workers are they cannot do much, lots of industries have strict requirements of not sending their data over to randos in the internet.

Microsoft and Google have the upper hand here with their workspace offerings and could easily position themselves as secure enclaves where you can use local LLMs where your data never leave your premises and is never used for training.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: