Hacker Newsnew | past | comments | ask | show | jobs | submit | more unstablediffusi's commentslogin

that's not scraping, that's web search.

their scrapers wouldn't identify themselves


do you give uninvited strangers in your home the same level of trust you give your family and friends?


silencing the opposition creates an illusion of consensus. in the deluded minds of the terminally online, it is paramount to maintain that illusion.

in every remotely political discussion here, reddit opinions are allowed to be expressed as non-constructively as you please, but all dissent, no matter how factual and constructive, gets flagged within minutes.


They're not deluded. They're evil. By faking consensus you mint new converts because the false consensus affects the opinions of everyone new to the subject. The platform designers, moderators, etc, etc, know this and that's why they do it.


Apparently it's because the original headline had the unfortunate juxtaposition of "homelessness" & "experimentation"?

I wouldn't be so quick to call delusion/dissent when designers of our spaces have simply made it far too easy to turn private affects into public effects..

(& It might be rude of me to be so concrete.. so.. apologies)

https://news.ycombinator.com/item?id=44213954

& What if everyone starts camping on pristine beaches? That'd be something! To marvel at!


their consent was not required. https://en.wikipedia.org/wiki/Transformative_use

petabytes of training data are transformed into mere gigabytes of model weights. no existing copyright laws are violated. until new laws declare that permission is required, this is a non-argument.

>If this AI worked without training, no one would say anything.

adobe firefly was trained on licensed content, and rest assured, the anti-AI zealots don't give it a pass.

the copyright is just one of the many angles they use to decry the thing that threatens their jobs.


There is no final word on the matter yet and there are counterpoints to the "Transformative use" argument.

https://www.reuters.com/legal/litigation/judge-meta-case-wei...

> "You have companies using copyright-protected material to create a product that is capable of producing an infinite number of competing products," Chhabria told Meta's attorneys. "You are dramatically changing, you might even say obliterating, the market for that person's work, and you're saying that you don't even have to pay a license to that person."

> "I just don't understand how that can be fair use," Chhabria said.

https://ipwatchdog.com/2025/05/12/copyright-office-weighs-ai...

> Stylistic imitation even without substantial similarity would likely be implicated under such a [market-dilution] theory, which could be considered as a market effect under factor four that diminishes the value of the original work used to train the model.


that's one sympathetic judge's opinion vs written law.

I'm not American, but it's clear to me that it will be the supreme court that ultimately decides whether licensing is necessary or not. the parties involved - the megacorps with infinite money and the litigious publishers - won't settle for less.

and given that the ruling in favor of the publishers would all but kill the American AI efforts (good luck licensing millions of works needed to train a coherent model) and greatly benefit China (who doesn't give a fuck about IP), I find it highly likely that it will not happen.


yes, it's obvious to anyone who paid attention in the past century, and particularly in the past ten years or so.

the oligarchs, who are the power behind the throne in every country on this planet, don't care what color the underclass is - they simply want lower wages for their peasants and higher value for their properties. the politicians and the media they own will always find a rationale to do their bidding.


if that is any consolation, no one gives a shit about xitter's ToS either. it will continue to be scrapped by every major player.


How exactly is it being scraped? My understanding is Twitter and LinkedIn are both huge pains in the ass to scrape right now.


There's a number of companies out there, like "brightdata", which pay a small amount to app developers to install a native "sdk". That SDK mimics a browser, and makes requests as if the user's device is doing it.

Since it's using a large number of real user's devices, and closely mimicing real web browsers, it ends up looking incredibly similar to real user traffic.

Since twitter allows some amount of anonymous browsing, that's enough to get some amount of data out. You can also pay brightdata for one large aggregated dataset.

https://bright-sdk.com/

This is part of the AI revolution, user's devices being commandeered to DDoS small blogs and twitter alike to feed data to the beast.


>As a society, we have power if we want to use it

you have no power over China, which is what the GP implies


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: