Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

[flagged]


He's talking about Apache Spark.

While it can handle 100 MB easily there probably are faster ways to handle that small amount of data. But yes, Spark can handle many PB and doesn't require a ton of changes in the code as you scale up from say 10 TB to 100 PB. The underlying cluster would change, and the performance profile would change a lot (10 TB can be done in-memory ... many PB, not so much)


He's right. Pretend you have 100PB. Write code for that. It'll work for 100MB but have terrible overheads.


He's talking about writing it for Apache Spark.

Which really isn't intended for 100 MB (I bet I could write a unix pipe & filter script that's faster than Spark), but is intended for 10 TB through several PB.


Please don't make technical comments (or any comments) in this inflammatory, nasty way.

We're lucky that you didn't spark a horrible flamewar, but instead got patient, factual replies.


I know. And I'm sorry, bitterly sorry, but I know that... no apologies I can make can alter the fact that in our thread you have been given a dirty, filthy, smelly piece of technical argument.


I'm not following you here. The point is that you have a history of doing this and if you won't stop doing it you can't comment on HN.


What can I say, being an asshole is one of my diversity dimensions. We just had that training at work, it was very helpful.

Also, I hope you aren't a Python dev in your spare time. Guido would be appalled.

https://www.youtube.com/watch?v=EdzqTGmEcZE




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: