Hacker Newsnew | past | comments | ask | show | jobs | submit | capten's commentslogin

Imagine telling someone they're wrong without providing any evidence or context.

Welcome to the internet.

I expect this decupling will affect certain users. Can certain users speak up about this? I'm certain that certain users use Manus and are users on this certain site.


Likely a bad translation from the original Chinese version?

That being said it’s the first time I see Manus, they seem to have the perfect AI startup tagline for a new season of HBO’s Silicon Valley:

“Our Mission - To extend human reach by giving everyone the code to leverage their life. Our Product - We build general AI agents as the Action Engine for life.”


> Likely a bad translation from the original Chinese version?

No, most likely they're referring to users who created accounts etc on the meta infrastructure which is being spun down and erased and accounts created prior were never in their purview.


It's so weird to me that the benchmarks remain so low, but the models are marketed as revolutionary. And if you say that low coding capabilities aren't a problem, say that to the token price hike and 'general use' model setup.

Why not sell it as a math agent? Why do I have to set up 4 agents to check each others' work?


from what I understand, it's because unlike the other models, MAI models haven't yet fine-tuned against the synthetic datasets specifically designed to boost the benchmark scores.


It’s about bang for buck. That high a score for 5B params is pretty good, nigh unbelievable a short while ago.

It is my belief that smaller models will get better and better, and even cloud SOTA models will shrink.

Yet another reason the current buildout will feel like the railroads.


It's 5B active params in MoE, not 5B total params (total is 137B).


> It’s about bang for buck.

Hard to know when they don't give the price per token. Presumably it will be comparable to a low-mid range model in terms of price. But otherwise their 'Ideal Zone' is meaningless without factoring in the price per token. I don't how much tokens are being used, that's an implementation detail to me. I care about price / performance / latency.


https://docs.github.com/en/copilot/reference/copilot-billing...

Model Input Cached input Output MAI-Code-1-Flash $0.75 $0.075 $4.50


Yeah the future is probably a number of highly specialised small models you can run on your own hardware rather than massive frontier models in the cloud.

That's what I'm betting on anyway.


That seems to be what Microsoft is betting on also based on what was shown at the BUILD keynote today + that new surface ultra and the surface mini PC with the new Nvidia chip. Nadella really played up local AI as the main use case they have in mind.


MOE basically work that way already, QWEN/etc with low active params (A-number in name) allows to inference big models locally (only active params have to fit into memory)


Step 3.7 Flash on my Asus GB10 based mini pc is incredibly close to that today. I’m very impressed, and that’s without MTP to boost performance


The SOTA models will not shrink, because the problems will get bigger, from "write me a C compiler" to "clone Stripe business and run it".


There will always be tasks that are withing reach of whatever the SOTA models are, but not of the cheaper, perhaps locally runnable ones. It seems that already people are finding Qwen 3.6 27B sufficient for many coding tasks (the llama.cpp author is now using it exclusively).

As models get better and smaller, I expect that we will rapidly (within a year?) get to the point where SOTA models are not needed for the vast majority of coding tasks, and even today it seems many people are just using them for the planning phase.

How many people drive Ferraris vs Fords? How many people driving a Ford would, on a utilitarian basis, be any better off driving a Ferrari?

So far there seems to be mainly two high volume use cases that have been found for LLMs - coding and business flow automation, and it seems neither of these need SOTA models. I wonder if there will continue to be enough market demand for massive expensive SOTA models to make them worthwhile developing?


I don't think the paper is saying hallucination is limited to LLM's. The decoupling that needs to happen is the subconscious notion that LLM's are computers.

Computers are predictable and calculated (still based on a person's design), but LLM's are unreliable and unpredictable. If the general populous assumes LLM's are as trusted as a calculator, we're in for a bad time.

I hope there's a better way to prove that beyond unrealistic stress testing while speaking in absolutes, but honesty needs to come first on all sides.


I've been working on a ttrpg site for the last year with my own IP. The intent is to create an experience that makes an intuitive UI that minimizes tedious tasks.

You need to know "does this guy look hurt"? The enemy HP bar can be set to either an actual percentage, or set to have cracks in the bar to signify a range of damage. Does only person take notes? Personal notes are shareable and there's a section for community notes. Do you have enough perception to notice a hidden door? The UI can be set to go off passive perception and give you those notifications automatically.

It's still in early alpha testing with friends, but it should eliminate general GM pain points to encourage more groups to form.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: