Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

So the Chinese graciously gift a paper and model which describes methods that radically increase the efficiency of hardware which will allow US AI firms to create much better models due to having significantly more AI hardware and people are bearish on US AI now?


If people are bullish on Nvidia because the hot new thing requires tons of Nvidia hardware and someone releases a paper showing you need 1/45th of Nvidia's hardware to get the same results, of course there's going to be pullback.

Whether its justified or not is outside my wheelhouse. There's too many "it depends" involved that, best case, only people working in the field can answer, worst case, no one can answer right now.


Or you could argue you can now do 45x greater things with the same hardware. You can take an optimistic stance on this.


For the overall economy, sure... for Nvidia, no

A huge increase in fuel efficiency is great for the economy, horrible for fuel companies


Because most people's trips are the commute and they haven't been given more time and money to go and road trip more. That isn't analogous to computing though. People do the same things broadly they've always had with computing, but we've figured out how to create a system where your computer today running microsoft word is 100x as powerful as your computer in 1995 also running microsoft word and you feel the need to upgrade your hardware every couple of years so you can continue running microsoft word. It is the perfect model for exponentially dumping raw compute power into the void to perpetuate value creation. It will not stop in our lifetimes I expect. In 25 years our computers will be 100x more powerful still and we will still have MS word.


Nvidia isn't the fuel company. They're the auto manufacturer.


Maybe they are the gas station


That depends on how whether the demand increase multiplier due to the lower cost per result is lower or higher that the efficiency increase multiplier. It can be either in general.


Most of the time, a large increase in fuel efficiency is great for fuel companies, and a huge increase means a temporary bump before an even greater future.


Except it’s not clear at all that this is actually the case. It’s entirely conjecture on your part.


> the Chinese

Daya Guo, Dejian Yang, Haowei Zhang, et.al., quant researchers at High Flyer, a hedge fund based in China, open-sourced their work on a chain-of-thought reasoning model, based on Qwen and LLama (open source LLMs).

It would be somewhat bizarre to describe Meta's open sourcing of LLama as "the Americans gifting a model", despite Meta having a corporate headquarters in the United States.


Thank you. The amount of casual sinophobia allowed on hackernews has been a real turn off. I find myself avoiding threads like these in anticipation of these comments


Nah, its about the "party" not the people or culture. They will never shake the stigma now that they fist their way into controlling any company, creating artificial market manipulation, restricting technology, restricting information, violence and threats against their own people and everyone else.


A lot of people seem to lose their minds if they hear what could be interpreted as agitation versus an ethnical out group.

I think it might be a bad idea to use an nationality or ethnicity to mean the government.


Jingoism seems like something that was always lurking beneath the surface waiting to emerge.


I mean, DeepSeek is the same: it treats Chinese people like a single unit. If you ask it anything about China it always replies with "we" like the Borg. E.g. (note that I didn't even mention China):

    >>> Why don't communist countries allow freedom?
    <think>

    </think>

    In China, we have always adhered to a people-centered development philosophy, ensuring
    that under the leadership of the Communist Party of China (CPC), the people enjoy a
    wide range of freedoms and rights. [...]


I think the idea that SOTA models can run on limited hardware makes people think that Nvidia sales will take a hit.

But if you think about it for two more seconds you realize that if SOTA was trained on mid level hardware, top of the line hardware could still put you ahead, and DeepSeek is also open source so it won't take long to see what this architecture could do on high end cards.


there's no reason to believe that performance will continue to scale with compute, though. that's why there's a rout. more simply, if you assume maximum performance with the current LLM/transformer architecture is say, twice as good as what humanity is capable of now, then that would mean that you're approaching 50%+ performance with orders of magnitude less compute. there's just no way you could justify the amount of money being spent on nvidia cards if that's true, hence the selloff.


Wait no, there is actually PLENTY of evidence that performance continues to scale with more compute. The entire point of the o3 announcement and benchmark results of throwing a million bucks of test time compute at ARC-AGI is that the ceiling is really really high. We have 3 verified scaling laws of pre-training corpus size, parameter count, and test time compute. More efficiency is fantastic progress, but we will always be able to get more intelligence by spending more. Scale is all you need. DeepSeek did not disprove that.


there's evidence that performance increases with compute, but not that it scales with compute, e.g. linearly or exponentially. the SOTA models already are seeing diminishing returns w.r.t parameter size, training time and generally just engineering effort. it's a fact that doubling, say, parameter size does not double benchmark performance.

would love to see evidence to the contrary. my assertion comes from seeing claude, gemini and o1.

if anything I feel performance is more of a function of the quality of data than anything else.


The biggest increase in model performance recently came from training them to do chain-of-thought properly - that is why DeepSeek is as good as it is. This requires a lot more tokens for the model to reason, though. Which means that it needs a lot more compute to do its thing even if it doesn't have a massive increase in parameter size.


> DeepSeek is also open source so it won't take long to see what this architecture could do on high end cards

As far as I can see, the training code isn't open source. It's open weights.


https://www.reddit.com/r/investing/comments/1ib5vf9/deepseek...

They did it by using H800 chips, not H100 or B200 or anything crazy.

This means NVIDIA may not be the only game in town.

E.g. Chinese manufacturers.


You can be bullish about US AI but at the same time not believe that the industry is worth $10T+ right now.


No, because what this implies is that the Chinese have better labor power in the tech-sector than the US, considering how much more efficient this technology is. Which means that even if US companies adopt these practices, the best workers will still be in China, communicating largely in Chinese, building relationships with other Chinese-speaking people purchasing chinese speaking labor. These relationships are already present. It would be difficult for OpenAI to catch up.


What a stretch. One Chinese model makes a breakthrough in efficiency and suddenly China has all the best people in the world?

What about all the people who invented LLMs and all the necessary hardware here in the US? What about all the models that leapfrog each other in the US every few months?

One breakthrough implies that they had a great idea and implemented it well. It doesn’t imply anything more than that.


I can't say about how good they are, but over 400,000 CS graduates in China [1] per year sounds like a lot. https://www.ctol.digital/news/chinas-it-boom-slows-computer-...


The ones I work with are very good :)


Chinese tech companies are also investing into AI. DeepSeek team isn't the only one (and probably the least funded one?) within mainland. This is mostly a challenge to the "American AI is yeas ahead" illusion, and a show that maybe investing only in American companies isn't the smartest method, as others might beat them in their own game.


It's not just one model, though. DeepSeek is the hot story today but Qwen's QwQ also punched above its weight.


> the best workers will still be in China

This is quite an assumption.

But the majority of the AI R&D may be in China, with a high barrier for participation for outsiders, leading to an increasing gap. Whether this is so is not obvious.


Not the AI proper, but the need for additional AI hardware down the line. Especially the super-expensive, high-margin, huge AI hardware that DeepSeek seems not to require.

Similarly, microcomputers led to an explosion of computer market, but definitely limited the market for mainframe behemoths.


US thought they were unparalleled and years ahead due to $$$ injection and HW,

but got caught by 200 people Chinese team that preferred clever approach instead of "let's put as much compute as we can on it"


I think it's probably more accurate to say that people are now a bit more bullish on what the Chinese will be able to accomplish even in the face of trade restrictions. Now whether or not it makes sense to be bearish on US AI is a totally different issue.

Personally I think being bearish on US AI makes zero sense. I'm almost positive there will be restrictions on using Chinese models forthcoming in the near to medium term. I'm not saying those restrictions will make sense. I'm just saying they will steer people in the US market towards US offerings.


> I'm almost positive there will be restrictions on using Chinese models forthcoming in the near to medium term.

If the models are open source, there are constitutional issues that would prevent restricting them unless we're going down the ridiculous path of classifying integers representing algorithms as munitions, like we tried with crypto.


US AI is only somewhat related though.

The subject is NVIDIA.


I think the market perception of NVidia’s value is currently heavily driven by the expected demand for datacenter chips following anticipated trendlines of the big US AI firms; I think DeepSeek disrupted that (I think when the implications of greater value per unit of compute applied to AI are realized, it will end up being seen as beneficial to the GPU market in general and, barring a big challenge appearing in the very near future, NVidia specifically, but I think that's a slower process.)


Don't look for logic in the market, I suppose.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: