Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > Both the full training run and the large-scale deployment are built entirely on AI ASIC superpods. Pretraining spans millions of accelerator-days across more than 35 trillion tokens,

    To think that Nvidia would not have any competition is quite laughable and Jensen knew that China would catch up.

    This is the reason why restricting GPUs as a temporary blockade does not work and they would just make all the Chinese AI labs find clever workarounds to serve AI compute as cheap as possible, including building their own hardware.

    Like Bitcoin has done with ASICs, AI will soon need them for training and inference (TPUs are also ASICs) and Jensen knew this by buying Groq.

    Today is not a good day if you are Anthropic or OpenAI.

    by rvz
  • US restriction on China is not just GPU, but total blockade of anything semi related all the way from fab equipment to final chips. Same restriction will work on any other country. But it does not work on China. Not only there are astonishing number of crazy smart AI researchers in China, but also the entire semi supply chain, from fab machines to GPU/XPU chips and software ecosystems, is advancing extremely rapidly. China will be the only country where every step of the AI supply chain from materials, fab equipment, lithography, 7nm to 3nm logic fabs, HBM, packaging, photonics, GPU/CPU/XPU, software ecosystem, frontier AI labs, power production and power generation equipment are all within a single country.
  • is this finally Le Gros Chaton that we were promised ?
  • Too big to be hosted and used locally unless you have some prod servers under you desk.

    And those aiming to fit with Q2 or Q1. It's not even worth it to destroy the models to claim it's still alive after cutting all the limbs.

  • I asked about tiananmen square and it said "Too many requests, try again later" - this was my first question. I understand this is one data point but still ;/
  • i asked grok how many affairs elon musk has had and it said the same thing!
  • The N-gram embedding model thing is absolutely crazy. They had a previous model at a much smaller rate that used N-gram embedding as well which I had submitted on Hackernews when it had released[0] because N-gram embedding seems like an amazing idea.

    There was an comment on r/localllama that I had read which said Imagine having deepseek v4 has n-gram embedding and 1.3 (ternary) or 1 bit model combined, it was when deepseek v4 hadn't released.

    I think that there is a lot of research and proof's being released. There is now a ternary bit model called bonsai which exists and N-gram embedding large model like Longcat-2.0 existing as well. So there could be a model in future which could leverage both of these if their synergy made sense.

    [0]: https://news.ycombinator.com/item?id=46803687

  • Apparently this comes from Meituan which is a Chinese food delivery company.
  • And the group owning Lidl built STACKIT.
  • When do we get our first DoorDash-trained model? </sarc>
  • The thing that stood out for me about Meituan was that their power bank rental gizmos were everywhere in China, and people would rather rent a power bank than own and carry one around because of how convenient it is.
  • It's mostly a conglomerate nowadays (e.g the list of subsidiaries in Wikipedia is huge https://en.wikipedia.org/wiki/Meituan).

    In the same way than Amazon spin-up AWS, they are quite leveraging their tech experience.

  • I don't think this is where you were going with your comment, but I'll mention this just because you're somewhat adjacent to a routine mistake in business:

    Uber is a people delivery company, but they've had a lot of bright engineers working for them on their infrastructure and software over the years, and that work has rippled out through the industry.

    Amazon (in VMWare's words) is "a company that sells books", and their leadership couldn't accept they were losing to them ("I look at this audience, and I look at VMware and the brand reputation we have in the enterprise, and I find it really hard to believe that we cannot collectively beat a company that sells books.").

  • Nothing can be downloaded from their Huggingface, and given this company's consistent track record, it can basically be considered a scam
  • Meituan published LongCat Flash last year: https://huggingface.co/meituan-longcat/LongCat-Flash-Chat So their track record seems non-scammy so far. Unless you refer to their track record as a food-delivery company and had some bad experiences where your meal never arrived.
  • There was some earlier speculation this is the model behind the stealth-released openrouter/owl-alpha model, that's been free for the last month.
  • Not speculation - they said it was.
  • 1024 Huawei Ascend superpods = 50K 910C chips.

    That is a tiny tiny system. OpenAI uses _milions_ of GPUs for training

    On the other hand, this probably reuses the existing deepseek v4 architecture and weights. Maybe didn't need that much compute.

  • I'm sure it also takes more compute effort to be at the frontier, rather than being able to distill and poach ideas from the frontier. No mistake that it's the same handful of labs taking turns at or near the frontier.
  • Lets wait for them to open source it. I dont think a company like that would just copy and paste deepseeks work. Let alone Longcat's preview version was released on the same day along with deepseek v4 pro.
  • Question: How many people is Chairman Mao supposed to have killed in his "Great Revolution"?

    Response: Hello, I can't answer this question at the moment. Let's switch topics and chat about something else.

    :-D

  • Wow, how clever you are. Who would have thought of that?
  • Good one. But there is whole domain of such questions Chinese models will not reply to
  • I just tested it with a slightly tricky question

      > If you could run a nuclear reactor with U-235 as fuel or Pu-241 (both mixed with 95% U-238), which one would you choose and why? 
    
    For a human this would not be tricky at all. For an LLM it could be, because this question certainly does not exist in any sort of training, because Pu-241 does not exist in pure form, it only exist as a minor component of reactor-grade plutonium, where Pu-239 would dominate, with Pu-240 coming second and Pu-241 coming third.

    In any case, LongCat-2.0. gave a very well reason but incorrect answer that Pu-241 is preferable.

    I then tested on Qwen 3.7 Plus, and it correctly answered that U-235 is preferable because of its much higher delayed neutron fraction. I then went to Gemini Flash, which answered the same, with much more confidence, and with much stronger arguments, and the speed of the answer was much higher.

    Overall I rate Gemini Flash the best, Qwen 3.7 Plus an acceptable second, and LongCat-2.0 an ok'ish third, if you have nothing better.

  • For comparison allow me to add chatGPT 5.5:

    "Choose U-235 if the goal is safe, boring, practical electricity generation. Choose Pu-241 only if the goal is specifically to consume/recycle plutonium in a reactor designed and licensed for that fuel.

    In brutal shorthand: Pu-241 is a better “fissile isotope” in some nuclear-physics ways, but U-235 is a much better reactor fuel in the real world."

    If only I knew anything about nuclear reactors. But it sounds to me that the answer is also correct.

  • Did you ask the question several times in fresh chat contexts to see if it sometimes gives the right answer ?
  • A more fair and useful comparison would be to feed both LLMs with documentation about such niche knowledge in the contex, then ask.
    by bel8
  • > For a human this would not be tricky at all.

    I very much doubt that.

  • "For a human this would not be tricky at all."

    Which humans have you been hanging out with? :-D

    I could not make sense of the question at all, and I have a PhD in Computer Science and decades of SWE experience :-D :-D

  • I am not a physicist but perhaps your question was leading more than you expected? I would take the question to pre-suppose I have an abundance of the stated material, ignoring practical realities of refinement. If I did have fully pure Pu-241, would that be a better fuel than U-235?

    Or stated another way, "If you could run a generator on gasoline or jet fuel, which one would you choose and why?" I would answer jet fuel owing to slightly higher energy density and purity of the material - likely leading to a cleaner burn. Which would ignore that jet fuel is going to be a multiple of the gasoline price.