Notes on DeepSeek

Notes on DeepSeek

211 pointsby vinhnx141 comments

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Why would the agent send the results of the query "Show me my recent transactions" to LLM? This pretty deterministic results which involve no LLM interpretation or decision making.
  • wrong thread?
  • The more competitive the field is - the lower the prices should become anyways
  • US AI is almost a religious cult. It's devastating that they are treating it as a petty commodity
  • I would argue the US providers have gone full tilt into sales culture with respect to AI. Anything is said on a whim to redirect attention back from whomever is in the limelight. Initially I thought Anthropic was more pragmatic, but the constant release cycles of things that don't exist for most people, the gatekeeping, the statements made by Dario, it's all a part of large brand toxic sales and marketing.

    From the notes this part sat with me as the real difference:

    > As a whole, China seems to treat AI as just another technology, rather than as some kind of singularity moment. National attention is still on basic needs and infrastructure buildouts, and on providing more medicines for people. The “dreams of singularity" seem like a luxury or distant consideration.

    Meanwhile... In the fantasy land over here in the US we're constantly being told that it's "coming", "almost here", "too powerful for us to give you access to", "of national security importance!". Or... FUD.

    And while there may be trace amounts of truth in those overzealous statements we haven't seen a significant improvement in much outside of software development comparative to the spend and environmental impact.

  • Altman used to talk about making a religion and Dario Amodei constantly talks about "building a God" and meets with religious leaders including the Vatican.

    > It got me thinking, though--the most successful founders do not set out to create companies. They are on a mission to create something closer to a religion, and at some point it turns out that forming a company is the easiest way to do so. [1]

    [1] https://blog.samaltman.com/successful-people

  • Funny this was posted here the same day the Anthropic CEO posted a doomsday prediction begging for government regulation. I was curious how the Chinese feel about AI risks considering I would expect them to be more cautious than the Americans, but they clearly aren’t. Which indicates to me that the Anthropic CEO is probably just pushing for regulatory capture. I mean, maybe he believes what he is saying, but I don’t.
  • Deepseek v4 Pro is the first model I've sat down with in pi.dev and haven't felt like I've had to fiddle with the knobs to get working results.
  • Post appears to have been removed, I caught a copy of it: https://pastebin.com/rcAqEFG1

    I assume it will get reposted at some point.

  • thanks. this really isnt that long, might as well paste in full here since OP deleted.

    Notes on DeepSeek:

    We visited the company HQ last Tuesday. It was founded in 2023 by Liang Wenfeng and operated out of his hedge fund, High-Flyer, until somewhat recently. The company released their R1 model in January 2025, so it was interesting to see what they’ve been doing

    The company is located in an unmarked, 12-story building in Hangzhou. There is no DeepSeek branding visible from the street or lobby. I asked why this is, and the team demurred and said, “Well, there are many companies in this building, and we are not special.” They want to keep a low profile.

    We met with their Head of Data and Head of Infrastructure. The company only has 300 employees. They are at least an order-of-magnitude smaller than Anthropic, and don’t care to scale further just yet. Their Head of Infrastructure, in particular, was young; maybe 30 years old and apparently one of the best AI buildout and energy experts in the country. (We briefly walked through the labs, and everybody seemed young. There was a lot of discussion; it felt like an exciting and energetic place.)

    Lots of competition is coming from Alibaba (Qwen), ByteDance, and Moonshot (Kimi). People in China seem to mostly use Kimi or Deepseek. Young people use VPNs to access Claude, though Anthropic has blockers around usage in China and make it difficult. Poaching between groups is common, just like in the U.S. DeepSeek has a reputation as being really smart and “cool,” maybe similar to Anthropic. Big labs are mostly in Beijing, near Tsinghua and Peking University, with Hangzhou as the main exception (DeepSeek and Alibaba/Qwen are there).

    The DeepSeek team reads western AI writers. They listen to Dwarkesh and read Gwern. The people we met with said they had never met with any employees from Anthropic. They were not at all concerned with some kind of hostile / AGI takeover scenario. They kept bringing up job loss (which is already high amongst youth in China) as their main concern. When we asked if they do red teaming on their models, they said no. In China, AI models are not regulated directly; the government instead has restrictions on how those models can be used in software, services, etc.

    As a whole, China seems to treat AI as just another technology, rather than as some kind of singularity moment. National attention is still on basic needs and infrastructure buildouts, and on providing more medicines for people. The “dreams of singularity" seem like a luxury or distant consideration.

    We asked the DeepSeek team: “What has the highlight been so far? What are your plans for an exit?” And they said that their highlight and great achievement was R1. They did not gesticulate at a future model or vision, but rather seemed proudest of what they’ve already done. They are content for now to remain ~6 months behind U.S. companies while maintaining a lower profile and team size.

    by swyx
  • From the notes, they seem humble and empathic.

    We're lucky to have China imposing competiton to the western AI megacorps.

    If it wasn't for China, I would probably have to spend $100/mo on AI instead of $10 like I do currently while using DeepSeek and MiMo (opencode Go plan).

    And while I could do so comfortably, I feel for those who can't. It must feel incredibly isolating to only watch others have access to expensive models to leverage their careers.

    I hope SoTA AI becomes an universal right because it will contribute to too much income disparity otherwise.

    by bel8
  • Ollama cloud is also a fantastic deal. They also don’t have a monthly limit, just session and weekly.
  • I see this problem already for me.

    I have unlimited tokens at work than i go home what do i do? Spend 200$ per month? No def not.

    When Anthropic increased the limits for their 20$ plan, i started again coding with it on a private project and it was fun and i did a lot in that 4 weeks.

  • Yep. After yesterday's moves around "Fable 5" even twice as much.

    We've had a taste, and damned if I'm going to have the "means of production" snatched from me already?

  • > We're lucky to have China imposing competiton to the western AI megacorps.

    The second they get a hold of the market, Chinese Big Tech will be as bad or worse than US Big Tech.

    We're lucky to have DeepSeek.

  • Ever since I found Opencode Go AI coding is fun. I always hate the feeling of working inside a fenced constraint where if I just go hard enough I suddenly hit a wall and have to pay up a LOT more.

    It's crazy how much you get out from Deepseek V4 Flash alone.

  • "As a whole, China seems to treat AI as just another technology, rather than as some kind of singularity moment."

    This is a refreshing perspective.

  • The CCP knows, whatever the heck this technology will bring with itself, the current power dynamic inside of the country is on their side, and AI will solidify it.

    I hypothesize that, rather than slowly having it disperse in society and allow people to harness it in ways they don't want, they might as well accelerate everything until AI becomes the totalitarian swiss knife - which they can make use of in the best way of course.

    Let's see what will happen.

  • The whole "AI race" is a construct of American startup founders trying to get more money. The government picked up that line because it seems fun and useful to be "wining a race" against China. It's all nonsense. China doesn't care if they get AI first or second, they can replicate anything in a few months. They know it's only an excuse to get more money in the hands of billionaires.
  • I may have some explaination.

    China is an atheist country. The whole "creation" thing didn't even mean anything special to normal Chinese people.

    Chinese viewer have a meh reaction to Edward Scissorhands, they don't have a Frankenstein Complex.

    by est
  • Especially here on HN, where AI anxiety (especially amongst those that are really nervous that it needs to succeed) is very, very tiresome.
  • China is probably more capitalist in many respects than the west these days. AI, robotics and automation is a way to push into the future. In the west we have endless researchers stuck in a psychosis that they are talking to a sentient being.
  • "National attention is still on basic needs and infrastructure buildouts, and on providing more medicines for people. The “dreams of singularity" seem like a luxury or distant consideration."

    Further on. Refreshing indeed.

  • The CCP is very active in the matter of AI. In fact, the DeepSeek moment was responsible for Xi calling for a private meeting with tech bosses, including the exiled Alibaba founder Ma. Which is practically unheard of in China politics.

    I don't have enough information to say whether the Chinese leadership sees AI "just as the next technology" or they are more cautious due to its double-sword nature. But the immense efforts for building their own AI/GPU chips plus government's billions fund pushed for AI build out, a directive for fast pace integration on large scale and a sweeping national education reform for AI, I don't think it can be seen as similar to other ordinary techs.

    [0] https://www.reuters.com/world/china/china-prepares-295-billi...

    [1] https://www.globalneighbours.org/en/articles/china-unveils-n...

    [2] https://english.www.gov.cn/news/202606/10/content_WS6a296017...

  • I can't recall the scientist's name, but he said months ago that DeepSeek is best for Physics (maybe it was on The Diary of a CEO podcast). So I had a long chat about the Simulation Hypothesis, and I was really surprised by how good, deep, and straight to the point it was.

    What's brutal is that Google, which started this AI revolution, has literally the worst coding model! I tried 3.5 Flash last week (the stupid still pays for Ultra due to Google One's storage), and before I gave up on 3.1 Pro, I saw a coding agent hallucinate for the first time in months, even at the highest effort level!

    Meanwhile, I've tried DeepSeek with the DeepSeek TUI (now CodeWhale), and it didn't do any worse than Codex or Claude Code. I know there are benchmarks and all, some of them gamed, I'm sure, but in real-world experience, DeepSeek is absolutely amazing for its price! If you have software engineering skills and are not an accidental vibe-coder, honestly, try it out and stop burning money. I'm sure you will get even better results with OpenCode! Human Intelligence + Artificial Intelligence beats the highest AI model without the guidance of a HI!

    Meanwhile, I burned through my entire budget on the $200 Max for Fable 5, for a modest-amount project in Python using its own CLI coding agent. What a waste!

    I keep hearing "always use the bestest model" - no, always use the most practical one for the job! I got so many issues with Fable on a very small project that even Copilot found that it's simply not worth it for 99% of your tasks!

  • steve hsu?
    by HSO
  • Thanks for your comment, I especially like “If you have software engineering skills and are not an accidental vibe-coder, honestly, try it out and stop burning money.”

    I thought that using Opus with the Gemini Ultra subscription was in many ways awesome, but I simply feel happier using DeepSeek v4 flash with OpenCode (so fast!) of v4 pro when required.

  • Yeah Fable 5 is good but feels incremental and overhyped, also burned through my entire Cursor allowance in my Ultra plan in a single day. Ridiculous. They just want to create FOMO and appear mysterious so companies and users will feel so special for being allowed to use this model and pony up more money. After all they have to grow a few order of magnitude to pump their IPO valuation as much as possible, so I think this is just a strategy to justify their increased token pricing which starts to become absolutely insane. 10-20k per month per developer, do companies really think that's a good way to spend their IT budget? I assume 99 % of software shops wrtite run-of-the-mill web/mobile/desktop apps or some legacy backend APIs and CRUD code, you don't need a superintelligence to crank that stuff out. It sounds so ridiculous to have a model that supposedly can design biological weapons and then 99 % of users vibe code spaghetti Javascript with it. But the spice must flow!
  • It's just insane how different experiences are. I've let it spin for 2 days on difficult tasks for my job, it found very complex esoteric race condition bugs which were genuinenly the type of "will take you 5 days to figer out what went wrong".

    And on my personal account I was developing my own ORM in one session, on the other let it spin for 5h on "implement civilization II from scratch". It got quite far and everything it did genuinenly worked well, until they disabled Fable. The strange thing is that I was surprised by how little usage that actually costed, I didn't even hit the daily limit. Was expecting to be cut off directly with my subscription based on what people said yet it kept going and usage showed I had plently left.

    How come, we all have such different burn rates or even results?