Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > Reading everything becomes the default. At a cent per document, a model can read every paper in a field, every record in an archive, every email, or every message in a support queue as a matter of routine

    Pretty much already happening

  • A year ago I had an aha moment, when I realized that for my purposes, Gemini Flash was not only 9x cheaper, but 3x faster than Gemini Pro, while producing identical output. Who's the best model now!

    For a lot of tasks, even small models have saturated them a while ago, and then going cheaper and faster is just pure gains.

    For coding I also prefer to do it interactive/realtime, micro-prompting, surgical edits, which the small models can handle just fine.

    And then at the top, the real question is consistency. Not "can they do it" but "reliably enough that you don't need to constantly double check everything." (In my experience, not quite there yet, although it's getting way better.)

  • I've been working with the chinese open models for 4 months. They are more than capable for a tiny fraction of the cost of the frontier ones. And yet they also continue to get significantly better and (Deepseek's recent price increase aside) cheaper. Its hard to fathom how the truly frontier stuff will be able to compete long-term.
  • And why won’t the frontier models continue to become better? The open models are getting better but so are the frontier models. The frontier models might remain in a constant race to remain ahead
  • There’s still 50-500x cost reduction in “this is only an engineering problem” low hanging fruit from specialized chips to run the models + improved distillation.

    Entirely feasible that by 2031, Fable 5 (or greater) intelligence level models will run cool on smart phones, if not sooner.

  • If fable 5 could eventually run on a smartphone, I wonder what we'll get out of datacenters. Musk and others are trying to build 100gw of compute by 2030. Will we get something 1-10 million times better than Fable 5, or will the parity gap between local and datacenter capability pair down enormously?
  • You're betting on getting getting ridiculously powerful chips to run on batteries in a tiny housing without cooling, while we can't even get enough RAM? It would be a terrible waste of resources. Now we already have TFLOPs wasting in our pockets and backpacks, then we'll have PFLOPs idling, because there's so much time between prompts. Much more efficient to batch it on a server.
    by tgv
  • I saw a 1B model yesterday that was fine tuned on Fable output. I found that hilarious, but it did actually make all the scores go up.

    (Actually talking to it, it was about as coherent as you'd expect, i.e. 3/10)

    The floor for "actually usable model" keeps dropping though. (Seems to be about 27B right now?)

  • 100x seems like an underestimate. Even with no model improvements, we should see that sort of reduction. Looking at TSMC’s margins, Nvidia’s margins, and OAI/Anth (alleged) margins on inference, there is a room for a 100x reduction.

    Right now all three of those are at abnormally high levels. Competition will come for all three.

  • Personally I still see LLMs as very advanced search engines which lack intelligence. To me it seems that the cost of getting data is reduced by LLMs, not the cost of intelligence. I mean: we tell the model what we want to achieve, and the model responds with the right data in de form of code in seconds.

    That's why 'stackoverflow programmers' will have a hard time competing with LLMs but engineers are still needed for their intelligence.

    Well that's just my 2 cents.

  • You're correct on the technology assessment.

    But you forget all development today is just searching for a template, copy pasting, changing some small things. And a smart-ish search engine can do all that.

  • I think this viewpoint fails to understand what "intelligence" is. The idea must be that intelligence is some special thing that only humans have. So when machines couldn't do jack s... we said "it's the Turing test". When machines blew through the Turing test we said "that was just prediction..not really intelligence, that's different".

    It's not different. The delusion humans have is that intelligence is special and magical. It's not. It's just nature's prediction machine. A very fancy version to be sure. But not qualitatively different .

    All statements that "oh but it'll never be able to do that" will prove false.

  • Sooooo what does it matter if we do 99% of things a LLM can just solve as a 'advanced search engine'?

    Btw. an advanced search engine is probably the worst comparision i have read so far.

    A LLM is a latent space which is capable of a lot of things a search engine can't do. It can apply different type of patterns and flows onto data, it can combine these etc.

    My 'advanced search engine' was just able to create a working PR for exactly what i wanted it to solve (fixing a bug) by analysing the bug, finding a valid solution then commiting the solution itself.

  • I’m loving small models. The gemma4 26/31b models have been deeply impressive on weird prose analysis tasks that I am working on. Nova-micro is really stupid but is extremely fast when it’s smart enough to do something. I’m trying to be disciplined about able to evaluate quality vs cost everywhere for real systems built on this stuff. I probably need to get off Bedrock because it’s missing a lot of other little models that might be good competitors.
  • I think what a lot of people miss about jevon's paradox is the elasticity of demand of the underlying resource

    textiles had jevons paradox, and many more textile workers were employed even when textile machines were being created, until we saturated the demand for cheap clothing in the world and then textile workers were kaput (same for farming, and horses)

    software is currently undergoing jevons paradox, but it's very unknown how high the ceiling of demand for software is. web dev might be doomed, but software in general i think is probably limitless

    Intelligence is also probably unbounded (atm software and intelligence are very closely tied together). its very possible token spend rides up the curve forever.

  • Brings up the question of what the intelligence is used for. Humans exploited intelligence for competition. With each other to wipe out other Homo species, mate more and collect resources, with other animals to limit predator impact and gain food. Intelligence will be used offensively by corporations and their people to extract more from consumers (make pricing opaque, terms of service more complicated, etc.) and scams far more sophisticated. The "consumer" will need extra intelligence to fight all that off. There are only so many meals you can expertly produce, shirts to fold, itineraries to fun places you can execute, but there's a practical infinity of traps to set and avoid.
  • I think a big one is robotics. A robot can today fold your laundry. It takes ~10mins per item. Seriously. It takes a long time to process the image find the corner move the claw to the corner of the shirt and attempt to straighten before folding.

    Robots right now generally move at glacial speeds. You might have seen robots doing flips in semi controlled environments but watch how slowly they open doors etc. processing time is a major bottleneck.

  • Wait til the robots get on Cerebras, it'll set your pants on fire.
  • Have LLMs improved at being able to process physics-based problems and environments? I remember that issue being discussed around generative gaming a while ago but I hadn't heard much about it recently.
  • Shirt folding is piker stuff. The acid test (imho) will be to cleanse a bathroom & shower (or tub) top to bottom. Sparkling and provably sterile. Special treatment for mold on fixtures and walls, and for leftover body stuff (blood, mucus, fungus, what have you).
  • You must not have been paying attention to development with robots, there are many videos of robots moving really fast in "non controlled environments"
  • Xiaomi robots do this in double digit seconds.

    Still slow compared to humans, but Chinese robots will be as successful as Chinese EVs, phones and solar panels.

  • Sensors are a huge challenge for robotics. We have very precise force-feedback on our joints, pressure and heat (temperature gradient) sensors all over our body, and our hands have a sensor density that allows us to count needle heads and detect the exact grip strength needed by feeling the micro-slippage of objects in our hands. Robots don't have that.

    You can do backflips with pretty much just visual sensors for your environment, a good IMU for your spatial orientation, and some feedback on the position of a small number of really beefy joints and the force exerted on them. Folding laundry and opening doors is much more difficult, and trying to compensate with mostly vision requires going slow enough that things have time to move over appreciable distances before you take the next adjustment

  • I think speed is actually going to be a bigger factor than cost. Even projects where “money is no object” often hit a wall with LLM response times.

    Sure you can speed things up with parallel work under subagents, but as with parallelizing traditional computational tasks, there are diminishing gains.

    I keep hearing people saying just change the way you work to trust long-running agents and multi-task more, because they’re too slow to work with interactively for many use cases. I think that’s painful in a world where we expect humans to still heavily guide and interact with agents for their day-to-day work.

  • > they’re too slow to work with interactively for many use cases

    This just demonstrates how much we already take for granted the LLMs that we have now. If you compare it to what we had before (hand the task off to a junior dev and wait for them to complete the work) then it doesn't seem slow at all.

  • Given the 750 tok/sec GPT 5.6 Sol Ultrafast (via Cerebras), the many-1000 tok/sec Chinese models, and the 15000 tok/sec Taalas HC1, I think we're well on the way towards seeing that solved too. Combine the two, and yeah, wild ride incoming.

    What's especially bewildering to me is that translated back to raw bandwidth, even 15000 tok/sec is just like what, 75 KB/s? Extremely meager amounts of data, moving mountains.

    It's already kinda funny seeing LLMs throw out effort estimates in wall time terms. It's always some "hours, days, weeks" tier thing, when in reality, it's gone and done in minutes.

  • > Reading everything becomes the default. At a cent per document, a model can read every paper

    I love how in our day "reading everything" means "the computer reads it for me".

    I expect soon the computer will be able to go on bicycle rides, and spend time with my wife.

  • Futurama did it first!
  • The next generation of bike computers will probably have links to an LLM cloud subscription service for real-time coaching and route planning.
  • Hasn't the computer already been spending time with your wife?