Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • You might like webml-kit https://npm.im/webml-kit
  • Is there any use case/demo of lighter models like using embeddings for full text search, STT, OCR etc?
  • I love it! It works fine for me on Google Chrome on my MacBook Pro 14-inch M5 16GB RAM and 10 Cores CPU and 10 Cores GPU, with Metal 4 support
  • Oh nice, this actually works in Safari on my Mac! Prefill: 181.2 tok/s, Decode: 58.5 tok/s on my M4 Max -- not too impressive for 1B, but definitely impressive that it runs in my browser!
  • I really enjoy this engine. I’ve used it for personal projects, but it hasn’t been updated since Gemma 2. I suggest using Transformers.js instead these days.
  • It is kinda obvious, but maybe that's why it's not stated anywhere: each browser session will result in a download of 500 MB to ~1 GB, depending on your model selection. So, it's better to add a disclaimer if you end up using WebLLM in a customer-facing site.
  • Project is de facto dead, used it for many years and had to rip it out 6 months ago, don't waste your time.
  • This seems to be the demo:

    https://chat.webllm.ai/

    I am getting:

        WebGPUNotAvailableError: WebGPU is not supported in
        your current environment, but it is necessary to
        run the WebLLM engine.
    
    On both, FireFox and Chromium on Linux.

Explore Birbla archives