Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • What happens when AWS interrupts a Spot instance? I don't see an ASG in the architecture — does the controller detect the terminated worker and provision a replacement automatically?
  • You got me interested, a rough table of model size to instance type to spot $/hr would help a lot. The 0.5B example is CPU only, so it doesn't say much about what a 7B or 70B actually costs.
  • Basically the CLI

    1. compiles your model code and dependencies without requiring Dockerfiles or K8s

    2. provisions spot or on-demand instances directly in your own AWS/GCP account (via SkyPilot)

    3. it then spins up the runtime and gives you a production-ready URL endpoint

    I wrote this originally to replace Modal and Baseten, which I wasn't too pleased with. I needed to deploy open-source models inside my own VPC without vendor lock-in or proprietary Python decorators.

    Thought the community here might like it. AWS is an amazing cloud - with one-command deployment and scale-to-zero I was able to consolidate all my LLM serving into my own infrastructure for a fraction of the cost.

    Would love feedback. There are a few (many?) bugs and a lot of things to iron out, but I'm working with some close friends to make it awesome.

Explore Birbla archives

Self-host open-source LLMs on AWS with scale-to-zero · Birbla