Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • > The problem with a vector primary index

    We've realized this a long time ago at TopK and built a flexible serverless search engine from scratch. Supports dense/sparse vectors, late interaction, lexical search, indexed regex, filtering, and custom scoring in one query.

    - https://www.topk.io/blog/vector-dbs-are-the-wrong-abstractio... - https://www.topk.io/blog/topk-embed-v1

  • Dashboard was last updated on September 7, it started on September 5. Bug? Or no progress? It's linked in the blog post so would expect it to work: https://turbopuffer.com/v3
  • I've really liked lancedb for similar use cases. Not just that it is OSS. But Lance treats ANN as a secondary index similar to what turbopuffer v3 does. Rows sit in fragments, and the vector index never moves them.
  • I’m developing a local “code graph mcp tool” (not yet published) and followed a similar path, though I may have been able to go further since I have fewer vectors in my database (even on projects with 50M LOC).

    At first, I tried all those popular vector databases and was disappointed with their performance. In the end, the best and fastest solution turned out to be building a multi-database system on SQLite, compiled with everything related to multi-client operations removed. Only exclusive mode was left. Everything is as binary as possible. The index is completely separate — an IVF with pre-training — and is built on the GPU (250K vectors are built, processed, and saved in 4 seconds). Right now, my biggest problem is frequent data changes, and I need to implement optimizations to reduce recalculations.

    So far, I haven’t seen any vector database implementations that are heading in the right direction. Maybe only Lancedb looks promising, but it’s too heavy for my needs.

  • AI has some of the craziest up and down cycles of tech I've ever seen
  • Soon they’ll just sell you markdown.
  • Vector databases were always more about retrieval than either vectors or data storage. But the term stuck all too well and companies held on to it a tad too long. Sorry :)
    by gk1
  • > This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.

    > don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.

    This is a direct parallel to how Postgres and Mysql built indexes.

    Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.

    Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.

    Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.

    This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.

    I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).

    But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.

    [1] - https://www.uber.com/us/en/blog/postgres-to-mysql-migration/

Explore Birbla archives