Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • This has scrolled past in my feed and every time I've read it as "Io_uring without Radiohead", and I mean you could but what would be the point?
  • I’m curious why the choice is between syscalls and, specifically, io_uring with O_DIRECT. AFAIK Turso is like SQLite and supports multiple processes accessing the same database, and I would expect buffering to be a huge win in some workloads. What’s wrong with io_uring without direct? There’s also the middle ground of RWF_DONTCACHE.
  • You can use io_uring very well with sequential files, where the kernel will do read-ahead for you, possibly helped by you with fadvise.

    O_DIRECT is to be used with files that will be accessed in a random order (random meaning that the order is not predictable by the kernel) and with buffer sizes per access great enough that you want to avoid their copying between kernel and application (e.g. at least a few kByte per access).

    Whenever O_DIRECT is used, the programmer takes responsibility to implement an adequate form of read-ahead, based on the access pattern that is predicted for the application, and which cannot be guessed by the kernel.

    If the programmer did not implement read-ahead, like it was the case before doing the update described in TFA, that was an incompletely written program. One should not enable O_DIRECT without a complete implementation for its requirements, as that can lead only to lower performance than standard I/O.

  • If you're doing contiguous readahead in userspace, why not just use preadv? It'll limit you to doing readahead up until the next resident page, but at least in my experiments in Marginalia's index, preadv beats io_uring in all cases you can use a single preadv call to do the full read.
  • I found out that using mmap and just telling uring to read form there to beat anything else.
  • Not sure what you mean. Nothing about preadv lets you indicate you only want to read what's already in the page cache. And io_uring and preadv aren't orthogonal - you can give io_uring a preadv op to do the scattered read instead of issuing separate read OPs although I'm not 100% sure how much of a win that is in practice.

    Also, I think you misunderstood the blog as it's describing application read-ahead which is what you have to do when using O_DIRECT.

  • Interesting article but it gave me a bit of a panic attack. Benchmarking (with TPC or otherwise) is NOT the way to determine the correct approach here; that is strictly only to be used for databases (typically RDBMS) effectively “owning” the complete hardware they are running on. An embedded database might be used in that manner if it’s operating as the backend for a pure crud application that performs ~zero server side rendering, parsing, validation, etc and is essentially just an async http-to-SQLite interface. But more likely than not, an embedded db will be used and deployed on machines (not necessarily even servers) serving many a purpose, and need to perform best both within the confines of the resources available to the machine and in relative terms, necessarily making tradeoffs that might sacrifice performance for “value” in terms of CPU or memory usage.

    This isn’t just with regards to benchmarking, it’s an essential consideration *any* time you are taking ownership of the cache away from the kernel, which is the only piece in the stack that has viability into the global state and can be trusted to give back memory under pressure to ensure everything plays nice together. It’s not limited to just databases or even just memory, for example FreeBSD has had greater than its fair share of issues that trace back to the ZFS having a separate cache from the kernel (despite the much tighter integration between the two and presence of various mechanisms to address pathological cases). You can also refer to any comparison or benchmark between the use of spinlocks vs mutexes: spin locks consistently perform better in/on (micro)benchmarks but are almost always actually the worse choice in the grand scheme of things because the benchmarks falsely assume complete and uncontended ownership over system resources.

    This isn’t even io_uring specific and I’m hardly the first to bring this up in the context of O_DIRECT.

  • What would be the correct way to do benchmarking/profiling here? I ask as I've started to run some for Vinyl cache to see what performance issues I can find, and I've quickly learned just how hard it is to get good, repeatable benchmarks that isolate the right thing.
  • TPC-H is the recommended approach in Turso's CONTRIBUTING.md - https://github.com/tursodatabase/turso/blob/main/CONTRIBUTIN...