Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • Linux is unusual in OS kernels in that direct system calls from arbitrary userspace code are supported and ABI-stable. This model has always been a terrible idea. It robs the system of an ability to intercept system calls in userspace before doing an expensive privilege-mode transition.

    If, instead, as on OpenBSD, the kernel enforced the rule that all system calls had to go through libc (or perhaps a big ntdll.dll-like VDSO), then the whole problem the linked article tries in vain to solve would disappear. If you wanted to hook a system call, you'd just change the libc/VDSO dispatch. No need to rewrite any instructions.

    If I were Linus, I'd make a new rule: starting today, all new system calls must go through VDSO. No exceptions. SYSCALL from anywhere else? SIGKILL.

    This way, you can just LD_PRELOAD in front of the VDSO and system call interception in userspace Just Works.

  • > If I were Linus, I'd make a new rule

    Or, you know, just propose your idea to him

  • The amount of times we ran LD_PRELOAD in prod was vanishingly small and limited to debug so the OpenBSD solution seems to be just waste of CPU cycles
  • > all system calls had to go through libc (or perhaps a big ntdll.dll-like

    Which makes containers crap on Windows and *BSD as they have to run the currect libc or equivalent. Thus you need to build a different container per OS version which sucks compared to Linux.

  • > This model has always been a terrible idea.

    I disagree. It's an amazing idea. It allows me to write freestanding programs without any C libraries. It allows compilers to have Linux system call builtins that directly generate the calling convention. I created an entire lisp interpreter with nothing but Linux system calls, completely freestanding.

    I've written a sort of manifesto around this:

    https://www.matheusmoreira.com/articles/linux-system-calls

    > If I were Linus

    Good thing you aren't.

    As a kernel, Linux is completely independent from its user space. The instruction set is the correct abstraction for the system call entry point. There should be no "required C libraries". User space should be free to reinvent everything in Rust if it wants.

    There are various kernel mechanisms for system call interception if that's what you want. Tools like strace work just fine on my lisp interpreter, so libc is clearly not needed.

    LD_PRELOAD is a GNU ld feature. The linker is the exact sort of user space component that's supposed to be completely replaceable. None of this is any of Linux's business.

    Use of the vDSO is not even mandatory. All system calls in the vDSO are also available via the kernel entry point. The vDSO is just an optimization for frequently called system calls like gettimeofday. Forcing all programs to use the vDSO would force them all to not only implement the ELF spec but also to implement a small ELF linker. This is a significant blow if you want to create minimal freestanding Linux programs.

  • Direct system calls are an amazing idea. The NtDll and bsd models are worse. The whole libc becomes a security boundary without the protection of kernel space. So much windows malware and process tampering happens because now you have a library (ntdll) fully in userspace that is given special privileges, which now becomes a huge attack surface. Then you have to deal with breakages between the built in libc versions and the kernel

    This syscall overhead isn't as much as you suppose it is; for workloads where the syscall overhead actually makes a difference there are robust low-syscall paths for io/latency sensitive operations with DPDK, io_uring, and futex being a few examples.

    And there are robust performant methods on linux for syscall interception/tracing, see seccomp unotify, bpf tracepoints, ftrace.

  • > This model has always been a terrible idea. It robs the system of an ability to intercept system calls in userspace before doing an expensive privilege-mode transition.

    This model has always been a trade-off. It has downsides, but it also has upsides, including an immense boost in flexibility; decoupling from any particular userspace is useful.

    > This way, you can just LD_PRELOAD in front of the VDSO and system call interception in userspace Just Works.

    Can you LD_PRELOAD in front of the vDSO? I was under the (possibly mistaken) impression that the kernel injects it directly.