Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Why does the Linux kernel let the disk sit idle before the fsync? Why does fsync cause such a storm, if the kernel could write as it goes? Also, why does PostgreSQL do buffered writes instead of unbuffered? Seems like two wrongs make a wrong here...


Write-as-you-go isn't necessarily an optimal strategy for general purpose workloads with intermittent IO - batching writes lets you do more linear writes, do a better job of allocating blocks, hit the journal less often...

There is a need for an alternate API, but it's not like fsync is "broken" for the general case.


fsync has been broken since it was introduced in BSD4.x. It's never worked well, usually flushing the whole machine cache in one huge batch of i/o. I'm told current versions on Linux will restrict it to the one open file, but it's still a shotgun instead of a rifle.


The disk is almost certainly not idle, but instead the iops are consumed by read traffic.

Ideally, the kernel would know the deadline by which the data should be committed, manage accordingly, and the fsync would be simply acknowledged rather than inspiring extra IO.

This is pretty hard to work towards and the clearly specified narrow syscall interface leads to a great deal of difficulty in sharing this information with the kernel IO scheduling -- the way it is shared would naturally depend on the details of the kernel implementation, which would imply a tighter link between application and kernel version than is normally desirable.


The things that want sync don't usually want a promise for future time, they want positive ack the data is on the storage. So a deadline approach is pointless.

Async w/O_DIRECT is the correct approach, despite the complexity. The problem is handling table scans that exploit system buffer cache for read-ahead. The answer to that is to issue more i/os to bigger buffers, or to a set of buffers. Scatter-gather disk i/o would help with that.

It might be particularly sweet if well-behaved applications could rely on mixing O_DIRECT with bufferd i/o on the same files. That is, all writes done O_DIRECT on pages read with O_DIRECT, and reads might be done against page-cached buffered files.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: