Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I found out that using mmap and just telling uring to read form there to beat anything else.


Depends a lot on the memory pressure. If you can be fairly certain the data is (or will be) resident in memory, mmap is basically unbeatable. If you can't (because the data is larger than RAM or there's other stuff competing for RAM), mmap can have gnarly system-wide performance implications[1].

[1] Mandatory mmap=poop-emoji link: https://db.cs.cmu.edu/mmap-cidr2022/


In my case: its a PROT_READ, MAP_PRIVAT map and i cannot get a SIGBUS since i tell the kernel to handle it all for me, thanks to liburing, instead i get a short send in that case.


Why would you use uring to read from an mmap? Couldn't you just memcpy?


To be more precise: i use that map to send assets out directly to clients from a zip file.

Its a new web server i am building and its the fastest way i could find out.

Just switching from epoll to liburing made the server ~45% faster too, its ridiculous. It can serve 10 gigabyte per second with a single thread, or around 10 million responses per second with h2 and 32 multiplexed requests.

I had to write a new http load generator for that since i couldn't find one which could generate enough load to saturate my server or be fast enough to withstand it.


10 gb per second is pretty slow for a disk. You should be seeing much higher than that.


??? The only configuration that will allow disk reads at 10 GB/s is if you're using PCIe 5.0. PCIe 4.0 or lower, and SATA will not drop out long before that.

He's very clearly hauling data straight from the page cache.


...I'm still confused why you wouldn't use memcpy?


How can one memcpy from a fd to a socket? I don’t touch the bytes in userspace, I just tell the kernel to send them out.


So why would you mmap?


to not let the bytes i forward enter userspace, i serve from the page cache for the asset path, directly from a zip file.

i open the file, mmap it, close it and tell io_uring_prep_send which bytes from the mapping to send, this also saves me from a possible SIGBUS cause the access to the mmap happen inside the kernel, and when a SIGBUS would happen in userspace the kernel just reports a shorter send in cqe->res


Maybe they meant using IORING_OP_MADVISE with MADV_WILLNEED to bring ranges into memory?


The Linux-specific MADV_POPULATE_READ is much more reliable than MADV_WILLNEED if you want to implement read-ahead for a memory-mapped file.

In my opinion, the POSIX advices specified for madvise are useless or even dangerous.

On Linux, for precise control of memory-mapped files one should use only these 4 Linux-specific advices: MADV_COLD, MADV_PAGEOUT, MADV_POPULATE_READ & MADV_POPULATE_WRITE.

These should be used within io_uring, so that they will be executed asynchronously.

These have a well-documented meaning and using them carefully should be sufficient to reach optimum performance with mmap.


Though at that point you can just use FADV_WILLNEED.


I'm just steelmanning here ¯\_(ツ)_/¯


Databases vendors usually find that read outperforms mmap. Mmap being fast is a myth.


Tried every way to get contents of a file as fast to a client socket as possible.

mmap and io_uring_prep_send were faster than everything else, no matter the size as long as you keep the map around for the lifetime of the process.

for one off sends when a file is smaller than 256kb then io_uring_prep_read + prep_send are faster than everything else.


Databases don't send files to client sockets as fast as possible. They do computation and lots of random access.


Note that grandparent actually didn't say they used mmap to actually get the memory, just mapping it.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: