This is the most critical bit. We run our website (totally disconnected from our webservices, naturally) on a 256 MB VPS and frequently handle top HN stories on our wordpress blog (which runs on a host with several of our other PHP / static subsites).
Key points:
• Make sure that wordpress supercaching is on. You can verify this by looking at the last line of the HTML that comes back which has a timestamp for when it was generated.
• Turn off KeepAlives.
• Set MaxClients to 8.
• Use monit to check to make sure it can connect and restart apache with a kill -9 if it can't. (This is optional, but helps if you have some random thing that ends up taking a very long time to execute and eats up connections.)
With that in place even a much smaller server can easily handle a HN top story without breaking a sweat.
LOL what? what's your mpm? did you enable this to keep the backend from blowing up from too many queries? surely you can handle more than 8 connections at a time. is this the proxy layer, and if so are you using web caching on top of wordpress caching?
if monit is restarting apache every time it can't connect (i hope you have a long timeout) you're denying service to a lot of people. connections are supposed to queue so they don't get dropped.
As mentioned, this is on a single VPS with 256 MB of RAM. Each Apache process needs about 25 MB of (non-shared) RAM, so actually 8 is pushing it. We're using mpm_prefork. There's no additional proxy nor cache.
My point, specifically, was how low you can go with a cheapo VPS. We hold up fine during an HN spike. (We're B2B and not a destination site, so our usual load is trivial.) Even during an HN spike you're getting tops of 2-3 visitors per second, which can be dished out reasonably well with 8 workers.
The monit thing kicks in after a 30 second timeout. With the configuration above, we don't get that because of load, but rather when something else has gone wrong (specifically there's a wordpress plugin that our internal status blog uses that sometimes hangs). But given the original poster's issue of apache getting so out of control that it took him several minutes to get a live SSH connection and a system load of 60, having monit kill things (and restart them) is a preferable stop-gap.
(Note: Our actual customer facing stuff is quite different; there we're using multiple servers behind an nginx proxy and using a combination of Rails, Sinatra and Java services. The basic web stuff is segregated off from those primarily for security reasons.)
Ah. It's kind of scary that private RSS of each process would be 25MB, but it's certainly possible. I assume you've disabled every module you don't need?
If you have some free time try deploying your LAMP stack with Buildroot and uClibc. The application size ends up being around an order of magnitude smaller, but i've never bothered checking if private RSS on clunky apps like PHP or Perl is minimized at all.
edit Also for prefork we used to have some scripts that would monitor processes to see if they 'went crazy' and wouldn't ever return, and reap those processes so Apache would fork a new one so we didn't have to restart the whole server. When you're under peak load and you restart a server and it sends all those clients to all your other servers which are already almost at their peak things get very nasty very quickly. I know you only have the one VPS here but for applications with many servers it can be handy.
If you use ngix + php-fpm: the nginx fastcgi_cache module http://wiki.nginx.org/HttpFcgiModule#fastcgi_cache is really neat and a lot faster than the full-page caching plugins. I also found that an object cache does improve performance a lot for pages that are generated from wordpress, saving tons of database queries.
Do you use something automated to ensure things are cached? I use a different page caching plugin (W3 Total Cache), but it also puts the comment in the HTML.
No. Once I've seen that it's working properly at one point, I've never seen a case where SuperCache stopped working later (unless I changed the config).
Key points:
• Make sure that wordpress supercaching is on. You can verify this by looking at the last line of the HTML that comes back which has a timestamp for when it was generated.
• Turn off KeepAlives.
• Set MaxClients to 8.
• Use monit to check to make sure it can connect and restart apache with a kill -9 if it can't. (This is optional, but helps if you have some random thing that ends up taking a very long time to execute and eats up connections.)
With that in place even a much smaller server can easily handle a HN top story without breaking a sweat.