Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

That sort of uptime always scares the sh!t out of me when I see it.

Reboot at least once a month folks, even for those single node critical systems. Better it doesn't boot on a Friday night than mid morning on a Tuesday.

Also, monitor and send alerts :))



I manage a small cluster (< 500 nodes), and do staggered reboots every 3-6 months, mostly for security/firmware updates.

It's amazing how many servers that were seemingly running "fine" for months don't boot back up. Memory failures, disks that disappear, random power issues, motherboard/controller failures. As high as 1%.


There are a lot of "strains" that happen (current inrush, mechanical shock to drives, heat shock to parts) that happen on a system, particularly after a full power off and power up cycle.


For anything that is externally accessible, I imagine you'd want frequent security updates.


It’s interesting, we want reliability, but too much reliability makes something invisible, and then failures become catastrophic because there is no expertise in the organization.


Death and birth are evolved. Micro-organisms don't do it. A few multicellular organisms don't either (e.g. there's a kind of immortal jellyfish, but what it does is periodically revert to an amorphous mass and then grow back into a jellyfish.)


True reliability requires Chaos Monkeys


The thing about reliability is that not every single thing needs to be ‘reliable’. The reliability comes from knowing the failure points. Some issues you may not be able to fix right away but simply knowing why something fails is more reliable than not knowing. And that’s the scariest problem I’ve seen. Not knowing failures.


Moreso than this: redundancy.


It’s the same with wildfires. The obsession and panic around small fires makes me cringe.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: