Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

If I lose a link anywhere in the world then BGP reroutes within a second or so, and we get a notification in seconds or maybe minutes on slack depending on how busy slack is being.

If it's down for more than 15 minutes in country or 2 hours internationally it gets flagged for manual attention. Shorter outages are logged but only looked at monthly as part of the operational report process.

Now sure you can have two independent faults on the same day, but the reports are that they didn't know the backup line was down.



Meh, I knew people would just clutch to the monitoring ignoring everything else.

In your case you already have way more than monitoring. You have the infrastructure designed for resilience. You have that design implemented. You have the protocols to do if the things go north. You have automation to disregard minor events and to bring to the attention more serious things. You have way more than monitoring alone.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: