The site is going down this afternoon for extended maintenance. Running some tests, changing some options and install some new hardware. All designed to try and help deal with some of the recent outages. Running the backup first, then I will get started.
Update: Screw it I am done for tonight. Got about half of what I wanted figured out. Luckily the last half lets me keep RW up most of the time I am working. I will have to come back and keep working on this probably tomorrow which means don't freak if there is intermittent downtime for a minute or so every now and then for the next day or so.
Monday, September 21, 2009
Repairing a table in the database
Things will be locked up for a few minutes while the repair is run. I am aware of the situation and working on it, and hope to have things back up shortly.
UPDATE: Repair is done site is back online let me know if there are further issues.
UPDATE: Repair is done site is back online let me know if there are further issues.
Sunday, September 13, 2009
Server crash post-mortem
Time for the official post-mortem of what happened as far as the server crash goes. The official cause of the crash shall be listed as a failing power supply unit.
About a month ago the power supply for the server went completely dead. In order to get the server back up and running as quickly as possible I swapped in a spare unit I had from an older computer. It did the job beautifully. Seeing as how everything appeared to be working fine and there were no substantial problems I didn't replace it with a new unit.
About 4 days before I left for my trip back home the server shut itself down. When I booted it back up I had some problems with the MYSQL server and so chalked the problem up to that. Then I left and went away. We all know what happened next. When I got back to the server I found that it was in the same disabled state as the first crash. I got it back up and decided to watch and see what it would do. A day later the same thing happened.
So I ran some tests on the power supply and it was providing irregular power on the 12v rail, my guess is that was probably leading to a temperature triggered shutdown. Anyway, regardless, 2 days ago I purchased a high-end power supply unit and swapped it in. Everything seems to be running fine now.
Thanks to the donations everyone at RW gave or have promised to give, I have gone ahead and upgraded some of the networking hardware that was worrying me as well. I am also working on getting some hardware to allow for remote management of the server even if it is unresponsive, as well as server resets automatically if it becomes unresponsive.
So the whole thing is my fault for not swapping in a new power unit after the old one failed and instead relying on a spare one. Feel free to block me for some pi unit of time for my failing.
As a final note, if this is truly an act of God as more than one person posited, it is pretty convoluted and weak. A swarm of locust munching on my power cords would have been far more effective in both maintaining downtime and for the general "shock and awe" of it all.
And a few RationalWiki prods for the road:
No true Scotsman
Common descent
About a month ago the power supply for the server went completely dead. In order to get the server back up and running as quickly as possible I swapped in a spare unit I had from an older computer. It did the job beautifully. Seeing as how everything appeared to be working fine and there were no substantial problems I didn't replace it with a new unit.
About 4 days before I left for my trip back home the server shut itself down. When I booted it back up I had some problems with the MYSQL server and so chalked the problem up to that. Then I left and went away. We all know what happened next. When I got back to the server I found that it was in the same disabled state as the first crash. I got it back up and decided to watch and see what it would do. A day later the same thing happened.
So I ran some tests on the power supply and it was providing irregular power on the 12v rail, my guess is that was probably leading to a temperature triggered shutdown. Anyway, regardless, 2 days ago I purchased a high-end power supply unit and swapped it in. Everything seems to be running fine now.
Thanks to the donations everyone at RW gave or have promised to give, I have gone ahead and upgraded some of the networking hardware that was worrying me as well. I am also working on getting some hardware to allow for remote management of the server even if it is unresponsive, as well as server resets automatically if it becomes unresponsive.
So the whole thing is my fault for not swapping in a new power unit after the old one failed and instead relying on a spare one. Feel free to block me for some pi unit of time for my failing.
As a final note, if this is truly an act of God as more than one person posited, it is pretty convoluted and weak. A swarm of locust munching on my power cords would have been far more effective in both maintaining downtime and for the general "shock and awe" of it all.
And a few RationalWiki prods for the road:
No true Scotsman
Common descent
Saturday, September 12, 2009
The expanding face of RationalWiki
Woot! New network switch just arrived.
I am replacing the weird little neon hunk of plastic that I think is a network switch...but the Mandarin confuses me......with a solid linksys switch I ordered from newegg. Thanks to everyone that helped donate!
I think the switch was by far the "weakest" point in the network setup, and most likely to fail next. So this is a good upgrade.
A switch should take less than a minute to install so I am not bothering with an RW intercom message. But posting this here just in case I blow something up and the site stays down longer than expected.
After this just need to get some automatic/remote server monitoring hardware and we will be set!
I think the switch was by far the "weakest" point in the network setup, and most likely to fail next. So this is a good upgrade.
A switch should take less than a minute to install so I am not bothering with an RW intercom message. But posting this here just in case I blow something up and the site stays down longer than expected.
After this just need to get some automatic/remote server monitoring hardware and we will be set!
Friday, September 11, 2009
New status widget and update on google
New widget
So for fun I have setup a little widget on the blog to show the status of the RationalWiki servers.
If the server is down it just says server down. If the server is up and working it will display the number of hours and minuets that the server has been up "straight." That means no reboot or power down. I am also displaying the 15 minute running average for the CPU load so people can see how busy the server has been recently.
Google update
So based on my searching we are back on top for Andrew Schlafly and Poe's Law which were two of our bigger hitters for search engine referred traffic. Not everything is reindexed yet but it looks like we should recover all right from our downtime.
So for fun I have setup a little widget on the blog to show the status of the RationalWiki servers.
If the server is down it just says server down. If the server is up and working it will display the number of hours and minuets that the server has been up "straight." That means no reboot or power down. I am also displaying the 15 minute running average for the CPU load so people can see how busy the server has been recently.
Google update
So based on my searching we are back on top for Andrew Schlafly and Poe's Law which were two of our bigger hitters for search engine referred traffic. Not everything is reindexed yet but it looks like we should recover all right from our downtime.
Hardware replacement for real this time
Okay, so now that our magic new backup system appears to be working, I can actually do what I meant to do yesterday. So the site is down because I am replacing hardware.
Obligatory Google prod of the day:
Denyse O'Leary in honor of the first person to openly admit to considering a defamation lawsuit against us.
Esther Hicks just because.
UPDATE: It looks like everything went exactly as planned, smooth upgrade, site back online. I will continue to monitor the situation to make sure nothing weird happens.
Obligatory Google prod of the day:
Denyse O'Leary in honor of the first person to openly admit to considering a defamation lawsuit against us.
Esther Hicks just because.
UPDATE: It looks like everything went exactly as planned, smooth upgrade, site back online. I will continue to monitor the situation to make sure nothing weird happens.
Subscribe to:
Posts (Atom)
