Wednesday, December 9, 2009

Colbert Nation

The server is under tremendous load tonight, most likely from our top spots for relevant google searches after Andrew Schlafly's appearance the Colbert show.

I will do what I can behind the scenes to maximize our ability to weather the storm, but it is likely to be rough for a day or so. I imagine the initial rush should be only a few hours.

Thursday, November 26, 2009

Techinical issues...

If you have tried to access RW in the last hour or so you may have noticed we are having some technical issues.

Working on it while I type, and the good news is that the server was able to recover from a crash to a state that remote access was available. This is important because it means our fault protection appears to be working allowing for remote recovery from issues.

Hoping to have everything back and working with in the hour.

Update: Site recovery went fine, but rerunning back up since it got jacked up, so looking at another 20 minutes or so till the site is fully accessible again.

UPDATE: Okay site should be back up and running as normal. Its well past my bedtime.....have to teach tomorrow morning. OH well. Leave a message on my talk page or on here if you have problems with the site or you can contact Nx as he should be able to handle anything. He is still here right?

Sunday, November 1, 2009

October stats

October was the most active month ever for RW. There were multiple events that helped spur us on to nearly 100,000 unique IP address visits. We also appear to have recovered from our slump due to the RW outage in August/September.



Below is the daily traffic record. We had three major events that helped propel our traffic. I will talk about those below.



Tuesday, October 6, 2009

Traffic spike

Due to the coverage of The Conservative Bible Project RW is seeing a predictable jump in traffic. Its still going hot and in order to keep serving up are brand of the good news I am holding off on the back up for tonight.

Don't break anything.

Saturday, October 3, 2009

An act of "nature"

Sorry folks, based on the blinking clock above my computer it appears the power went out to my house. All the various fault protection equipment kicked in as appropriate but I can't sustain the server for longer than about half an hour with no power. Based on site monitoring it appears that it was inaccessible for only a few minutes at most.

Friday, September 25, 2009

Fault protection take 2

I have received the additional hardware I needed to get the fault protection working...I think. I will be setting that up and running some test on it today. It shouldn't effect RationalWiki at all as I can do my testing further down the network chain. If all goes well I will need to restart the server and that's about it. I will drop an intercom on RW when/if I do that.

Update: Okay, fault protection seems to be working well. I have implemented it live on the server now. I will continue to monitor everything and adjust as needed. But all appears well for now.

Wednesday, September 23, 2009

So just how bad has it been?

The recent downtime issues annoy me at least (probably more) than every user of RationalWiki. As always, I strive to offer the best service as far as performance, availability, and protection against catastrophe that I can. But I am not perfect, and I think it is important to keep in mind that this is hobby, and my "real life" as a graduate student is as overworked and under paid as most cliches depict it.

All that said and done, just how bad has it really been? After the great crash in August I started working with a service to help monitor RationalWiki's up time and server performance. With close to three weeks worth of data we can paint a picture for how bad things have been.



If you take a look at this analysis you can see that RW has actually been up 95 percent of the time. That is not bad all things considered. Take into account a few points: first that a good chunk of our downtime is packed together so it is mostly caused by 1 or 2 disasters that caused prolonged downtime, and second that our nightly backups cause time out errors for about 20-30 minutes. If you remove the few major disasters our uptime averages just over 99 percent a day, and without the backups you are looking at 100 percent coverage most days.

The key then is disaster recovery. To be able to quickly handle issues that cause long protracted downtime. Most of these are easily handled if I am awake and with in walking distance of the server. The issues today with the cable going down are very rare. So we are left with one major issues: server cop-outs that prevent remote log in and shutdown the site that occur when I am either asleep or traveling.

I am actively working on a solution that I think will greatly increase the servers ability to auto-recover from failure, and to expand the options for remote administration in the event of catastrophic failure when I am not present (ala what happened in August).

A lot of this is trial-by-error and learning as I go. I have never done a project like this before. All we can do is learn from our mistakes, and move forward with the goal of doing the best we can. That said what we do have is pretty good I think.