Tuesday, 14 July 2009

[INCIDENT 2009/001] July 14th 2009 - Unexpected server failure

Incident log for July 14th 2009
Attending: dwm, ncm
Status: Completed at 11:32hrs, July 14th 2009.
Summary:
  • The server kalimdor.tastycake.net stopped functioning correctly at or shortly after 0900.02hrs for reasons unknown. This was detected at 1045hrs, and normal service was restored at 1132hrs.

Transcript, times are in GMT+1:
  • [1132] Normal services restored.
  • [1130] All filesystems pass checks. Bring server up into normal multi-user mode.
  • [1113] Server booted into single-user mode using secondary kernel image. Checking all local filesystems for errors.
  • [1057] Reboot into single-user mode failed; initial ramdisk for primary kernel image found to be corrupted or truncated. Rebooting into backup kernel image.
  • [1045] Service failure discovered. Emergency reboot triggered after serial console found unresponsive.
  • [0902] Clients running on kalimdor.tastycake.net time-out from remote services.
  • [0900] kalimdor.tastycake.net stops logging to local system log.

Thursday, 28 February 2008

[AT RISK 2008/004] March 8th 2008 - Electrical works planned between 0001-0200hrs

Maintenance log for March 8th 2008
Attending: dwm
Status: Completed at 10:00hrs, March 8th 2008.
Summary:

  • The rack containing the Tastycake.net server kalimdor.tastycake.net will be briefly powered down so that the rack can be connected to a newly-installed power-distribution board.
  • As a result, no services will be accessible whilst the switchover is in progress. The at-risk period will last until 0200hrs, though the colo engineers hope to have normal services resumed by 0030hrs.
  • Works to be carried out:
    • Shut down kalimdor.tastycake.net. (Completed)
    • Wait whilst the co-location engineers switch the rack over to the new power-distribution feed. (Completed)
    • Boot kalimdor.tastycake.net. (Completed)
    • Verify services are running normally. (Completed)
Transcript, times are in GMT:

March 8th 2008
  • [0915] Services verified as functioning correctly. (There was a minor issue with the current experimental DNS service for the dwm.me.uk domain as a result of invalid zone configuration data, corrected. It turns out that you're not allowed a CNAME as well as an SOA for the root of a zone, but A and AAAA records are fine..)
  • [0015] Power restored, automated reboot in progress. All services runningas normal.
March 7th 2008
  • [2355] Automated scheduled shutdown executed.
  • [2230] Delayed-effect shutdown instruction executed; shutdown will occur at 2355hrs.
  • [2105] Reminder of pending works (tonight!) sent by email to all users.
February 28th 2008
  • [1420] Initial update to off-site status page.

Saturday, 16 February 2008

[INCIDENT 2008/003] February 16th 2008 - Filesystem corruption, suspect defective IDE channel.

Incident log for February 16th 2008
Attending: dwm, mark, ncm
Status: Completed at 2250hrs GMT.
Summary:

  • Tastycake.net server kalimdor.tastycake.net has suffered a filesystem corruption problem.
  • As a result, some disk / directory accesses are blocking indefinitely.
  • We suspect that this data corruption is occuring somewhere along the disk channel supporting /dev/hdg.
  • Works to be carried out:
    • Remove /dev/hdg from all RAID mirrors to prevent further filesystem corruption. (Complete)
    • Reboot machine into single-user mode. NOTE: No services will be available whilst in single-user mode. (Complete)
    • Run filesystem verification utilities on all disk filesystems. (Complete)
    • Restore any damaged files from backups as required. (Complete)
    • Reboot machine back into normal production operation. (Complete)
Transcript, times are in GMT:
  • [2250] Incident closed.
  • [2247] Summary: All of the recovered files were old transient copies of data that had been deleted deliberately, with the possible exception of some of ~anton's image files, which have been copied to his home directory for review.
  • [2230] Of the remaining files all owned by ~jeremy, all but one are old versions of existing mailboxes - probably an artifact of normal mailbox re-writing operation. (Checking unique message ids shows that the mail messages still exist in the live mailboxes.) The remaining file just contains the junk chars "|a:0:{}" and doesn't appear in my filesystem index comparison. Almost certainly junk, deleted.
  • [2221] Found that most the disconnected files owned by ~anton are temporary files generated by gallery; deleted. (Christ, ~anton, you've got over a gigabyte of temporary files in there going back years! Clear it out!) His remaining files appear to be old deleted .jpeg photos, but moved them to a RECOVERED_FILES directory in his $HOME to allow for inspection and recovery.
  • [2217] Filesystem checks on /dev/mapper/volume-recover complete, no errors. Re-mounting /vol/recover.
  • [2210] Generating home directory indexes of affected users on live system and offsite backup for comparison.
  • [2153] Picking through the disconnected files found in /home:
    • One mailspool index auto-generated by Dovecot; will be automatically regenerated: deleted.
    • 8.5MB junk mailbox owned by ~jeremy; expendable!
    • Remaining files owned by ~anton and ~jeremy, no other users affected.
  • [2152] Running xfs_repair check on /dev/mapper/volume-recover in the background.
  • [2151] Disconnected inode files in /var are all old Apache logfiles dating to July 2007, which is older than normal retention policy. Deleted.
  • [2146] Machine back in production. Checking contents of lost+found.
  • [2144] Reboot in progress.
  • [2138] All filesystem checks complete, bar /dev/mapper/volume-recover which can be done whilst online. Rebooting to normal production mode.
  • [2137] xfs_repair completed, no errors found.
  • [2135] Running full xfs_repair on /dev/mapper/volume-root.
  • [2134] Appoximately 50 disconnected inodes detected on volume-home, relocted to lost+found. These may be real files, or they may simply be historical artifacts.
  • [2132] Minor error (link count) detected on volume-var. Full repair run also detected some disconnected inodes; running full repair on volume-home for good measure.
  • [2128] Re-checking volume-home and volume-var with xfs_repair -n for good measure.
  • [2127] Second filesystem check of /dev/mapper/volume-root complete, no errors. We may have been fortunate and only had the kernel BUG trigger as a result of a read error and not an earlier write error as previously feared.
  • [2124] Filesystem check of /dev/mapper/volume-root complete, no errors. Checking result with xfs_repair -n (as opposed to xfs_check).
  • [2123] Filesystem check of /dev/mapper/volume-root running, at least minor errors expected.
  • [2121] Filesystem check of /dev/mapper/volume-home complete, no errors.
  • [2119] Filesystem check of /dev/mapper/volume-home running.
  • [2118] Filesystem check of /dev/mapper/volume-var complete, no errors.
  • [2117] Filesystem check of /dev/md6 (/boot) complete, no errors.
  • [2116] Machine rebooted into single-user mode. All services unavailable from this point.
  • [2057] Initial tastycake-status bulletin published.
  • [2025] walled all logged-in users to advise that emergency maintenance in progress.
  • [2033] /dev/hdg dropped from all RAID mirrors to avoid further disk corruption. The next step is to reboot the machine into single-user mode to conduct full filesystem checks and repairs.
  • [2033] Incident announcement sent to all admins via http://twitter.com/tastycake.
  • [2026] Determined that cause of fault is a faulty data channel to /dev/hdg resulting in incorrect data being written to disk. Begun dropping /dev/hdg from all RAID mirrors to avoid further corruption.
  • [2016] Filesystem corruption detected in /root/.wajig/kalimdor
  • [2014] Kernel BUG (internal error alert) spotted by inspection.
  • [2002] Host monitoring system generates another Critical warning.
  • [1007] Monitoring system downgrades previous critical warning to minor severity.
  • [1002] Critical warning generated by host monitoring system, indicating that a significantly higher than normal number of cron processes are running.
  • [0632] Minor warning generated by host monitoring system, indicating that a higher-than-normal number of cron processes are running concurrently.

Monday, 11 February 2008

[AT RISK 2008/002] February 11th 2008 - emergency kernel upgrade

Maintenance log for February 8th 2008
Attending: dwm
Status: Completed at 1405hrs GMT
Summary:
  • Tastycake.net server kalimdor.tastycake.net being rebooted (at least once) at approximately 1200noon GMT for an emergency kernel upgrade.
  • New kernel needed to patch local root escalation vulnerabilities (CVE-2008-0009, CVE-2008-0010).
  • No Tastycake.net services will be available whilst reboots are occurring.
  • Works to be carried out:
    • Build new linux kernel (2.6.24.2) to replace existing build (2.6.24). (Complete)
    • Install new kernel and set as default. (Complete)
    • Reboot machine to start using new kernel. (Complete)
Transcript:
  • [1405] All tests clear, at-risk period concluded.
  • [1400] Machine rebooted successfully into new kernel. Running final checks..
  • [1357] Machine rebooted.
  • [1353] Believed that I have corrected the booting problem (missing /dev/md0 entry in /etc/mdadm/mdadm.conf) and rebooting again. (Again, with 1-minute grace.)
  • [1347] Successfully rebooted using original kernel; will be fixing raid-auto start, then rebooting again.
  • [1334] Backup kernel not functioning; appears to not be auto-starting /dev/md0; will need to configure manually. This may take a few minutes..
  • [1331] Failed to boot using new kernel, power-cycled via power-switch interface.
  • [1328] Machine reboot.
  • [1326] Reboot triggered with 1-minute grace delay.
  • [1320] New kernel installed, ready to reboot. Warning sent via wall to all logged-in users.
  • [1313] Updated kernel package built, installed in Tastycake package repository.
  • [1107] Initial update of maintenance log.
  • [1010] Determined that 2.6.24.1 kernel that had been built overnight has been superceded by 2.6.24.2, building new kernel image.

Saturday, 9 February 2008

[AT RISK 2008/001] February 9th 2008 - scheduled maintenance

Maintenance log for February 8th 2008
Attending: dwm
Status: completed at 15:11hrs
Summary:
  • Tastycake.net server kalimdor.tastycake.net being taken offline at 1200noon GMT for maintenance.
  • No Tastycake.net services will be available whilst works are in progress.
  • Works to be carried out:
    • Install third 250GB hard-drive into RAID mirror. (Complete)
    • Install GRUB bootloader on third drive. (Complete)
    • Replace old 127GB hard-drive with new 250GB replacement. (Complete)
    • Install GRUB bootloader on replacement drive. (Complete)
    • Create new RAID mirror set on as-yet unallocated space. (Complete)
    • Expand LVM working set using new RAID mirror set. (Complete)
    • Upgrade local kernel to 2.6.24. (Complete)
    • Discontinue local NFS server, use read-only bind mount for /vol/recover instead.
      (New feature in 2.6.24.) (Cancelled)
Transcript, times are in GMT:
  • [1511] Final checks complete, at-risk period ends.
  • [1506] Performing final checks prior to announcing end of at-risk period.
  • [1505] Read-only bind mounts don't seem to be functioning, we perhaps need an updated mount-utils. This can be done safely at a later date.
  • [1456] Initialized /dev/md0 as new LVM PV and added PV to existing volume VG; total capacity: 232GB.
  • [1454] Added new second disk to main RAID mirror set /dev/md7.
  • [1451] Created new RAID mirror across previously-unused disk space.
    (NOTE: /dev/md0 is not the RAID mirror containing /boot, /dev/md6 is.)
  • [1440] Kernel installed, rebooting to verify correct operation and to reload DOS partition tables.
  • [1437] Installing updated kernel packages. (And SNMP security updates, whilst we're here.)
  • [1434] Partitioned unallocated space on all three disks. Leaving creation of new RAID mirror array until last, as it would only be interrupted by reboots anyway..
  • [1431] Partitioned drive 2 to match other disks. Added drive 2 partition to /boot RAID mirror volume. Installed bootloader on new drive.
  • [1423] Replacement drive 2 installed, booting.
  • [1415] Re-sync completed. Rebooting to replace drive 2.
  • [1402] Re-sync 90% complete. (Unfortunately, it seems to be slowing down to about 15MB/sec again.)
  • [1357] Installed spare GRUB bootloader on new disk.
  • [1350] Re-sync 80% complete.
  • [1334] Re-sync 66.6% complete.
  • [1318] Re-sync 50% complete. (Now peaking at ~22MB/sec; ETA at present rate: 44mins.)
  • [1301] Re-sync 33.3% complete. (Now peaking at ~19MB/sec; ETA at present rate: 67mins.)
  • [1250] Re-sync 25% complete. (Looks like its speeding up as it proceeds, probably due to disk geometry.)
  • [1230] Re-sync 10% complete.
  • [1219] RAID re-sync in progress; need to wait for it to complete before replacing disk 2. ETA @ ~15MB/sec: 115mins.
  • [1211] Adding new disk 3 partitions to RAID mirror sets.
  • [1209] Disk installed, server rebooted. Partitioned disk 3 to match existing layout.
  • [1201] Serial terminal up; sent reboot instruction with 2-minute grace.
  • [1159] Sent final warning via wall to save all state; disk installed in caddy and ready for reboot.
  • [1150] Readying disk three for hot-insertion. (Though, because we're running on IDE, this will require a reboot..)
  • [1147] Had to abort the transfer drive update; don't have access to the rear of the rack, and the front-side USB port is far too slow. Will just have to do today's work carefully..
  • [1105] Taking full filesystem image backup to spare transfer drive.
  • [1045] Initial update of offsite maintenance log.
  • [1035] Arrived at Telehouse Docklands.

Friday, 16 November 2007

[INCIDENT 2007/005] November 16th 2007 - Upstream DNS servers offline

Incident log for November 16th 2007
Attending: dwm, mark
Status: Completed
Summary:
  • HostEurope.com, who provide DNS hosting services for Tastycake.net, are offline and are not answering queries for Tastycake.net addresses.
  • This may result in difficulties accessing Tastycake.net services, and delay the delivery of email to Tastycake.net email addresses.
  • As the fault lies somewhere within HostEurope.com's facillities, and not with the Kalimdor server itself, there is little we can do to directly address this problem, and are waiting for HostEurope.com to implement a fix.
  • In the unlikely event that HostEurope.com do not restore service in a timely fashion, we are preparing contingency plans to move the Tastycake.net domain to another hosting provider.
Transcript, times are in GMT:
Sunday Nov 18th
  • [1200] Upstream DNS problems resolved.
Friday Nov 16th
  • [2045] Incident report posted to http://tastycake-status.blogspot.net/.
  • [2021] HostEurope Support facility identified as offline.
  • [2015] Both authoritive DNS servers (provided by HostEurope.com) identified as failed.
  • [2010] DNS resolution problems first reported for tastycake.net domain.

Wednesday, 19 September 2007

[INCIDENT 2007/004] September 19th 2007 - /var full

Incident log for September 19th 2007
Attending: dwm
Status: Completed at 1230hrs GMT+1
Transcript, times are in GMT+1:

  • [1230] Incident closed, all services appear to be running normally.
  • [1158] Upgraded clamav to latest-stable version to address database retrieval problem. Continuing to monitor status of system.
  • [1130] Increased size of /var filesystem online. Restarted failed services: sysklogd, exim, clamav.
  • [0726] /var filesystem on Kalimdor goes full resulting in service failures. Flash messages dispatched from service monitoring agent indicating that SMTP is unavailable.

Thursday, 2 August 2007

[INCIDENT 2007/003] August 2nd 2007 - Airconditioning failure

Incident log for August 2nd 2007
Attending: dwm, nick
Status: In progress
Transcript, times are in GMT+1:

  • [1126] Kalimdor rebooted, throttled CPU to 900MHz until aircon fixed.
  • [1005] Text to Jump support requesting urgent powercycle - worried about potential CPU busy-wait loop and resulting temperature spike.
  • [0955] Email to Jump support requesting powercycle.
  • [0950] Hard lock after attempted CPU throttling to avoid burnout.
  • [0940] Inspected temperatures on kalimdor - 80-85 degrees CPU. This with ondemand throttling!
  • [0156] Email from Jump informing of air conditioning failure in TFM8.

Friday, 27 July 2007

[INCIDENT 2007/002] July 27th 2007 - Filesystem failures

Incident log for July 27th 2007
Attending: dwm, nick, mark
Status: Completed at 22:24hrs, GMT+1
Transcript, times are in GMT+1:
  • [2224] Restored to full multi-user mode with all services. (MySQL proved a little problematic - the automatic table management tooling didn't fix absolutely everything up - but that should be fine now.) Hopefully that's the last we'll see of this particular class of problem, at least for a while - but we'll be keeping a close eye on things just to make sure.
  • [2201] Rebooting into full multi-user now.
  • [2144] New kernel is installed. Rebooting again into single-user to test.
  • [2136] Booted in single-user mode. Next step: download and install a new more-up-to-date kernel.
  • [2114] Okay, rebooting into single-user mode. We're going to move the UDMA-6-capable disk to /dev/hda, the first disk slot that's not attached to the Promise controller, in case it's the Promise IDE controller itself not liking UDMA-6 operation.
  • [2050] Quick SMART checks completed with no errors. Quick checks writing test files to temporary filesystems also passed.
  • [2047] Running SMART self-tests on both remaining IDE disks.
  • [2015] Hmm, re-checking the newly reconstructed XFS filesystems showed minor errors again on /, easily fixed - but it may be that we haven't fully eliminated the cause of any underlying IO problems.
  • [2012] All of the filesystems are back and intact. Next steps will be to reboot the machine into single-user and check that everything boots up that's supposed to. Also, I'm going to try some additional block-level tests to check that data being written to the drives is being recorded and read correctly. First, however, the curry Nick and I ordered has arrived at the Security desk..
  • [1956] rsync is showing some files which are absent without leave; these are most likely the files relegated to lost+found. Rather than manually re-sort these files, we're going to restore those that are affected from the backup image - if anything proves to be missing, we still have the file fragments in lost+found that we can dig through.
  • [1947] Okay, /home filesystem has been checked and had its errors fixed, with a few files getting relocated to lost+found. We're going to do some whole-filesystem comparisons between our backup image and the repaired /home to try to make sure that nothing vital is missing.
  • [1935] Root filesystem has been restored. /boot filesystem passes all checks. /var contained some errors, fixed. Now running xfs_repair on /home to see how well it can cope with the errors on that volume.
  • [1833] Okay, we've got the RAID mirror up on the two good disks and have the LVM volumes up. Mounted the root filesystem; it appears that the earlier xfs_repair run moved everything into lost+found, which is fairly amusing. Proceeding to create new root filesystem and restore the contents from the backup image on the transit disk.
  • [1820] Performed some further tests, and determined that the error is persistent on that specific disk (as opposed to the disk channel it was occupying.) This disk, one of the disks that dates back to Kalimdor's original installation, is now considered suspect and has been removed. We will be rebuilding on the two other disks only.
  • [1755] Manual re-assembly of the RAID mirror showed that disk 2 of 3 had an invalid and out-of-date RAID superblock. This raises the possibility that this disk, or its ribbon cable, is bad. We've removed it from the working set and proceeding on the other two disks.
  • [1732] Machine has passed Memtest86 up to and beyond pass #4; it's fine. Next step: reboot from a rescue disk and restore the root filesystem from our transit disk.
  • [1705] Console hooked up, now running extended hardware memory test.
  • [1650] dwm and nick now on-site and online in Telehouse. Next steps: re-establish a functional root filesystem and conduct a thorough memory hardware check.
  • [1448] Backup complete. Packing up ready to head into Telehouse.
  • [1410] More than 2/3 complete - currently about 170,000 files left to go.
  • [1343] Copy of backup image to transit drive now 50% complete - about 300,000 files left to go.
  • [1317] Returned with the USB-SATA adaptor. Transfer of backup image to the transit drive is approximately 1/3 complete. Nick is now also en-route from Southampton to Telehouse; ETA sometime after 1500hrs.
  • [1226] Copying of backup data to transit drive now underway (all 600,000+ files of it.) Departing now for the local Maplin to pick-up a USB-SATA adaptor, back soon..
  • [1208] Called ahead to a Maplin a couple miles away - they've got the USB-SATA adaptor that we need. Just finished wiring the transit SATA drive into the offsite-backup server; about to start data copy to transit drive.
  • [1204] Received authorization to enter Telehouse building.
  • [1144] Updated plan:
    1. Copy the offsite-backup image to removeable media - in practice, a spare 250GB SATA disk.
    2. Whilst that's running, I'll head out to procure a USB-SATA adaptor so that we can copy the data back onto Kalimdor.
    3. Request for authorization to enter Telehouse has been submitted, we'll hopefully get that by the time we're ready to head in. Nick is also heading up from Winchester to assist with the recovery.
    4. Once we're onsite, we should be able to restore Kalimdor to good working order - in the worst case, restoring absolutely everything from the offsite-backup image taken early this morning.
  • [1129] Okay, that unfortunately didn't work. xfs_repair nuked /sbin/init. Off-site recovery methods are now exhausted, physical local access will now be necessary.
  • [1125] Root filesystem repair complete. Rebooting again into single-user mode.
  • [1121] Okay, plan: attempt filesystem recovery of /, see if we can get the recovery tools properly functional. In parallel, also preparing for physical entry to Telehouse with mobile copy of offsite-backup image.
  • [1110] Sent request for physical Telehouse access to co-lo.
  • [1059] Planning next steps.
  • [1049] Read-only check of / (root) filesystem is showing fairly extensive corruption. Other filesystems may be similarly affected. It may be necessary to physically go to Telehouse to rebuild the host from offsite backup.
  • [1044] Read-only check of /var XFS filesystem failed to terminate. Rebooting again to return to ground state.
  • [1021] Checking /var filesystem.
  • [1018] Executed restart via serial-console.
  • [1013] First response to issue. SSH, Apache services malfunctioning.
  • [0939] Issue raised via text-message.
Comments:
  • The bad news is that this has been a several-hour-long outage, for which we deeply apologise. The good news is that recovery seemed to go well and we believe any data loss from this significant filesystem failure was very minimal.
  • We think the root cause of the recent problems was a faulty disk. Unfortunately, rather than failing and refusing to function, we believe that it was silently recording data incorrectly - causing problems when it was read from again in normal operation. We have removed this disk and will be replacing it with a fresh replacement.
  • You may find it amusing to learn that today is Sysadmin Appreciation Day. If only the machines themselves respected such hallowed events..

Thursday, 26 July 2007

[INCIDENT 2007/001] July 26th 2007 - Kernel error

Incident log for July 26th 2007
Attending: dwm, mark
Status: Completed at 16:45hrs, GMT +1
Transcript, Times are in GMT+1:
  • [1645] Back to normal operation.
  • [1643] Final checks complete, booting to full operation.
  • [1642] Reboot complete, executing final checks.
  • [1639] Filesystem check of / complete. Executing reboot to single-user.
  • [1635] Filesystem /home check complete. Proceeding to check / (root filesystem).
  • [1631] Filesystem check of /export/recover complete. Proceeding to check /home.
  • [1616] /var check complete. Now checking /export/recover.
  • [1614] kalimdor.tastycake.net rebooted via serial console via SysRq. Checking /var filesystem.
  • [1611] dwm logged in via remote root shell.
  • [1610] Alarm raised by mark; kernel OOPS reported in XFS filesystem code. SSH services unavailable.

Sunday, 22 July 2007

Kalimdor maintainance log - July 22nd 2007

Kalimdor maintenance log for July 22nd 2007.
Attending: dwm
Status: Completed at 23:30hrs, GMT+1.

Objectives:
  1. [ABORTED] Replace existing power-supply unit (PSU) with new more-efficient model (80PLUS-rated) provided by Jump Networks.
  2. [COMPLETE] Repair faulty inode on /home filesystem.
Transcript, times are in GMT+1:
  • [2330] Kalimdor.tastycake.net has been returned to full multi-user mode, and is running all services. This ends the at-risk period.
  • [2324] Spare disk re-added to RAID, rebuild in progress. Switching to multi-user mode.
  • [2320] Rebooted successfully. Satisfied that all is well. Rebooting again, this time to replace backup disk.
  • [2313] Minor housekeeping errors on / fixed. Now rebooting. (Still in single-user mode.)
  • [2308] Minor housekeeping errors on /var fixed.
  • [2304] Quota checks complete; now double-checking other filesystems.
  • [2302] xfs_copy of /home complete. xfs_check shows new filesystem is intact. Performing first mount; quotacheck running.
  • [2255] xfs_copy is now more than 80% complete.
  • [2249] xfs_copy is now more than 60% complete.
  • [2244] xfs_copy is now more than 40% complete.
  • [2239] xfs_copy is now more than 20% complete.
  • [2232] xfs_copy is now running, copying the contents of the previously-created backup volume to /home.
  • [2224] Okay, xfs_repair just isn't working, and my window for getting home tonight is closing. Going to reconstruct and repopulat e /home from scratch.
  • [2220] Despite xfs_repair fixing some specific issues, mounting /home and checking shows that the errors have not been corrected. This raises a new hypothesis: the RAID mirror isn't fully synchronized, or isn't syncing data correctly. Investigating.
    [2206] Hmm, given how the building alarm keeps coming and going (and started at 2200), it's probably a test. Carrying on..
  • [2201] And that's the building fire alarm.
  • [2159] Rebooted with one disk removed. xfs_repair has now run once successfully over /home with minor changes (removals to lost+found) - rerunning again to see if the FS has now settled to a good state.
  • [2145] xfs_repair reported and corrected some errors; however, re-running xfs_repair reported even more errors - I suspect that the /home filesystem is either suffering from a serious problem, or the underlying LVM is malfunctioning badly -- most likely the former. However, to be sure, I'm going to pull one of the RAID mirror disks and keep it in reserve. In the worst case, I will be able to repopulate any broken filesystems from the spare disk.
  • [2142] Comical error message of the day: bad (negative) size -2500720168097138090 on inode 580671.
    Fsck continues..
  • [2133] Backup complete. Double-checking integrity of backup FS, then will re-run fsck on /home.
  • [2002] /home filesystem backup running. It's only completed a couple of GB so far, so it'll take a good few minutes to complete. Taking advantage of the delay to go and fetch some food before I pass out!
  • [1953] The filesystem check has turned up the expected single-inode error; however, xfs_repair is unable to fully repair the filesystem. Now making a seperate copy of /home before continuing, just to be on the safe side.
  • [1936] Old power supply has been replaced and Kalimdor has been re-installed in the rack. Now rebooting to single-user mode to perform the planned filesystem checks.
  • [1908] Aha: it turns out this particular sub-variant of PSU doesn't include a particular -5v line necessary for correct operation. (We've got a ATX12V PSU, and the new one we have is an ATX12V v2.2 PSU. Frustratingly, they're not backwards compatible. ) Now going through the delicate process of removing the new PSU and threading the old one back in.
  • [1851] The reinstalled machine is failing to power-up with the new PSU, though it's able to drive its networking status lights, none of the fans are running and it fails to respond to the power-switch. Working to identify the fault now, though if we can't fix this very quickly we'll have to fall back to our older (working) PSU.
  • [1827] New power supply installed, machine re-assembled. Getting a power cable to the optical drive was indeed very fiddly, but achieved now. Unfortunately, the new PSU doesn't have a seperate IEC break-out socket for mounting on the rear of the case, and there's nowhere to physically attach the new PSU inside the rack itself. About to reinstall in the rack now.
  • [1757] Swapping out the PSU. Cable-running and re-mounting on the inside of the case is a little fiddly, so will take a few more minutes.
  • [1705] Obtained access to TFM-8 server room containing Kalimdor. Proceeding to execute a clean shutdown.

Thursday, 19 July 2007

[AT RISK] Kalimdor downtime THIS SUNDAY from 1700hrs

Hello all,

Kalimdor.tastycake.net will be taken out of service for on

Sunday 22nd July (THIS SUNDAY) from 1700hrs

... in order to carry out preventative maintenance. Whilst these works are in progress, NONE of the tastycake.net services will be available. I anticipate that the works will only take approximately 30 minutes to complete.

Status updates will be posted to our new offsite status page, http://tastycake-status.blogspot.com.

The works to be carried out are as follows:
  • Replacement of Kalimdor's internal power-supply unit (PSU) with a new, more-efficient equivalent that conforms to the 80PLUS specification.
  • A filesystem check on /home in order to clear a stuck inode.
I apologise for the short notice; if these planned works present a problem, or if you have any other comments, concerns or queries, please contact us the administrators via all-heroes([a])tastycake.net.

Cheers,
David