Today mdadm send me a mail to warn that one of my hard drive ( /dev/hdd1 ) was ejected from my RAID-5 array. After some manipulations (no writes, just reads on the file system to get information) and reboots, I ended up with a file system in a strange state: the folder structure was totally messed up and lots of files disappeared.

Assuming that this situation was about an inconsistent file index, I decided to reset the superblocks of the remaining physical disks:

1$ mdadm --zero-superblock /dev/hdc1
2$ mdadm --zero-superblock /dev/hdb1

I don’t know why I decided to do so, but it was the stupidest idea of the week. After such a violent treatment, my array refused to start:

1$ mdadm --assemble /dev/md0 --auto --scan --update=summaries --verbose
2mdadm: looking for devices for /dev/md0
3mdadm: no RAID superblock on /dev/hdc1
4mdadm: /dev/hdc1 has wrong raid level.
5mdadm: no RAID superblock on /dev/hdb1
6mdadm: /dev/hdb1 has wrong raid level.
7mdadm: no devices found for /dev/md0

At this moment I was sure that all my data assets were lost. I was desperate. My only alternative was to ask Google. So I did.

I spend several minutes browsing the web without hope. I finally found someone in the same situation as mine (sorry, in french) on debian-user-french mailing list.

The solution was to recreate the RAID array. This sound counter-intuitive: if we recreate a raid array over an existing one, it will be erased! Right? Wrong! As it is said on debian-user-french  , mdadm is smart enough to “see” that HDD of the new array were elements of a previous one. Knowing that, mdadm will try to do its best (i.e. if parameters match the previous array configuration) and rebuild the new array upon the previous one in a non-destructive way, by keeping HDD content.

So, here is how I finally recovered my RAID array:

1$ mdadm --create /dev/md0 --verbose --level=5 --raid-devices=3 /dev/hdc1 missing /dev/hdb1
2mdadm: layout defaults to left-symmetric
3mdadm: chunk size defaults to 64K
4mdadm: size set to 312568576K
5mdadm: array /dev/md0 started.

Of course this doesn’t solve my initial problem about the /dev/md0 file system: it is still in an altered state. Maybe it’s too late to recover data. But at least I reverted all my today’s mistakes, and the situation will not deteriorate until I power up my RAID! :)

31 archived comments

Comments are closed. These were posted on Disqus, and are archived here.

  1. Thank god I wasn't the only one who did that. Thanks so much for putting this out there - I'd rather not think of how much time you saved with this!
    There's nothing quite as bad as having to try to restore original research data from a hundred different places! Thanks again!

    Tim

  2. Vasichkin

    Thanks!!!
    I fixed RAID0 after upgrading kernel..

  3. So has anyone gotten this to work on a raid 5 and seen the data afterwords? I have 5 drives that would love to back up an running again :( Only one of my superblocks is corrupted, (as might be the drive). Is there a way to get four drives out of a a five drive array up and running again when mdadm has forgotten they exist?

  4. Thanks very much.

    I didn't have your problems, but your tutorial helped me getting up a raid which superblock was not found.

  5. Thanks for sharing this tip!

    You may have saved me hours of recovering from backup!

  6. Thanks so much for this! You just save me 750 GB of data!

    It's hard to believe that this simple fact (that the create option does not overwrite data) has been left so poorly documented. In fact, as good as mdadm is, it is very poorly documented.

    It seems every Fedora upgrade I go through it's a battle to bring in my old existing RAID 1 arrays. With Core 8 it would have been impossible if not for this write up.

  7. i have a different situation if any one can help,
    i have this NAS it is the intel ss4000-e. when i changed the authentication from local to active directory, a message showed up saying that all data will be destroyed,i clicked ok by mistake. i have very important data on this array and need to get it back. can anyone help please?

  8. Thank you :)))

    With your french link I just rebuild my RAID1 that *lost* superblocks.

    This sound counter-intuitive: if we recreate a raid array over an existing one, it will be erased ! Right ? Wrong !

    I did

    mdadm -Cv /dev/md0 -l1 -n2 /dev/sda /dev/sdb
    

    and I have all my data now.

    Note:
    I do not understand why --zero-superblock on /dev/sda, but it worked.
    (It's not a productive server, just a fresh install, so I dared to do so)

    1. can this

      mdadm -Cv /dev/md0 -l0 -n2 /dev/sda /dev/sdb
      

      recover RAID level zero SUPER-BLOCK after new installation of my linux

      NOTE: I was use ubuntu and my new installation was Fedora9

      1. -l0 is correct for raid level 0.

        But no idea if

        I was use ubuntu and my new installation was Fedora9

        will work

        man mdadm
        
               -l, --level=
                      Set  raid  level.  When used with --create, options are: linear,
                      raid0, 0, stripe, raid1, 1, mirror, raid4, 4, raid5,  5,  raid6,
                      6,  raid10,  10, multipath, mp, faulty.  Obviously some of these
                      are synonymous.
        
                      When used with --build, only linear, stripe,  raid0,  0,  raid1,
                      multipath, mp, and faulty are valid.
        
                      Not yet supported with --grow.
        
  9. Thanks so much. Short and and just to the point, excellent. Helped me a lot :).

  10. Hello,

    I am hoping this thread is still somewhat alive.

    I have a four drive raid5 mdadm array. I recently tried to move the raid array to a new computer and somehow lost 3 out of the 4 superblocks in the process.

    I am wondering if someone could answer a question for me.

    If I run mdadm --create on this array, but the physical order of the drives is different than it was (ie. sdb may now be connected to sdc), will mdadm properly reconstruct this array?

  11. hi there,
    i use a 4 drive raid 5 wich now is almost fucked up. i just zeroed my superblocks becouse of trying to recover a bad array. now i created a array with same parameters as used in old.

    mdadm --create /dev/md0 --verbose --level=5 --raid-devices=4 --spare-devices=0 /dev/sda /dev/sdb /dev/sdc /dev/sdd
    

    but now i get in mdadm --detail /dev/md0

     Active Devices : 3
    Working Devices : 4
     Failed Devices : 0
      Spare Devices : 1
    

    i had no spare in previous array. is everything over now? i cannot mount /dev/md0 because

    mount: wrong fs type, bad option, bad superblock on /dev/md0,
           missing codepage or helper program, or other error
           In some cases useful info is found in syslog - try
           dmesg | tail  or so
    

    and fsck reports "unsupported features"?!?

    1. the problem i had some posts above was my controler. i build up a new pc with my existing devices and used a SATA 1 Controller with 4x1.5 TB. the sad thing was, that my controller accepted the devices and mdadm could rebuild my array, but all data was smashed with CRC error. the TOC was okay and i could see all fileinformation, but 3/4 files had CRC errors. shame on me cuz i used a SATA 1 controller and shame on the controller to accept my devices. i lost all data!

  12. Hi,

    This also works if you have a two disk RAID1, one disk goes faulty, the other drops out of the array and you re-add the one that drops out of the array as a spare thinking that it is the faulty one...

    Long story short - data is back :D

  13. Dude, awesome!
    Been looking for hours for help with this.

    Had a RAID1 array where one of the disk started failing, and the other one was marked as spare, therefore not allowing the md-device to be created. It seems to be working now. Great undocumented feature of mdadm!

    I recreated the array from just the working disk (with "missing") and then added another drive.

    Greets from Sweden

  14. This is an interesting find... the mdadm superblocks exist (at least partially) to allow you to reconstruct the array without knowing all of the exact details about which components live where. If you have a set of RAID5 components where you have all drives except one AND all the superblock info is present on those components, you can just

    mdadm --assemble /dev/md0
    

    and mdadm will have enough info to know where all the data lives. However, if you do NOT have the superblock info, mdadm has to be explicitly told where each component drive exists in the array.

    If you're recreating a RAID5 array, the key piece of information here is that if you do NOT have the superblocks, the order of the components in the NEW array has to be the same as the order of the components in the OLD array.

    So if you originally did

    mdadm --create /dev/md0
    

    but lost all the superblocks, as long as you can run the same create command with the same drives (or the "missing" keyword in place of one of them) then mdadm will create the same superblock/slot assignment and you're good to go.

    Only in this case can mdadm read the data that is on the components of the array. This implies that once you have a functioning RAID5 array, you should get the output of

    mdadm --examine <all components>
    

    and store that somewhere, because it will tell you which "slot" the various components go into.

  15. I'm trying currently to rebuild two arrays that somehow got messed up. In one case, the array initially just showed one drive missing. Since smartctl showed no actual drive problems, I added it back into the array. Once it finished resynching, mdadm immediately threw out a massive number of error messages. After that the array showed two drives missing.

    My other array was corrupted in the same incident that took out the one above. The difference is, it always showed two drives missing.

    I tried getting the drives to re-assemble but they won't, so I wrote a script to cycle through the permutations of drives, with and without one missing, to see if I could create a mountable volume.

    When that failed, I tried it again with a --zero-superblock before slotting in a drive. That also failed. Each time I try to mount, I get:

    mount: wrong fs type, bad option, bad superblock on /dev/md2,
           missing codepage or helper program, or other error
           In some cases useful info is found in syslog - try
           dmesg | tail  or so
    

    I'm guessing there is some corruption in the file system. However I don't want to try fixing it until I can get the drive order correct and there are a lot of choices. I don't want to even think about what a fsck might do if the drive order was wrong. :)

    Can anyone suggest a method for determining the correct drive order in a RAID 5 array other than attempting to mount a file system?

    1. Nevermind, it turns out that the issue was trying mount an ext3 file system as ext2. I'd always thought that you could do it, but apparently that is not true. When I reran my script with the mount going as ext3, I got a successful mount.

    2. Hello. I have the same problem described by @error and @Gary. However since I actually posted all the codes and etc in ubuntuforums, I was wondering if you were kind enough to have a look at this and tell me if there is anyway that I could mount a newly created md0 without needing to reformat it?

      Here is the link:
      http://ubuntuforums.org/showthread.php?p=10775382#post10775382

      Thank you so much,
      Mo

  16. Thanks for the info. It was very helpful :)

  17. Kev,

    thank you for your helpful remarks!

  18. Thanks man! This saved me when my raid decided to up and crash on me.

  19. this is my current super-block can some one help me with the order to create it in is from /dev/sd[bcdefghi]1 but if failed during an reshape from 6-8 drives

    /dev/sdf:

              Magic : a92b4efc
            Version : 0.91.00
               UUID : 01986f9c:c5d44da8:5df300a1:eb89baa4
      Creation Time : Thu Oct 22 18:25:01 2009
         Raid Level : raid5
      Used Dev Size : 976751872 (931.50 GiB 1000.19 GB)
         Array Size : 5860511232 (5589.02 GiB 6001.16 GB)
       Raid Devices : 7
      Total Devices : 7
    Preferred Minor : 0
    
      Reshape pos'n : 1664256 (1625.52 MiB 1704.20 MB)
      Delta Devices : 1 (6->7)
    
        Update Time : Sat Jan 30 19:30:41 2010
              State : active
     Active Devices : 7
    Working Devices : 7
     Failed Devices : 0
      Spare Devices : 0
           Checksum : 8accf9eb - correct
             Events : 219187
    
             Layout : left-symmetric
         Chunk Size : 128K
    
          Number   Major   Minor   RaidDevice State
    this     6       8      128        6      active sync   /dev/sdi
    
       0     0       8       33        0      active sync   /dev/sdc1
       1     1       8       49        1      active sync   /dev/sdd1
       2     2       8       65        2      active sync   /dev/sde1
       3     3       8       81        3      active sync
       4     4       8       97        4      active sync   /dev/sdg1
       5     5       8      113        5      active sync   /dev/sdh1
       6     6       8      128        6      active sync   /dev/sdi
    
  20. Anonymous

    thanks for saving 4.5TB...

  21. Hi there,

    Thank you for posting this information. We have had a distantly related problem that this post may have resolved.

    We've been trying to recover a s/ware raid-5 array from a QNAP 409P NAS where the superblocks of sufficient member drives to render the array inactive were corrupted during a power fault.

    We had tried -everything- to get the raid to assemble in a separate system with no success before coming across this page. Having decided that we were unlikely to recover the array, we thought we'd try deliberately zeroing the superblocks for all the member drives and then seeing if re-creating as you describe above would get us back online. The array is resyncing at the moment :D Hopefully the filesystem in it will be intact.

    Many thanks!

    - AJ

  22. have zeroed superblocks, rebuilt 4 2tb raid 5 twice now, replacing different 2tb drives in different orders (twice) both times once superblocks have been zero'd (which is interesting as original superblocks would be recognized briefly if each 'missing' superblock 2tb drive was hot unplugged; then plugged - superblock would show perfect 'clean' md info!) but they have been zero'ed twice now; first time out of order of original layout (by mistake) and second time zero'd and rebuilt in orderd as remembered. No avail.

    No recognized fs before array rebuild, or after. all drives have always been stated as 'clean' even before superblocks have been zero'd.

    Any ideas would be great. approx 5tb and 20years of data lost at moment. Thank you for great help so far!

  23. Thanks for posting your notes. I found my Fedora 16 x86_64 system unresponsive this morning and had to hard boot it.

    During bootup dracut reported that /dev/md1 couldn't be found.

    Boot into rescue and the auto discovery process for Linux partitions couldn't locate the RAID 5 with it's Logical Volumes, either.

    mdadm --create /dev/md1 --verbose --level=5 --raid-devices=3 /dev/sda1 /dev/sdb1 /dev/sdc1
    

    This found that each of the devices were "part of a raid array" I answered yes to the Continue creating array after which cat /proc/mdstat showed my raid5 array in recovery.

    Whoop.

    This has happened before on this system, usually over a weekend where the desktop environment sits idle for a couple of days. I suspect some kind of power saving is kicking in on the Fedora 16 desktop. Perhaps it's trying to sleep the hard disks and somethings getting mucked.

    I also noticed I had CStates enabled in the bios, disabling that for future testing.

    Thanks again for posting your notes.

  24. Thanks for this post, I've zeroed my superblock to shrink my RAID-1 array in 2 steps. 1 hard drive after the other.
    After a reboot in rescue mode, I couldn't assemble the array. But creating with --create made it perfect :) Now it is synchronizing :)

  25. you just saved my bacon.. I saved my data the same way :)

  26. Many, many, many thanks for your info!!
    I didn´t have a --zero-superblock option but it was useful to recover my Raid6 crashed in a power fault.
    I could recover all info :-)