Showing posts with label storage. Show all posts
Showing posts with label storage. Show all posts

Sunday, March 12, 2017

New NAS Build Update 1

After working with Rockstor and weighing my options with HDDs, I've decided to go a slightly different direction...

NAS OS
I'm a big CentOS fan.  I also have been wanting to learn more about Docker and BTRFS.  This makes Rockstor with its "Rock-ons" and BTRFS filesystem the perfect choice.  Or does it?

BTRFS
BTRFS seems to be Linux's answer to ZFS.  However, it is not nearly as mature, feature complete or ready for primetime in my opinion.  I wanted to work with 4+ drives which implies a RAID5/6 type of setup.  I quickly found that BTRFS has some serious problems with its RAID5/6 implementation to the point where it corrupts data and therefore is not recommended for production.  I could buy a true hardware-based RAID card for around $100, but I'm trying to avoid spending additional coin on this build (see Hardware section).

XPEnology
I'm "upgrading" from a Synology NAS.  I love the simplicity and the 'it just works' factor it provides.  Finding a NAS OS that provides a similar experience has proven to be difficult.  Rockstor will not let you use any non-BTRFS filesystems on data drives.  Per the RAID issue above and the fact that I did not get very good support from their somewhat inactive forums rules out Rockstor for this build.  About the same time I started running into obstacles with Rockstor, I stumbled upon XPEnology.

While not officially supported by Synology, it is based on their OSS code and the forum appears fairly active.  I've been testing it in a VM while waiting for some additional hardware parts to arrive (see hardware section).  This is basically DSM on your hardware of choice.  The real test will be the installation on the TS140-based system.  This is in no way my final decision on the NAS OS as further testing is required.

Hardware
The more I worked with the TS140's case, the less I liked it.  Both internal drive trays require an L-shaped SATA connector since they're so close to the side of the case.  Lenovo only provides one.  That was the final straw.  Luckily, I had a storage server case, an ARK 4U-500-CA that I was about to list on Craigslist but decided to use for this build.  It fits nicely with my 'low cost, greater flexibility' theme.  The rackmount hardware is removable.  Just add some rubber feet on the bottom and presto-chango - it's a server tower case!  It can hold at least 10 drives, has great air-flow and best of all - its free (to me).

In the next update, I'll detail the fun of moving the TS140 guts into a new case and some additional decisions required based on the number of drives needed.


Wednesday, March 1, 2017

Time for a (NAS) Change

I have been running the DVR feature of Plex for several months now.  During the same time the message "this server is not powerful enough to convert video" kept popping up intermittently preventing some videos from playing.  Oddly enough, it doesn't happen all of the time on all shows, but enough to be more than annoying.  Another thing I couldn't understand is why it was trying to transcode video at all when I have this option turned off for local streaming.

I was finally able to determine that the tuners record TV shows using the older MPEG-2 video standard which is still used to broadcast HD television today.  On the other hand, Plex wants to stream video using the newer H.264 standard.  Therefore, Plex will convert recorded TV shows on-the-fly (transcode) when streaming to Plex clients.

My existing NAS is a Synology DS713+ running Plex Media Server along with several other applications.  Upon researching this problem, I found a list of compatible NAS devices along with performance notes here: Plex NAS Compatibility List.

What this list tells me is that I have to spend around $1500 to get a Synology box that does not have any problems streaming 1080P HD video (mostly due to the requirement a modern Intel Core I3 CPU or higher).  I love my Synology but not that much - its hard to justify that kind of coin when I can build a system twice as powerful for half the price (or better).

So now I've decided to build my own.  I used to build and repair systems over twenty years ago and its interesting how little has really changes with system components.  I have spare disks, a RAID controller (if needed) and so my requirements are relatively light:
Form-factor: Tower
CPU: Xeon proc (no more problems with transcoding)
Memory: At least 4GB ECC RAM but expandable
Disk: capable of at least 4 disks, hot-swappable nice-to-have]
Quiet (will live in my office)
Cheap (I'm poor... and cheap)

I narrowed the system choice down to the Dell PowerEdge T20 or the Lenovo ThinkServer TS140.  These units are nearly identical and I wanted to go with the Dell but the ThinkServer prides itself in being quiet (26 decibels) and has an extra PCI slot that pushed it over the 'edge (hidden pun).  I ordered the ThinkServer from eBay/Newegg here: Lenovo ThinkServer TS140 70A4003AUX Tower Server

A note on memory:  I'm going to try the 4GB of memory to start.  It comes as a single DIMM (as opposed to two or more DIMMs) which leaves three open slots.  So if more memory is needed I can add another 4GB DIMM later.  I had tried running a custom built storage server three or four years ago using a desktop board with non-ECC memory running FreeNAS.  I had read on the forums multiple warnings that only ECC memory should be used due to the risk of bit rot and eventual data corruption.  Well, it only took a couple weeks before VMs stored on that systems starting becoming corrupted.  Having been bit by this bug, I won't run a storage server without ECC memory and highly, highly, highly recommend you do the same.

A note on storage:  I'm going to reuse the two 4TB drives in my Synology for this system in addition to two more 4TB drives I'll purchase for this build.  The TS140 has an integrated hardware RAID controller.  So for the data volume, I'll install the two new drives and pull one of the Synology drives for a total of three.  Then transfer the data off the Synology to the new system.  Once I'm satisfied everything is in good working order, I'll pull the fourth drive and install into the new system, expand the volume, etc.  Then I'll have a very nice DS713+ for sale!

A note on the OS: I will miss the easy UI experience Synology DSM provides.  However, I have found a decent alternative: Rockstor.  I looked at several options: OpenFiler (dead) NexentaStor, FreeNAS, etc.  But Rockstor offers features that I really want: web-based interface, based on CentOS 7, Docker-based plugin system called Rock-ons, optimized to run from an USB drive and Apple Time Machine Support.  I also use the Synology for security cameras but found Rockstor also supports ZoneMinder video surveillance system.
It's not all rainbows and unicorns though.  I'm not very confident about BTRFS (pronounced 'butter FS').  Rockstor recommends against using RAID5 or 6 for production systems.  No worries for me since I'm using hardware-based RAID.  I've also read that newer versions are much more stable than early releases.  So while they tout BTRFS as a big feature, it may actually be their biggest drawback.  Time and testing will tell.  There must be a reason that many if not most NAS companies are now offering this filesystem as an option on their devices.

Also note that the TS140 has an internal USB port - perfect for a USB flash drive to host the OS.

Here's the parts list:
System: Lenovo ThinkServer TS140 70A4003AUX Tower Server = $320
4TB Enterprise SATA x 2 = $177
16GB USB Flash Drive: Transcend 16GB JetFlash 820 = $10
Operating System: Rockstor = Free
RAID Adapter (if needed): PERC H310 = Free
NIC Adapter = Intel Quad port = Free

Application Mapping:
Plex Media Server (Synology package) = PMS Rock-on (Rock-on)
Surveillance Station (Synology package) = ZoneMinder (Rock-on)
CrashPlan (Community Package) = CrashPlan (Rock-on or native)
Photo Station (Synology Package) = TBD
Download Station (Synology Station) = CouchPotato?  TBD

So for $100 less I'll have a systems that's much more powerful and expandable  than a Synology 916+ with twice the capacity.  I'll follow-up with some additional thoughts once I have run this build through its paces.

Thursday, May 16, 2013

Re-Purposing an Old EMC Celerra

One of the reasons I was hired by my current employer was to implement an EMC Celerra NS352 NAS/SAN that was purchased the year before.  That was six years ago.  Since then I have moved all workloads off of this storage and on to other SANs.  So what to do with this device?  We tried selling it to four different vendors that specialize in used hardware - none of them would take it, not even for free.  Ouch!  Instead of paying a recycling company to come pick it up, I decided we could re-use it for tier II or III data.  Note that the Celerra has a CLARiiON SAN back-end.  I was never impressed with the Celerra's performance (putting it nicely) but connecting the CLARiiON directly to the SAN fabric and shutting-down/disconnecting all of the NAS head components should be worth the effort.  It does have roughly 13TB of 2Gb fibre drives after-all.

A word about maintenance:  This NAS/SAN system EOL'ed in February of this year.  Now the data I'm hosting on here will not be very important - backups, temporary data, test VMs, etc, but if a disk or other component fails (and it will), it would be more than nice to get the item replaced in a timely manner.  Enter third-party maintenance.  The Celerras were already under third-party maintenance, I simply worked with the existing vendor to convert the asset from Celerra to CLARiiON.  This also had the unexpected benefit of further reducing the cost of the maintenance contract.

The following sections document the process and procedures I used to re-purpose the CLARiiON back-end of two Celerras, giving a little more life to what would have otherwise been sent to the scrap heap.

Shutdown and Cabling

  1. Power off Celerra NAS head (not CLARiiON disk shelves or Service Processor bay)
  2. Disconnect Ethernet and fiber cables from SP bay.

Re-IP The Service Processors

  1. Connect a laptop to the Ethernet port on SPA
  2. Change laptop's IP to 128.221.252.111/255.255.255.0
  3. Browse to 128.221.252.200 (SPA)
  4. Logon using nasadmin/nasadmin
  5. Drill down to SPA, right-click and choose properties
  6. Under the Network tab, change IP and SP Network Name.
  7. Unplug the Ethernet cable from SPA an connect to SPB.
Repeat steps above but Browse to 128.221.252.201 (SPB).
Connect both SPs to Ethernet switch.

Now the remaining steps can be performed remotely.

User Management

I recommend setting up an administrative user to coexist with nasadmin. There is also an existing "admin" user that I would just leave there. I create an account named "Administrator" with a unique password.
  1. Go to the Tools menu and select Security\User Management\Add.

Clean-up Existing LUN

There will likely be many LUNs that were configure for and consumed by the Celerra NAS part of the system.  I simply deleted all of these, event the DART/OS LUNs and RAID Groups.  I did not have any hosts defined/registered.

Enable Access Logix

  1. Right-click on the CLARiiON (serial number) and choose Properties
  2. In the Storage Access tab, in the Data Access pane, check "Access Control Enabled"

Disk Layout/Hot Spares

Next I generate a disk layout report to see what disks are assigned as hot spares. Go to the Reporting node and generate a Configuration\Available Storage report.  Check to see where the current hot spare disks are located. I typically like to have these at the end of the disk cabinet/bus enclosure. Move the spare(s) to the appropriate drive(s) as needed.

The last EMC recommendation I'm aware of is 1 hot spare per 30 disks.

SAN Fabric Configuration

Time to connect the CLARiiON to your fibre fabric. 
  • Each SP should have 2 ports. Connect one port to each fabric.
  • Both SPs should be connected to both fabrics.
  • Create an alias for each port/WWN in your SAN switch.
  • Create the zones in your SAN switch to allow the hosts to "see" the CLARiiON.
  • Save and enable the new configuration.

Host and LUN Creation, Masking

Now that the CLARiiON can see the hosts, its time to register the hosts with the CLARiiON:
  1. Right-click on the CLARiiON (the serial number) and choose Connectivity Status
  2. Note that you will need to know the WWNs of each FC port of your hosts. All of my hosts are ESXi servers, so the following instructions will be for these types of hosts.
  3. Highlight the Initiator Name and click Register
  4. The Initiator Type should be "CLARiiON Open"
  5. Enter the HBA information
  6. Enter the Host information
  7. When finished click OK
  8. Click the Refresh button and you should see the host name appear in the "Server Name" column
  9. Complete the remaining initiators.
After you have all of the hosts registered, its time to create the LUN(s) you want the hosts to consume:
  1. Create a RAID Group -
    1. Right-click on the RAID Groups node and choose 'Create RAID Group'
    2. RAID Group ID: leave the default value ("0" for the first one)
    3. Number of disks: this is the stuff that starts religous wars. To keep it simple I use groups of 5 using RAID5.
    4. RAID Type: RAID 5
    5. Disk Selection: you could leave the default of "Automatic" and Navisphere will choose for you. However, I've never liked it's choices, so set it to manual and choose the disks that make the most sense (for example, the first 5 disks in enclosure x).
  2. Create the LUN -
    1. Right-click on the RAID Group you just created and choose "Bind LUN"
    2. RAID Type - default should match the RAID Group setting
    3. RAID Group - should be the same ID as the one you selected
    4. Rebuild Priority - leave default
    5. Verify Priority - leave default
    6. Default Owner - Choose Auto
    7. LUN Size - Choose MAX from the drop-down list

Connecting It All Together

  1. Create the Storage Group -
    1. Right-click on Storage Groups and choose "Create Storage Group"
    2. I used the name of my cluster since this storage will be shared among all hosts in that cluster
  2. Add hosts to the storage group - 
    1. Right-click on the storage group name and choose "Connect Hosts"
    2. Select all of the hosts that should be a member of this group and move them to the right-hand pane
    3. Click OK
  3. Add LUNs to the Storage Group -
    1. Right-click on the storage group and choose "Select LUNs"
    2. Expand out the appropriate component and select the LUN(s)
    3. Click OK
That's it! Your hosts should now be connected to the CLARiiON storage and consuming LUNs per what youy have configured above.

Other Notes

  • I recommend documenting the config as your setting this up. This includes network names, IPs and SP WWNs which can be found at:
         Storage Domains\LocalDomain\[SerialNo]\Physical\SPs\SPA(B)\Ports
         (It's the second half of each WWN)
  • Don't forget to recan your ESXi hosts and check to make sure the paths are using RoundRobin, VMW_SATP_CX and IOPS=1
  • I would also generate a new Available Storage report after each RAID Group/LUN creation event. I save these to an Excel spreadsheet for future reference
  • Use a tool such as Solarwinds Storage Manager to monitor SAN usage and availability.

Monday, April 29, 2013

Cool Tool: Seagate Wireless Plus

Okay, the Seagate Wireless Plus is probably closer to a cool "toy", but it is cool nonetheless.  I was skeptical at first - why do I need yet another storage device for media when I have Spotify for music (or one of the countless online streaming music providers)?  Why not just use Amazon Video On Demand (or one of the countless streaming movie providers)?  Well, it stores movies, music and pictures.  It can stream up to four HD movies simultaneously.  Built-in 10hr battery, Wi-Fi broadcasting, apps for iDevices and Android devices, hmmm...

Then the rationalization, er... use case, hit me - since my kids have Android tablets, I could replace the bulky CD carrier and flaky behind-the-seat DVD players in the mini-van.  Sweet!  The kids keep pulling out the cables to the point where I'm re-splicing them once every couple of months.  With this thing I can turn it on, put it in the glove box and forget about it.  Nice!

Problem!  How do I get all of those movies ripped and copied on to the drive.  Well, its going to take two more (cool) tools to get this done.  First up, Slysoft's AnyDVD HD.  One year of updates is roughly $80 and 2 years is around $103 (I'd get 2 years).  It's worth every penny.  This handy little software sits between your DVD/Blu-ray drive and your software player such that the disc appears to be unencrypted.  We need the disc to appear this way for the ripping tool to work.

Now that we have unfettered access to our movie, how do we rip and compress it into one file?  Note that if we don't compress it we'll fill up the wireless drive much sooner than we might have otherwise.  Also remember we're streaming moving to 7" Android tablets, we don't need the HD detail Blu-ray (1080p) or even DVD (480p) gives us.  After searching the Interwebs for some time I discovered the answer: DVD Catalyst 4.  As of this writing its on sale for $10 but it's worth even the full price of $20.  This product is simply amazing - the number of supported devices is staggering.

This actually presented another problem - my kids have two different devices, an Acer A100 and a Samsung Galaxy Tab 2 (which has much better hardware specs).  Which device do I choose?  I tried a couple of different options but found the best compromise to be the Acer A101.  Using this profile, movies stream and playback fine on both devices.  They even work on my Samsung Galaxy S3 phone nicely.

With these two tools working together, ripping a disc to an MP4 file is an easy 2-3 click process.  After that, connect the drive via the USB port to the computer and copy the files to the "Video" folder.  That's it!

I'm about two-thirds through the discs and just now starting to break 100GB (out of 900GB+ free).  The kids have used it a dozen times on various trips w/o issue - they're happy so I'm happy!

I can highly recommend this product.  Note that I haven't discussed all of the features such as wireless Internet connectivity pass-thru.  I'm sure there are other "use cases" that I didn't even touch on.  If you need a portable drive that can be accessed wirelessly, this is it.





Wednesday, April 10, 2013

ERROR: Call fails for “HostDatastoreSystem.QueryVmfsDatastore- CreateOptions”

I've run into this several times now when re-deploying servers as iSCSI SAN storage systems for vSphere.  What happens is there's an old filesystem partition (or two) on the device\disk so ESXi refuses to configure it as a datastore.

To fix this problem you have to delete these partition(s) from the device\disk.

Word of warning:  make sure you delete the correct partition!  If you delete the wrong partitions, you may have to recover/re-install ESXi.  The correct partitions should not be difficult to find, but now if you screw something up you can't blame me - you've been warned!

Use the vSphere client - on the ESXi host go to the Configuration tab, Storage, Devices.  Take note of the device name your trying to configure as a new datastore.

SSH into the ESXi host.
Run the following command:
fdisk -l

This will list the partitions on that disk device.

Now you need to delete these partitions:
fdisk /dev/disks/[DEVICE_NAME]

When prompted, delete each partition.  Press "d" for delete, then "1" for partition 1.  Do this for all partitions on this device.

When finished, press "w" to write the changes to disk.

You should now be able to go back into the vSphere client and create a new datastore using this device!

Tuesday, March 5, 2013

Using the vSphere 5 CLI for Storage Configuration

When Adding a new host, SAN or LUN to your vSphere environment there are some CLI commands and configuration settings you should consider.  Since Windows 7 is my primary desktop OS, I use the VMware vSphere CLI for Windows.  You can download the latest version from VMware's downloads site.

Before using any of the commands below, please review your SAN manufacturers best practices documentation for vSphere environments.  Can't find any?  Then you bought the wrong SAN!  Seriously, most manufacturers provide something around documentation - if Google doesn't help try using Bing... better yet, try contacting your VAR.

View Devices and Their Settings
To view devices and some of their related settings, use the following command:
esxcli -s [HOSTNAME] -u root -p [PASSWORD] storage nmp device list

Set Storage Array Type and Path Selection Policy
One of the first things you'll want to do is change the DEFAULT Path Selection Policy (PSP) for whatever storage array types (SATP) you're using.  This way, when adding a new LUN/device, the type will be set to the proper type automatically.  Most modern arrays support ROUND ROBIN.  To change the array type, use the following command:
esxcli -s [ESXHOST] -u root -p [PASSWORD] storage nmp satp set --default-psp VMW_PSP_RR --satp VMW_SATP_ALUA (or VMW_SATP_CX, etc.)

Note that this will not change/update existing LUNs/devices.  I recommend using the vSphere client for this unless you have many that need to be updated.  In this case, use the following command:
FOR /F %G IN ('esxcli --server [HOSTNAME] --username root --password [PASSWORD] storage nmp device list ^| findstr naa.600') DO esxcli --server [HOSTNAME] --username root --password [PASSWORD] storage nmp device set --device %G --psp VMW_PSP_RR

Note 1: In the above command, I loop through and set the RR policy for all devices that begin with "naa.600" - set the NAA ID for your environment as needed.

Note 2: the Linux command uses grep but Windows has FINDSTR.  Took a bit for me to figure out but if nothing else, this is the big find in this article - your welcome!

Set the IOPs Value
Finally, most array manufacturers recommend setting IOPS to "1".  To change the IOPs parameter, use the following command (again, thank you FINDSTR!):
FOR /F %G IN ('esxcli --server [HOSTNAME] --username root --password [PASSWORD] storage nmp device list ^| findstr naa.6005') DO esxcli --server [HOSTNAME] --username root --password [PASSWORD] storage nmp psp roundrobin deviceconfig set -d %G --iops 1 --type iops

Note:  As in the previous command, I loop through and set the IOPs parameter for devices that start with "naa.600" - set this for your environment as needed.

Now the "devices" hosted on your array should be optimally configured for use by vSphere 5.  Don't forget to go through and check each host.

Other Helpful Commands

esxcli storage vmfs extent list
esxcli storage core device detached list
esxcli storage core device detached remove -d [NAA ID]
esxcli storage core adapter rescan [ -A vmhba# | --all ]
esxcli storage filesystem list
esxcli storage filesystem unmount [-u <UUID> | -l <label> | -p <path> ]



Monday, January 14, 2013

Extending A vSphere Replicated Virtual Disk

I recently had a VM that needed one of its virtual disks extended.  vSphere Replication needed to be disabled for this disk before vCenter would allow this operation.
When reconfiguring replication for this VM/disk, it detected the original/smaller disk:
Duplicate File Found. Do you want to use this file as an initial copy?
If you choose yes, it will try to use this file instead of re-sending the entire virtual disk.  This option didn't work. I think the files are just too different (size) and it doesn't know how to handle it.

If you choose no, you can configure a different datastore, but it won't let you use the same one.  This would leave the original replicated virtual disk out there unnecessarily taking up space.

I ended up manually deleting the original replicated VMDK and had VR resend the virtual disk again.  No big deal this time as it was just Disk 0/C: drive but this could be a real PitA if you need to extend a larger disk.

Lessons learned:

  1. Size your drives properly the first time.
  2. Consider creating a new disk instead of extending an already large virtual disk (30, 40, 50GB+).





Monday, June 25, 2012

ERROR: VMware ESXi with 3PAR SAN and Dead LUN 254

Problem

Last week I discovered  a couple of error messages in the vmkernel.log file that caused me some concern:

2012-06-22T18:18:45.806Z cpu17:939673)WARNING: vmw_psp_rr: psp_rrSelectPathToActivate:972:Could not select path for device "Unregistered Device".
2012-06-22T18:18:45.806Z cpu17:939673)WARNING: NMP: nmpPathClaimEnd:1195:Device, seen through path vmhba2:C0:T2:L254 is not registered (no active paths)


I did a "esxcfg-mpath -l" and found 4 dead paths - 2 to each fibre HBA.  I never provisioned a LUN with an ID of 254.  Maybe this was a "special use" LUN?  I doubted it because my HP EVA has one of these and the device is listed in vCenter.  However, there were no devices with LUN ID of 254 listed anywhere in vCenter, only 4 dead paths.

Solution

After a focused Google search, I found the answer.  The HP 3PAR guy that came out and did the installation had us use a host persona of "1 - Generic" when we should have used "6 - Generic-legacy".  Luckily these can be changed on the fly via the InForm Management Console.

After making the change in the IMC, I rescaned each of the hosts and the messages stop appearing in the vmkernel.log and the dead paths were no longer listed in the vSphere Client.

Here are the relevant sites I found per the Google search:

And, of course, I always recommend following the manufacturer's best practices:

Looks like the guide was recently updated and it does recommend using the persona of "6 - Generic legacy".

Another mystery solved.

Thursday, May 31, 2012

VMware ESXi and HP 3PAR Storage

We recently added a new HP 3PAR F400 SAN to our existing VMware cluster.  If you haven't read about this SAN and are interested in SAN storage, I highly recommend you take a look at HP's site for more information: 3PAR Storage.

One thing I've found this SAN lacking in is documentation.  It's easy to find, but bits of information are scattered across several documents.  The HP implementation consultant handed me a thumb drive with 3GBs of data before he left, half of which are PDFs (to give you an idea).

The point of this post is to bring all of these bits together in one place to include settings and best practices for VMware ESXi 5.0 along with the relevant references.  Here's what I've found so far - I will continue to add and update settings as I find them where relevant:
ESXi Advanced Settings:
DiskMaxIOSize   = 128 (Source)
QFullSampleSize = 32  (Source)
QFullThreshold  = 4

Default Path Selection Policy (PSP) = VM_PSP_RR (Source and Powershell Tip)
IOPS = 1

LUN Size = As large as ESXi can handle, currently 2TB (Source)
VMDK format = Eager Zeroed Thick

VMs - zero-out free space (Source):
   For Windows = SDELETE (Download)
   For Linux = DD  (section 3.2.1 in source doc linked above)

Thursday, March 15, 2012

vSphere 5 Upgrade: VMFS Datastores

Per Its Time vSphere 5 Upgrade, time to upgrade VMFS datastores.  Like the last couple of steps this one completed w/o issue.  Note that it is better to create VMFS datastores because you'll get a 1MB block size regardless of what size datastore you're creating, optimizing disk space.  Compare this to the 2-8MB block sizes required based on datastore size in ESXi 4.1 and earlier.

All of my datastores now report to be VMFS version 5.54.


ERROR: ExtPart - Cannot find C: Drive

ExtPart threw this error on a Windows Server 2003 R2 64bit system that just had logical disk 0 extended.
ExtPart does support 32 and 64bit Windows Server 2003 R1/R2 (needs to be extracted via 32bit Windows or a tool such as 7zip).

My guess was that the command interpreter was not "seeing" the newly extended local disk space.  Sure enough, a reboot fix 'er right up!

I highly recommend both tools be a part of any administrator's toolbag:
Dell's ExtPart Tool
7-zip

Friday, March 2, 2012

Removing a VMFS Datastore

To remove VMFS datastores in the past, I always made sure there were no VMs left on the datastore, right-clicked and chose "Delete".  Well, apparently thats the wrong way to do it!  I must have been lucky, as VMware claims this method could result in an APD (All Paths Down) state.  If you don't know what that is, let me tell you it's bad (I have experienced this but for a different reason).  Your host(s) will lose access to storage.

I stumbled upon this vSphere blog post that has the procedure to remove a datastore the right way: Best Practice: How to correctly remove a LUN from an ESX host

UPDATE: For ESXi 5.0, I found it better to follow the KB mentioned in that blog post:
Unpresenting a LUN in ESXi 5.x

For ESXi 5.0, here's how I do a slightly modified vesion of the procedure listed in the post above with more information on items such as HA:
  1. Make sure all VMs are evacuated from the datastore/LUN.
  2. Using Datastore Browser, make sure there aren't any left-over directories or files.  If there are, delete them (make sure they can safely deleted first, of course).  Exception: the HA directory - you'll remove that a different way later.  Also note that you won't be able to delete a file if it's in use.
  3. Make sure all vSphere features are disabled for the datastore (i.e. SIOC, Storage DRS)
    1. HA Datastore Heartbeating:  There may be some cases where HA has chosen the datastore you're trying to remove.  In this case, edit cluster settings and change Datastore Heartbeating to "select only from my preferred datastores", then select at least three datastores other than the one you're trying to remove.  I highly recommend changing this setting back to "Select any of the cluster datastores" after you're finished removing this one.
  4. Next, for each host, go to Configuration\Storage\Datastores View, right-lick on the datastore and choose "unmount":
    1. The "Unmount Datastore Wizard" appears. Make sure all hosts are selected.
    2. Click next.  Hopefully the next step will look like this:
    3. Click "Next" then "Finish".  After a minute the datastore will be grayed-out and italicized.
  5. Then go to Configuration\Storage\Devices View, right-lick on the datastore and choose "detach"
    1. You get a pop-up dialog window similar to the one show above.  Choose "OK.  The device will be grayed-out and italicized after a few seconds.
  6. Go in to your SAN and remove LUN masking/unpresent the LUNs to the ESXi hosts.
  7. Finally, right-click on your cluster and choose "Rescan for Datastores".  You may get a StorageConnectivityAlarm alert for every host in your cluster unless you disable this alert first.

I like to go into every host and check the storage adapter for the LUN just to make sure but I haven't found a problem with this procedure yet.