Showing posts with label Error. Show all posts
Showing posts with label Error. Show all posts

Wednesday, April 10, 2013

ERROR: Call fails for “HostDatastoreSystem.QueryVmfsDatastore- CreateOptions”

I've run into this several times now when re-deploying servers as iSCSI SAN storage systems for vSphere.  What happens is there's an old filesystem partition (or two) on the device\disk so ESXi refuses to configure it as a datastore.

To fix this problem you have to delete these partition(s) from the device\disk.

Word of warning:  make sure you delete the correct partition!  If you delete the wrong partitions, you may have to recover/re-install ESXi.  The correct partitions should not be difficult to find, but now if you screw something up you can't blame me - you've been warned!

Use the vSphere client - on the ESXi host go to the Configuration tab, Storage, Devices.  Take note of the device name your trying to configure as a new datastore.

SSH into the ESXi host.
Run the following command:
fdisk -l

This will list the partitions on that disk device.

Now you need to delete these partitions:
fdisk /dev/disks/[DEVICE_NAME]

When prompted, delete each partition.  Press "d" for delete, then "1" for partition 1.  Do this for all partitions on this device.

When finished, press "w" to write the changes to disk.

You should now be able to go back into the vSphere client and create a new datastore using this device!

Monday, June 25, 2012

ERROR: VMware ESXi with 3PAR SAN and Dead LUN 254

Problem

Last week I discovered  a couple of error messages in the vmkernel.log file that caused me some concern:

2012-06-22T18:18:45.806Z cpu17:939673)WARNING: vmw_psp_rr: psp_rrSelectPathToActivate:972:Could not select path for device "Unregistered Device".
2012-06-22T18:18:45.806Z cpu17:939673)WARNING: NMP: nmpPathClaimEnd:1195:Device, seen through path vmhba2:C0:T2:L254 is not registered (no active paths)


I did a "esxcfg-mpath -l" and found 4 dead paths - 2 to each fibre HBA.  I never provisioned a LUN with an ID of 254.  Maybe this was a "special use" LUN?  I doubted it because my HP EVA has one of these and the device is listed in vCenter.  However, there were no devices with LUN ID of 254 listed anywhere in vCenter, only 4 dead paths.

Solution

After a focused Google search, I found the answer.  The HP 3PAR guy that came out and did the installation had us use a host persona of "1 - Generic" when we should have used "6 - Generic-legacy".  Luckily these can be changed on the fly via the InForm Management Console.

After making the change in the IMC, I rescaned each of the hosts and the messages stop appearing in the vmkernel.log and the dead paths were no longer listed in the vSphere Client.

Here are the relevant sites I found per the Google search:

And, of course, I always recommend following the manufacturer's best practices:

Looks like the guide was recently updated and it does recommend using the persona of "6 - Generic legacy".

Another mystery solved.

Thursday, June 14, 2012

ERROR: The query service is not available or was restarted

I wasn't getting any results when going to the Hardware Status tab for all hosts in my recovery site.  When clicking "update", I'd get the error:
The query service is not available or was restarted. Please retry.

Of course retrying doesn't work.  Hardware status worked fine for all hosts in my protected site.  I thought maybe it was a problem with linked mode in vCenter so I logged on to the vCenter server in the recovery site, fired up the vSphere client, opened the Hardware Status tab and got the same results.

I then started another instance of the vSphere client and logged directly on to the ESXi host.  The hardware sensor data worked fine here (it's not in a separate hardware tab, but looks nearly the same).  Hmmm.... must be something with vCenter?

I Googled the error and found this VMware KB:

Well it's the exact same error message so this must be the fix, right?  Wrong!
First of all, step 14 is incomplete.  Please follow these steps to reset the vCenter Inventory database:

Secondly, this was not the only problem and probably didn't ultimately fix the issue.  I found this link in the same Google search:

The above forum posting had a link to the following web site with instructions on updating the ADAM instance vSphere uses for linked mode:

While this was for 4.1, the same settings apply to 5.0.  The only thing I would recommend is checking all of the common name (CN) properties to make sure the FQDNs are correct.  You do not need to change these to IP addresses!

I did have to reboot the vCenter server in the recovery site after making the changes.  Even then, it didn't seem to start working until the following morning so it may take some time for the changes to propagate.  Not certain about that but now the hardware status tab works for all servers in the recovery site.

ERROR: vSphere Replication shows Not Active

Another strange one.  Existing VM replications appear to be working based on "last sync completed" time stamps.  However, setting up a new replications result in a status of "Not Active".  Right-clicking on a VM and choosing "sychronize now" results in this error:
Call "HmsGroup.OnlineSync" for object "[some long GID]" on Server "[server name/IP]" failed.  An unknown error has occurred.
I Googled the error and found this VMware communities forum thread:
SRM5 using vSphere replication, status shows 'not active'

Read through it but note that you shouldn't have to reboot everything like one poster did.  I rebooted the VRMS server in the recovery site and replications started working for all VMs again.  YMMV.  This happened to me after having rebooted the vCenter server also at the recovery site.

ERROR: A general system error occurred

Recently, when trying to logon to SRM 5, I got the following error:
A general system error occurred: Internal error

Wow, that's real telling!
I tried restarting the SRM servers but no dice.  I then opened a ticket with VMware support.  I started a WebEx with the tech and after reviewing several SRM and vCenter logs, he really couldn't find the root cause of the problem.  However, he did say that they've only seen this generic error with vCenter, not SRM.

We rebooted the recovery side vCenter and viola, I was able to login again successfully.  It's the old saying - if all else fails, reboot!