<html>

  <head>

    <meta content="text/html; charset=ISO-8859-1"

      http-equiv="Content-Type">

  </head>

  <body text="#000000" bgcolor="#FFFFFF">

    Hi Andrew,<br>

    <br>

    this is something that I saw in my logs too, first on one node and

    then on the other three. When that happend on all four of them,

    engine was corrupted beyond repair.<br>

    <br>

    First of all, I think that message is saying that sanlock can't get

    a lock on the shared storage that you defined for the hostedengine

    during installation. I got this error when I've tried to manually

    migrate the hosted engine. There is an unresolved bug there and I

    think it's related to this one:<br>

    <br>

    [<a href="https://bugzilla.redhat.com/show_bug.cgi?id=1093366"><b>Bug&nbsp;1093366</b></a>

    -<span id="summary_alias_container"> <span

        id="short_desc_nonedit_display">Migration of hosted-engine vm

        put target host score to zero</span>]</span><br>

    <a class="moz-txt-link-freetext" href="https://bugzilla.redhat.com/show_bug.cgi?id=1093366">https://bugzilla.redhat.com/show_bug.cgi?id=1093366</a><br>

    <br>

    This is a blocker bug (or should be) for the selfhostedengine and,

    from my own experience with it, shouldn't be used in the production

    enviroment (not untill it's fixed). <br>

    <br>

    Nothing that I've done couldn't fix the fact that the score for the

    target node was Zero, tried to reinstall the node, reboot the node,

    restarted several services, tailed a tons of logs etc but to no

    avail. When only one node was left (that was actually running the

    hosted engine), I brought the engine's vm down gracefully

    (hosted-engine --vm-shutdown I belive) and after that, when I've

    tried to start the vm - it wouldn't load. Running VNC showed that

    the filesystem inside the vm was corrupted and when I ran fsck and

    finally started up - it was too badly damaged. I succeded to start

    the engine itself (after repairing postgresql service that wouldn't

    want to start) but the database was damaged enough and acted pretty

    weird (showed that storage domains were down but the vm's were

    running fine etc). Lucky me, I had already exported all of the VM's

    on the first sign of trouble and then installed ovirt-engine on the

    dedicated server and attached the export domain.<br>

    <br>

    So while really a usefull feature, and it's working (for the most

    part ie, automatic migration works), manually migrating VM with the

    hosted-engine will lead to troubles.<br>

    <br>

    I hope that my experience with it, will be of use to you. It

    happened to me two weeks ago, ovirt-engine was current (3.4.1) and

    there was no fix available.<br>

    <br>

    Regards,<br>

    <br>

    Ivan<br>

    <div class="moz-cite-prefix">On 06/06/2014 05:12 AM, Andrew Lau

      wrote:<br>

    </div>

    <blockquote

cite="mid:CAD7dF9d9WB7-mK43+gpfkDVYbykvDZUM3+-GqqAF5Lh9B7=1MA@mail.gmail.com"

      type="cite">

      <pre wrap="">Hi,

I'm seeing this weird message in my engine log

2014-06-06 03:06:09,380 INFO

[org.ovirt.engine.core.vdsbroker.VdsUpdateRunTimeInfo]

(DefaultQuartzScheduler_Worker-79) RefreshVmList vm id

85d4cfb9-f063-4c7c-a9f8-2b74f5f7afa5 status = WaitForLaunch on vds

ov-hv2-2a-08-23 ignoring it in the refresh until migration is done

2014-06-06 03:06:12,494 INFO

[org.ovirt.engine.core.vdsbroker.vdsbroker.DestroyVDSCommand]

(DefaultQuartzScheduler_Worker-89) START, DestroyVDSCommand(HostName =

ov-hv2-2a-08-23, HostId = c04c62be-5d34-4e73-bd26-26f805b2dc60,

vmId=85d4cfb9-f063-4c7c-a9f8-2b74f5f7afa5, force=false,

secondsToWait=0, gracefully=false), log id: 62a9d4c1

2014-06-06 03:06:12,561 INFO

[org.ovirt.engine.core.vdsbroker.vdsbroker.DestroyVDSCommand]

(DefaultQuartzScheduler_Worker-89) FINISH, DestroyVDSCommand, log id:

62a9d4c1

2014-06-06 03:06:12,652 INFO

[org.ovirt.engine.core.dal.dbbroker.auditloghandling.AuditLogDirector]

(DefaultQuartzScheduler_Worker-89) Correlation ID: null, Call Stack:

null, Custom Event ID: -1, Message: VM HostedEngine is down. Exit

message: internal error Failed to acquire lock: error -243.

It also appears to occur on the other hosts in the cluster, except the

host which is running the hosted-engine. So right now 3 servers, it

shows up twice in the engine UI.

The engine VM continues to run peacefully, without any issues on the

host which doesn't have that error.

Any ideas?

_______________________________________________

Users mailing list

<a class="moz-txt-link-abbreviated" href="mailto:Users@ovirt.org">Users@ovirt.org</a>

<a class="moz-txt-link-freetext" href="http://lists.ovirt.org/mailman/listinfo/users">http://lists.ovirt.org/mailman/listinfo/users</a>

</pre>

    </blockquote>

    <br>

  </body>

</html>