Known Issues in Lightbits 3.18.4

AI Tools

ID

Description

48646

A false UnrecoverableDataIntegrityError event with no volume ID and "Rebuild read detected a damaged object" in the frontend log could be reported during rebuild when a volume with old deleted snapshots is deleted.

48546

node-manager can fail to start after losing a data drive when journaling is enabled and the journal devices are selected from a device pool that overlaps with the data devices. This is a related but distinct trigger from the general "missing drive" restart failure: it can still occur even once that issue is fixed. No workaround is currently available.

48452

The installer's etcd storage-performance check is mostly void and passes for disks with IOPS much lower than 10000 required.

48435

On AMD servers, certain NVMe SSDs could report timeouts. It is recommended to disable PCIe relaxed ordering in the BIOS.

48205

When a large number of clone volumes are created from a small number of shared snapshots and a node is then restarted or replaced, some of those clone volumes can become stuck retrying volume creation, which can leave the node's protection groups unable to complete their rebuild. No workaround is currently available.

47941

A snapshot creation can fail and report an "error" state if the same source volume had a management operation (such as an attach, detach, or another snapshot) interrupted shortly beforehand; for example, by a cluster-manager restart. The source volume is unaffected (no data loss) and stays available. Retrying the same request does not help; the failed snapshot should be deleted manually, and a new one should be created instead. If the new snapshot also fails, contact Lightbits Support.

47940

During periods of frequent etcd leadership change (for example, due to network problems, slow disks, or node restarts), cluster-manager could repeatedly restart. While this is happening, healthy storage nodes can be briefly marked Inactive, protection groups that rely on them can become Degraded (reduced redundancy), and some client operations such as volume attach could fail and need to be retried. There is no data loss, and the cluster recovers automatically once etcd leadership stabilizes; no manual recovery is required. If the restarts persist, contact Lightbits Support.

47792

A volume can be reported as stuck in an "Updating" state indefinitely, even though the underlying operation actually completed successfully and the cluster is healthy. This can happen if the cluster-manager process restarts while a volume update (such as an attach or detach) is in progress; for example, during an upgrade, node failure, or other control-plane disruption. There is no data loss, and affected volumes continue to serve I/O. No customer-side workaround is available; contact Lightbits Support to recover an affected volume.

47753

Changing a node's Permanent Failure (PF) threshold can behave unexpectedly if the node is already Inactive; lowering the threshold to a value that has not yet elapsed does not take effect until the original threshold time passes, and raising the threshold on an already-inactive node can prevent it from reaching Permanent Failure at all. Workaround: Change a node's PF threshold before performing maintenance that will make it inactive. If you need to shorten the threshold after a node is already inactive, use a value that has already elapsed.

47300

If the cluster manager's connection to the cluster's internal coordination service is interrupted (for example, due to a leader change, network disruption, or a brief service restart), the cluster manager's view of cluster servers can stop updating entirely, with no automatic recovery. Workaround: Restart the cluster-manager service to restore normal operation.

46209

The bundled Grafana Alloy log-streaming agent ships with Grafana's anonymous usage statistics reporting enabled by default. The report is sent out on a fixed four-hour schedule to https://stats.grafana.org/. In air-gapped or firewalled deployments, these attempts fail and produce recurring error entries in the journal. In connected deployments, they could represent an unsanctioned outbound connection to a third party. The exact list of fields collected is documented at https://grafana.com/docs/alloy/latest/data-collection/.

45729

When SSD journaling is configured with multiple NVMe devices, cluster cleanup does not remove the RAID (mdadm) arrays created by Lightbits. Before re-deploying on the same servers, manually remove these arrays using mdadm.

45712

In rare cases, if the internal connection used to track volume or snapshot state changes is interrupted (for example, due to a leader change, network interruption, or a brief service restart), a volume or snapshot update can remain stuck showing as still in progress even though the underlying change has completed. Workaround: Restart node-manager on the affected server to resolve the stuck state.

45202

lbcli drops part of a label key that contains a hyphen. When creating a volume, if a label key contains a hyphen (-), lbcli incorrectly drops everything before the hyphen. For example: lbcli create volume ... --labels="pve-vmid=100" is saved as vmid=100 instead of pve-vmid=100. Workaround: Avoid hyphens in label keys, or create the volume via the REST API, which is not affected.

45157

During a node power-up after an abrupt shutdown, a node's reported power-up progress could reach 100%, drop back to 50-60%, and then climb to 100% again. This is because the last of three internal recovery phases is not included in the calculation. This affects the reported percentage only. Recovery completes correctly and the node returns to service only when it is actually done.