Known Issues in Lightbits 3.18.4
ID | Description |
|---|---|
48646 | A false |
48546 | node-manager can fail to start after losing a data drive when journaling is enabled and the journal devices are selected from a device pool that overlaps with the data devices. This is a related but distinct trigger from the general "missing drive" restart failure: it can still occur even once that issue is fixed. No workaround is currently available. |
48452 | The installer's etcd storage-performance check is mostly void and passes for disks with IOPS much lower than 10000 required. |
48435 | On AMD servers, certain NVMe SSDs could report timeouts. It is recommended to disable PCIe relaxed ordering in the BIOS. |
48205 | When a large number of clone volumes are created from a small number of shared snapshots and a node is then restarted or replaced, some of those clone volumes can become stuck retrying volume creation, which can leave the node's protection groups unable to complete their rebuild. No workaround is currently available. |
47941 | A snapshot creation can fail and report an "error" state if the same source volume had a management operation (such as an attach, detach, or another snapshot) interrupted shortly beforehand; for example, by a cluster-manager restart. The source volume is unaffected (no data loss) and stays available. Retrying the same request does not help; the failed snapshot should be deleted manually, and a new one should be created instead. If the new snapshot also fails, contact Lightbits Support. |
47940 | During periods of frequent etcd leadership change (for example, due to network problems, slow disks, or node restarts), cluster-manager could repeatedly restart. While this is happening, healthy storage nodes can be briefly marked Inactive, protection groups that rely on them can become Degraded (reduced redundancy), and some client operations such as volume attach could fail and need to be retried. There is no data loss, and the cluster recovers automatically once etcd leadership stabilizes; no manual recovery is required. If the restarts persist, contact Lightbits Support. |
47792 | A volume can be reported as stuck in an "Updating" state indefinitely, even though the underlying operation actually completed successfully and the cluster is healthy. This can happen if the cluster-manager process restarts while a volume update (such as an attach or detach) is in progress; for example, during an upgrade, node failure, or other control-plane disruption. There is no data loss, and affected volumes continue to serve I/O. No customer-side workaround is available; contact Lightbits Support to recover an affected volume. |
47753 | Changing a node's Permanent Failure (PF) threshold can behave unexpectedly if the node is already Inactive; lowering the threshold to a value that has not yet elapsed does not take effect until the original threshold time passes, and raising the threshold on an already-inactive node can prevent it from reaching Permanent Failure at all. Workaround: Change a node's PF threshold before performing maintenance that will make it inactive. If you need to shorten the threshold after a node is already inactive, use a value that has already elapsed. |
47300 | If the cluster manager's connection to the cluster's internal coordination service is interrupted (for example, due to a leader change, network disruption, or a brief service restart), the cluster manager's view of cluster servers can stop updating entirely, with no automatic recovery. Workaround: Restart the cluster-manager service to restore normal operation. |
46209 | The bundled Grafana Alloy log-streaming agent ships with Grafana's anonymous usage statistics reporting enabled by default. The report is sent out on a fixed four-hour schedule to https://stats.grafana.org/. In air-gapped or firewalled deployments, these attempts fail and produce recurring error entries in the journal. In connected deployments, they could represent an unsanctioned outbound connection to a third party. The exact list of fields collected is documented at https://grafana.com/docs/alloy/latest/data-collection/. |
45729 | When SSD journaling is configured with multiple NVMe devices, cluster cleanup does not remove the RAID (mdadm) arrays created by Lightbits. Before re-deploying on the same servers, manually remove these arrays using mdadm. |
45712 | In rare cases, if the internal connection used to track volume or snapshot state changes is interrupted (for example, due to a leader change, network interruption, or a brief service restart), a volume or snapshot update can remain stuck showing as still in progress even though the underlying change has completed. Workaround: Restart node-manager on the affected server to resolve the stuck state. |
45202 |
|
45157 | During a node power-up after an abrupt shutdown, a node's reported power-up progress could reach 100%, drop back to 50-60%, and then climb to 100% again. This is because the last of three internal recovery phases is not included in the calculation. This affects the reported percentage only. Recovery completes correctly and the node returns to service only when it is actually done. |
© 2026 Lightbits Labs™