[Bug 2146964] Re: [SRU] Prevent masakari HA from creating duplicated notifications

Vladimir Petko 2146964 at bugs.launchpad.net
Fri Jul 3 00:08:31 UTC 2026


Hi,

would it be possible to make small changes to the patches before the upload:
- add Bug: https://review.opendev.org/c/openstack/masakari/+/978343 and Bug-Ubuntu: https://bugs.launchpad.net/masakari/+bug/2146964 DEP-3 headers
- git format-patch upstream commit, there is a minor difference in the provided patch vs. upstream commit:
 --- a/masakari/conf/api.py
 +++ b/masakari/conf/api.py
-@@ -94,6 +94,33 @@ unchanged.
+@@ -98,6 +98,33 @@
  
-   ``masakari-api``
- 
-+* Related options:
-+
-+  None
-+"""),
+   None
+ """),

Please ping me on matrix/MM once the package is ready for upload.

-- 
You received this bug notification because you are a member of Ubuntu
OpenStack, which is subscribed to Ubuntu Cloud Archive.
https://bugs.launchpad.net/bugs/2146964

Title:
  [SRU] Prevent masakari HA from creating duplicated notifications

Status in Ubuntu Cloud Archive:
  In Progress
Status in Ubuntu Cloud Archive caracal series:
  In Progress
Status in Ubuntu Cloud Archive dalmatian series:
  New
Status in Ubuntu Cloud Archive epoxy series:
  New
Status in Ubuntu Cloud Archive flamingo series:
  In Progress
Status in Ubuntu Cloud Archive gazpacho series:
  In Progress
Status in Ubuntu Cloud Archive ussuri series:
  Won't Fix
Status in Ubuntu Cloud Archive yoga series:
  In Progress
Status in masakari:
  Fix Committed
Status in masakari package in Ubuntu:
  In Progress
Status in masakari source package in Focal:
  Won't Fix
Status in masakari source package in Jammy:
  In Progress
Status in masakari source package in Noble:
  In Progress
Status in masakari source package in Questing:
  In Progress
Status in masakari source package in Resolute:
  In Progress
Status in masakari source package in Stonking:
  In Progress

Bug description:
  [ Impact ]

   * Older versions of Masakari are lacking a proper mechanism to prevent
     concurrent record inserts into the database, resulting duplicated
     event processing that can potentially lead to disastrous outcomes.

   * Affected versios include:

     - Focal/Ussuri 9.0.0-0ubuntu0.20.04.5 / 9.0.0-0ubuntu0.20.04.5~cloud0 (UCA)
     - Jammy/Yoga 13.0.0-0ubuntu1 / 13.0.0-0ubuntu1~cloud0 (UCA)
     - Jammy/Caracal 17.0.0-0ubuntu1 /	17.0.0-0ubuntu1~cloud0 (UCA)
     - Noble 17.0.0-0ubuntu1
     - Resolute 21.0.0-0ubuntu1
     - Stonking 21.0.0-0ubuntu1

   * This is a known issue and reported by many users also presented in
     LP#2028450.

   * This requires an HA Masakari deployment without a coordinator
     configured.

   * None of the current Charmed Masakari deployments support using a
     coordinator and are all affected.

  [ Test Plan ]

   * Testing requires an OpenStack environment deployed with HA Masakari.
     Masakari deployed in HA mode is a hard requirement as the issue only
     affects HA deployments due to workers competing with each other.

   * Juju bundles will be crafted and attached to this SRU bug to make
     reproducer deployments easier. The plan is to cover all affected
     versions. Detailed instructions will be provided.

   * The LP#2028450 bug includes a reproducer script (2), which is handy
     to reproduce the issue as well as validate the fix.

   * The reproducer script simulates concurrent insertion situations
     using OpenStack CLI.

  [ Where problems could occur ]

   * The nature if this change is very simple. In the HA API, the
     create notification function will be modified to introduce an
     artificial pause, significantly reducing the chance of
     concurrent insertion if event records in database.

     This will only be invoked if coordination back-end is not
     configured. Please see (3) for more details:

     if not CONF.coordination.backend_url:
             time.sleep(random.uniform(1, 5))

   * The probable risk here is a delay between 1-5 seconds during
     event creation and further processing. These events happen
     for example when a compute node becomes inaccessible and the
     instances need to be evacuated and launched on a new node.
     In other words the recovery will be delayed for 1-5 seconds.

  [ Other Info ]

   * There is currently an effort to merge this change upstream (4).

  (1) https://bugs.launchpad.net/masakari/+bug/2028450
  (2) https://launchpadlibrarian.net/849768758/lp2028450-reproducer.bash
  (3) https://review.opendev.org/c/openstack/masakari/+/978343/3/masakari/ha/api.py
  (4) https://review.opendev.org/c/openstack/masakari/+/978343

To manage notifications about this bug go to:
https://bugs.launchpad.net/cloud-archive/+bug/2146964/+subscriptions




More information about the Ubuntu-openstack-bugs mailing list