Only the ovn-controller that runs a BFD session writes its status in the
Southbound BFD table, and only when the state of the session changes.
When that ovn-controller exits with cleanup, it releases the gateway
port binding and deletes its Chassis row, but the status stays as it
was, usually "up". If no other chassis can take the gateway over,
nobody writes the row again. ovn-northd then keeps the route's nexthop
in the ECMP set, and the traffic hashed to it is lost.
During a cleanup exit ("exit" without "--restart", with
ovn-cleanup-on-exit not set to false), ovn-controller now first sets
the status to "down" for each BFD session that it runs, whose status is
"up" or "init", and whose gateway has no other registered chassis: the
port is an l3gateway port, or the HA chassis group of its
chassisredirect binding has no other member whose Chassis row exists.
chassis_name is never written.
The update has its own transaction, committed before the cleanup loop
releases the bindings. New attempts start only in the first second,
and there are at most three: all rows, all rows again after a transient
failure, and after any other failure, such as an SB RBAC rejection,
only the rows whose chassis_name is this chassis. ovn-controller stops
waiting for the Southbound database after 1.5 seconds, so the release
is delayed by at most that much. An INFO message reports the sessions
marked down, and a WARN names those that could not be marked.
A stale "up" still stays when:
- ovn-controller is killed or stopped by a signal, including SIGTERM,
or the host fails;
- "exit --restart" or ovn-cleanup-on-exit=false keep the binding;
- another member of the HA chassis group is registered but dead, or
exits before it has taken the gateway over;
- the HA chassis group has another registered member that cannot
really take the gateway over, as in the per-router groups that
Neutron fills with every gateway chassis.
If a CMS moves a gateway router to another chassis when its chassis
exits, as Neutron does for a router pinned with options:chassis, the
router's routes are removed until the sessions on the new chassis come
up. With SB RBAC, the new chassis can write "up" only once the row's
chassis_name names it.
Reported-at: https://github.com/ovn-org/ovn/issues/320
Submitted-at: https://github.com/ovn-org/ovn/pull/332
Assisted-by: Claude Opus 5.5 (claude-opus-5-5), Cursor Grok Bot / Ultimum
harness assistants
Signed-off-by: Premysl Kouril <[email protected]>
---
NEWS | 8 +
controller/ha-chassis.c | 25 ++
controller/ha-chassis.h | 4 +
controller/ovn-controller.8.xml | 30 ++-
controller/ovn-controller.c | 192 ++++++++++++++
controller/pinctrl.c | 90 +++++++
controller/pinctrl.h | 10 +
ovn-sb.xml | 9 +
tests/ovn.at | 441 ++++++++++++++++++++++++++++++++
9 files changed, 805 insertions(+), 4 deletions(-)
diff --git a/NEWS b/NEWS
index 7f94d0b14..dece17b50 100644
--- a/NEWS
+++ b/NEWS
@@ -19,6 +19,14 @@ Post v26.09.0
- Removed OVN's ovs-bugtool plugin and helper scripts.
- Removed ovn-sim utility scripts.
- Removed ovn-docker utility scripts.
+ - ovn-controller: A cleanup exit ("exit" without "--restart") now marks
+ down the BFD sessions that the chassis runs when no other chassis in
+ the gateway's HA chassis group is registered, before it releases the
+ port bindings. Their routes are then removed instead of staying in
+ place with a stale "up" status. A packaged "systemctl restart" of
+ ovn-controller also does a cleanup exit, so the only gateway's routes
+ are briefly removed until its sessions come up again; use
+ "ovn-ctl restart_controller" to avoid that.
OVN v26.09.0 - xxx xx xxxx
--------------------------
diff --git a/controller/ha-chassis.c b/controller/ha-chassis.c
index ad0b3ef0b..d6ca5e52e 100644
--- a/controller/ha-chassis.c
+++ b/controller/ha-chassis.c
@@ -217,6 +217,31 @@ ha_chassis_group_contains(
return false;
}
+/* Returns true if 'ha_chassis_grp' has an HA chassis other than
+ * 'local_chassis' that is registered, that is, whose Chassis row still
+ * exists. HA_Chassis refers to its Chassis weakly, so the reference is
+ * cleared when a chassis exits with cleanup or is removed with
+ * "ovn-sbctl chassis-del". This tells nothing about whether that chassis
+ * is alive or reachable. */
+bool
+ha_chassis_group_has_other_registered(
+ const struct sbrec_ha_chassis_group *ha_chassis_grp,
+ const struct sbrec_chassis *local_chassis)
+{
+ if (!ha_chassis_grp) {
+ return false;
+ }
+
+ for (size_t i = 0; i < ha_chassis_grp->n_ha_chassis; i++) {
+ const struct sbrec_chassis *chassis
+ = ha_chassis_grp->ha_chassis[i]->chassis;
+ if (chassis && chassis != local_chassis) {
+ return true;
+ }
+ }
+ return false;
+}
+
struct ha_chassis_ordered *
ha_chassis_get_ordered(const struct sbrec_ha_chassis_group *ha_chassis_grp)
{
diff --git a/controller/ha-chassis.h b/controller/ha-chassis.h
index c7c91e001..fe6c3ea62 100644
--- a/controller/ha-chassis.h
+++ b/controller/ha-chassis.h
@@ -41,6 +41,10 @@ bool ha_chassis_group_contains(
const struct sbrec_ha_chassis_group *ha_chassis_grp,
const struct sbrec_chassis *chassis);
+bool ha_chassis_group_has_other_registered(
+ const struct sbrec_ha_chassis_group *ha_chassis_grp,
+ const struct sbrec_chassis *local_chassis);
+
struct ha_chassis_ordered *ha_chassis_get_ordered(
const struct sbrec_ha_chassis_group *ha_chassis_grp);
diff --git a/controller/ovn-controller.8.xml b/controller/ovn-controller.8.xml
index 3c33654ff..b7cc5628d 100644
--- a/controller/ovn-controller.8.xml
+++ b/controller/ovn-controller.8.xml
@@ -454,7 +454,9 @@
<ref table="Chassis_Private" db="OVN_Southbound"/>,
<ref table="Encap" db="OVN_Southbound"/>, and
<ref table="IGMP_Group" db="OVN_Southbound"/> rows when it terminates
- through the <code>exit</code> runtime command. The default value is
+ through the <code>exit</code> runtime command. Before that, it
+ marks down the BFD sessions that no other chassis could take over,
+ as described for <code>exit</code> below. The default value is
<code>true</code>.
</p>
<p>
@@ -907,9 +909,29 @@
<dl>
<dt><code>exit</code> [<code>--restart</code>]</dt>
<dd>
- Causes <code>ovn-controller</code> to gracefully terminate. With
- <code>--restart</code>, it skips the database cleanup normally done
- on exit.
+ <p>
+ Causes <code>ovn-controller</code> to gracefully terminate. With
+ <code>--restart</code>, it skips the database cleanup normally done
+ on exit.
+ </p>
+
+ <p>
+ The cleanup first sets the <ref table="BFD" column="status"
+ db="OVN_Southbound"/> of a BFD session that this chassis runs to
+ <code>down</code>, if the status is <code>up</code> or
+ <code>init</code> and no other chassis in the HA chassis group of
+ the gateway port is registered. Nobody else could take such a
+ session over, so its status would otherwise stay unchanged after
+ the exit. This update is committed before the port bindings are
+ released. <code>ovn-controller</code> gives it at most about 1.5
+ seconds: if the Southbound database rejects it or does not answer
+ in time, <code>ovn-controller</code> logs a warning and continues
+ the exit. With role-based access control in the Southbound
+ database, the update is allowed only for rows whose
+ <ref table="BFD" column="chassis_name" db="OVN_Southbound"/> is
+ this chassis. Nothing is updated when <code>ovn-controller</code>
+ terminates any other way, for example on <code>SIGTERM</code>.
+ </p>
</dd>
<dt><code>ct-zone-list</code></dt>
diff --git a/controller/ovn-controller.c b/controller/ovn-controller.c
index d98cb721c..1a397a08e 100644
--- a/controller/ovn-controller.c
+++ b/controller/ovn-controller.c
@@ -29,6 +29,7 @@
#include "chassis.h"
#include "command-line.h"
#include "compiler.h"
+#include "coverage.h"
#include "daemon.h"
#include "dirs.h"
#include "openvswitch/dynamic-string.h"
@@ -108,6 +109,9 @@
VLOG_DEFINE_THIS_MODULE(main);
+COVERAGE_DEFINE(bfd_exit_mark_down);
+COVERAGE_DEFINE(bfd_exit_mark_down_failed);
+
static unixctl_cb_func ct_zone_list;
static unixctl_cb_func extend_table_list;
static unixctl_cb_func inject_pkt;
@@ -7846,6 +7850,187 @@ ovsdb_idl_loop_next_cfg_inc(struct ovsdb_idl_loop
*idl_loop)
}
}
+/* During a cleanup exit, bfd_exit_mark_down() makes at most
+ * BFD_EXIT_MAX_ATTEMPTS attempts, starts them only within
+ * BFD_EXIT_NEW_ATTEMPT_MSEC, and stops waiting for the Southbound database
+ * after BFD_EXIT_HARD_CAP_MSEC, so that releasing the port bindings is
+ * delayed by at most that much. */
+#define BFD_EXIT_MAX_ATTEMPTS 3
+#define BFD_EXIT_NEW_ATTEMPT_MSEC 1000
+#define BFD_EXIT_HARD_CAP_MSEC 1500
+
+/* Marks down the BFD sessions that this chassis runs and that no other
+ * registered chassis could take over (see pinctrl_bfd_exit_mark_down()),
+ * in a Southbound transaction of its own that completes before the cleanup
+ * loop releases the port bindings.
+ *
+ * The first attempt covers all such sessions. After a transient failure
+ * (TXN_TRY_AGAIN) it is retried once. After any other failure, for example
+ * an SB RBAC rejection, one more attempt covers only the sessions whose
+ * chassis_name is this chassis, because SB RBAC allows only those.
+ *
+ * On return no transaction is open on either IDL loop and none is in
+ * flight on the Southbound one, so the cleanup loop starts from a clean
+ * state, and at once. */
+static void
+bfd_exit_mark_down(struct ovsdb_idl_loop *ovs_idl_loop,
+ struct ovsdb_idl_loop *ovnsb_idl_loop,
+ struct ovsdb_idl_index *sbrec_chassis_by_name,
+ struct ovsdb_idl_index *sbrec_port_binding_by_name,
+ struct shash *vif_plug_deleted_iface_ids,
+ struct shash *vif_plug_changed_iface_ids)
+{
+ if (!ovsdb_idl_has_ever_connected(ovnsb_idl_loop->idl)) {
+ return;
+ }
+
+ long long int start = time_msec();
+ int n_all = 0; /* Attempts on all rows. */
+ int n_subset = 0; /* Attempts on the chassis_name subset. */
+ bool subset = false; /* Whether the next attempt is on the subset. */
+ bool in_flight = false; /* Whether our attempt is being committed. */
+ bool stop = false; /* Whether no new attempt may start. */
+ bool dropped = false; /* Whether our attempt got no answer in time. */
+ size_t n_sent = 0; /* Rows in the last attempt. */
+ size_t n_marked = 0; /* Rows marked down by successful attempts. */
+ const struct sbrec_chassis *chassis;
+
+ for (;;) {
+ update_sb_db(ovs_idl_loop->idl, ovnsb_idl_loop->idl,
+ NULL, NULL, NULL, NULL);
+ update_ssl_config(ovsrec_ssl_table_get(ovs_idl_loop->idl));
+
+ ovsdb_idl_loop_run(ovs_idl_loop);
+ struct ovsdb_idl_txn *ovnsb_idl_txn
+ = ovsdb_idl_loop_run(ovnsb_idl_loop);
+
+ const char *chassis_id
+ = get_ovs_chassis_id(ovsrec_open_vswitch_table_get(
+ ovs_idl_loop->idl));
+ chassis = (chassis_id
+ ? chassis_lookup_by_name(sbrec_chassis_by_name, chassis_id)
+ : NULL);
+ long long int now = time_msec();
+
+ /* The outcome of our attempt, if it completed in this pass. It is
+ * taken from the transaction: while the transaction is in flight,
+ * the IDL still shows the old status. */
+ enum ovsdb_idl_txn_status status = TXN_INCOMPLETE;
+ if (in_flight) {
+ /* ovsdb_idl_loop_run() reaps only a successful transaction. On
+ * a transaction that was already sent, ovsdb_idl_txn_commit()
+ * just returns its status. */
+ status = (ovnsb_idl_loop->committing_txn
+ ? ovsdb_idl_txn_commit(ovnsb_idl_loop->committing_txn)
+ : TXN_SUCCESS);
+ } else if (!stop) {
+ if (!chassis || now - start >= BFD_EXIT_NEW_ATTEMPT_MSEC
+ || n_all + n_subset >= BFD_EXIT_MAX_ATTEMPTS) {
+ stop = true;
+ } else if (ovnsb_idl_txn) {
+ n_sent = pinctrl_bfd_exit_mark_down(
+ ovnsb_idl_txn, sbrec_bfd_table_get(ovnsb_idl_loop->idl),
+ sbrec_port_binding_by_name, chassis, subset);
+ if (!n_sent) {
+ stop = true;
+ } else {
+ in_flight = true;
+ if (subset) {
+ n_subset++;
+ } else {
+ n_all++;
+ }
+ /* Send it now, so that a failure that is known at once,
+ * for example without a connection, is seen here. */
+ status = ovsdb_idl_txn_commit(ovnsb_idl_txn);
+ }
+ }
+ }
+
+ /* Always commit, so that no transaction is left open. */
+ if (!ovsdb_idl_loop_commit_and_wait(ovnsb_idl_loop)) {
+ /* After a failure the IDL does not ask to be woken up. */
+ poll_immediate_wake();
+ }
+ int ovs_txn_status = ovsdb_idl_loop_commit_and_wait(ovs_idl_loop);
+ if (!ovs_txn_status) {
+ vif_plug_clear_deleted(vif_plug_deleted_iface_ids);
+ vif_plug_clear_changed(vif_plug_changed_iface_ids);
+ } else if (ovs_txn_status == 1) {
+ vif_plug_finish_deleted(vif_plug_deleted_iface_ids);
+ vif_plug_finish_changed(vif_plug_changed_iface_ids);
+ }
+
+ if (in_flight && status != TXN_INCOMPLETE) {
+ in_flight = false;
+ if (status == TXN_SUCCESS || status == TXN_UNCHANGED) {
+ /* The next pass finds what is left, normally nothing. */
+ COVERAGE_ADD(bfd_exit_mark_down, n_sent);
+ n_marked += n_sent;
+ } else if (status == TXN_TRY_AGAIN && !subset && n_all < 2) {
+ /* Try all rows again once the IDL has caught up. */
+ } else if (!subset) {
+ subset = true;
+ } else {
+ stop = true;
+ }
+ poll_immediate_wake();
+ }
+
+ /* Also wait for a transaction that the main loop left in flight. */
+ if (!ovnsb_idl_loop->committing_txn) {
+ if (stop) {
+ break;
+ }
+ } else if (now - start >= BFD_EXIT_HARD_CAP_MSEC) {
+ /* The server does not answer. Forget the transaction, so that
+ * the cleanup loop does not inherit it; a late reply is
+ * ignored. */
+ ovsdb_idl_txn_destroy(ovnsb_idl_loop->committing_txn);
+ ovnsb_idl_loop->committing_txn = NULL;
+ dropped = in_flight;
+ break;
+ }
+
+ poll_timer_wait_until(start + (stop || in_flight
+ ? BFD_EXIT_HARD_CAP_MSEC
+ : BFD_EXIT_NEW_ATTEMPT_MSEC));
+ poll_block();
+ }
+
+ long long int elapsed = time_msec() - start;
+ if (dropped) {
+ VLOG_WARN("Southbound database did not answer the BFD status update "
+ "within %d ms; dropped it and continuing exit.",
+ BFD_EXIT_HARD_CAP_MSEC);
+ }
+
+ /* Nothing is in flight now, so the IDL shows what the database has. */
+ struct ds rows = DS_EMPTY_INITIALIZER;
+ size_t n_left = (chassis
+ ? pinctrl_bfd_exit_pending(
+ sbrec_bfd_table_get(ovnsb_idl_loop->idl),
+ sbrec_port_binding_by_name, chassis, &rows)
+ : 0);
+ if (n_left) {
+ COVERAGE_ADD(bfd_exit_mark_down_failed, n_left);
+ VLOG_WARN("Could not mark %"PRIuSIZE" BFD session(s) down during "
+ "exit cleanup (%"PRIuSIZE" marked, %d attempt(s) on all "
+ "rows, %d on rows with this chassis_name, %lld ms); "
+ "continuing exit. If SB RBAC is enabled, check BFD "
+ "chassis_name: %s",
+ n_left, n_marked, n_all, n_subset, elapsed, ds_cstr(&rows));
+ } else if (n_marked) {
+ VLOG_INFO("Marked %"PRIuSIZE" BFD session(s) down before releasing "
+ "gateway bindings (no other registered HA chassis).",
+ n_marked);
+ }
+ ds_destroy(&rows);
+
+ /* Let the cleanup loop start at once. */
+ poll_immediate_wake();
+}
+
int
main(int argc, char *argv[])
{
@@ -8816,6 +9001,13 @@ loop_done:
/* It's time to exit. Clean up the databases if we are not restarting */
if (!restart) {
+ /* Before the bindings are released, so that nobody sees them
+ * released while the BFD sessions still look up. */
+ bfd_exit_mark_down(&ovs_idl_loop, &ovnsb_idl_loop,
+ sbrec_chassis_by_name, sbrec_port_binding_by_name,
+ &vif_plug_deleted_iface_ids,
+ &vif_plug_changed_iface_ids);
+
bool done = !ovsdb_idl_has_ever_connected(ovnsb_idl_loop.idl);
while (!done) {
update_sb_db(ovs_idl_loop.idl, ovnsb_idl_loop.idl,
diff --git a/controller/pinctrl.c b/controller/pinctrl.c
index 2a29c5e5a..9fcd42c0d 100644
--- a/controller/pinctrl.c
+++ b/controller/pinctrl.c
@@ -8307,6 +8307,96 @@ bfd_monitor_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
}
}
+/* Returns true if a cleanup exit of 'chassis' should mark the BFD session
+ * of 'bt' down: 'chassis' runs the session, its status is "up" or "init",
+ * and no other chassis that could take the gateway over is registered.
+ * With 'subset', also requires 'bt''s chassis_name to be 'chassis'. */
+static bool
+bfd_exit_should_mark_down(struct ovsdb_idl_index *sbrec_port_binding_by_name,
+ const struct sbrec_bfd *bt,
+ const struct sbrec_chassis *chassis, bool subset)
+{
+ if (strcmp(bt->status, "up") && strcmp(bt->status, "init")) {
+ return false;
+ }
+ if (subset && strcmp(bt->chassis_name, chassis->name)) {
+ return false;
+ }
+
+ const struct sbrec_port_binding *owner_pb
+ = bfd_session_owner_pb(sbrec_port_binding_by_name, bt, chassis, NULL);
+ if (!owner_pb) {
+ return false;
+ }
+
+ /* An l3gateway binding has no HA chassis group, and a chassisredirect
+ * binding may have none either. Then 'chassis' is the only candidate. */
+ return !ha_chassis_group_has_other_registered(owner_pb->ha_chassis_group,
+ chassis);
+}
+
+/* Called by ovn-controller during a cleanup exit, after its main loop has
+ * ended and before it releases its port bindings. Sets to "down" the status
+ * of every BFD session that 'chassis' runs, if the status is "up" or "init"
+ * and no other chassis of the gateway's HA chassis group is registered.
+ * Nobody could take such a session over, so without this its status would
+ * stay stale after the exit.
+ *
+ * With 'subset', considers only the rows whose chassis_name is 'chassis':
+ * with SB RBAC, those are the only rows 'chassis' may update.
+ *
+ * Returns the number of rows set. This reads only Southbound data, so it
+ * does not need 'pinctrl_mutex'. */
+size_t
+pinctrl_bfd_exit_mark_down(struct ovsdb_idl_txn *ovnsb_idl_txn,
+ const struct sbrec_bfd_table *bfd_table,
+ struct ovsdb_idl_index *sbrec_port_binding_by_name,
+ const struct sbrec_chassis *chassis, bool subset)
+{
+ if (!ovnsb_idl_txn) {
+ return 0;
+ }
+
+ size_t n = 0;
+
+ const struct sbrec_bfd *bt;
+ SBREC_BFD_TABLE_FOR_EACH (bt, bfd_table) {
+ if (bfd_exit_should_mark_down(sbrec_port_binding_by_name, bt, chassis,
+ subset)) {
+ VLOG_DBG("Marking BFD session down before exit cleanup: "
+ "logical_port %s, dst_ip %s, status %s",
+ bt->logical_port, bt->dst_ip, bt->status);
+ sbrec_bfd_set_status(bt, "down");
+ n++;
+ }
+ }
+ return n;
+}
+
+/* Returns the number of BFD rows that pinctrl_bfd_exit_mark_down() would
+ * still set to "down" for 'chassis' without 'subset', and appends a short
+ * description of each of them to 'rows'. */
+size_t
+pinctrl_bfd_exit_pending(const struct sbrec_bfd_table *bfd_table,
+ struct ovsdb_idl_index *sbrec_port_binding_by_name,
+ const struct sbrec_chassis *chassis, struct ds *rows)
+{
+ size_t n = 0;
+
+ const struct sbrec_bfd *bt;
+ SBREC_BFD_TABLE_FOR_EACH (bt, bfd_table) {
+ if (bfd_exit_should_mark_down(sbrec_port_binding_by_name, bt, chassis,
+ false)) {
+ ds_put_format(rows, "%slogical_port %s dst_ip %s (status %s, "
+ "chassis_name \"%s\")", n ? "; " : "",
+ bt->logical_port, bt->dst_ip, bt->status,
+ bt->chassis_name);
+ n++;
+ }
+ }
+ return n;
+}
+
static uint16_t
get_random_src_port(void)
{
diff --git a/controller/pinctrl.h b/controller/pinctrl.h
index a638fe29f..235b0e5de 100644
--- a/controller/pinctrl.h
+++ b/controller/pinctrl.h
@@ -23,6 +23,7 @@
#include "openvswitch/list.h"
#include "openvswitch/meta-flow.h"
+struct ds;
struct hmap;
struct shash;
struct lport_index;
@@ -63,6 +64,15 @@ void pinctrl_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
int64_t cur_cfg);
void pinctrl_wait(struct ovsdb_idl_txn *ovnsb_idl_txn);
void pinctrl_destroy(void);
+size_t pinctrl_bfd_exit_mark_down(
+ struct ovsdb_idl_txn *ovnsb_idl_txn,
+ const struct sbrec_bfd_table *,
+ struct ovsdb_idl_index *sbrec_port_binding_by_name,
+ const struct sbrec_chassis *chassis, bool subset);
+size_t pinctrl_bfd_exit_pending(
+ const struct sbrec_bfd_table *,
+ struct ovsdb_idl_index *sbrec_port_binding_by_name,
+ const struct sbrec_chassis *chassis, struct ds *rows);
void pinctrl_seqno_run(void);
void pinctrl_seqno_flush(void);
diff --git a/ovn-sb.xml b/ovn-sb.xml
index 2096fc3e3..d7ca33891 100644
--- a/ovn-sb.xml
+++ b/ovn-sb.xml
@@ -5590,6 +5590,15 @@ tcp.flags = RST;
</li>
</ul>
</p>
+
+ <p>
+ The <code>ovn-controller</code> that runs the session updates this
+ column when the state of the session changes. During a cleanup
+ exit, it also sets <code>down</code> for the sessions it runs if no
+ other chassis in the HA chassis group of the gateway port is
+ registered. See <code>exit</code> in
+ <code>ovn-controller</code>(8).
+ </p>
</column>
</group>
</table>
diff --git a/tests/ovn.at b/tests/ovn.at
index 783c822a3..9c5d6ebd5 100644
--- a/tests/ovn.at
+++ b/tests/ovn.at
@@ -47545,3 +47545,444 @@ AT_CHECK([grep "skipping output to input port" \
OVN_CLEANUP([hv1])
AT_CLEANUP
])
+
+dnl Helpers for the "BFD - cleanup exit" tests. No BFD peer answers in
+dnl these tests, so they set the SB status to "up" as if the session had
+dnl come up.
+m4_define([BFD_EXIT_HELPERS], [
+# bfd_exit_add_hv HV IP [rbac]
+#
+# Adds sandbox HV. Unless "rbac" is given, its ovn-controller then uses an
+# SB connection without an RBAC role. (With SSL, the default connection has
+# role ovn-controller, see ovn_start.)
+bfd_exit_add_hv() {
+ local hv=${1}
+ sim_add $hv
+ as $hv
+ check ovs-vsctl add-br br-phys
+ ovn_attach n1 br-phys ${2}
+ if test "${3}" != rbac && test X$HAVE_OPENSSL = Xyes; then
+ check ovs-vsctl set open . \
+ external-ids:ovn-remote=unix:$ovs_base/ovn-sb/ovn-sb.sock
+ OVS_WAIT_UNTIL([grep -q 'ovn-sb.sock: connected'
$hv/ovn-controller.log])
+ fi
+}
+
+# bfd_exit_add_gw PORT N HV[:PRIORITY]...
+#
+# Adds gateway port PORT (172.16.N.1/24) to lr0, with the given gateway
+# chassis, and a route to 10.N.0.0/16 via 172.16.N.100 monitored by BFD.
+bfd_exit_add_gw() {
+ local port=${1} n=${2} hv prio
+ shift 2
+ check ovn-nbctl lrp-add lr0 $port 00:00:00:00:ff:0$n 172.16.$n.1/24
+ check ovn-nbctl ls-add ls-$port
+ check ovn-nbctl lsp-add-router-port ls-$port ls-$port-lr0 $port
+ for hv in "${@}"; do
+ prio=${hv#*:}
+ test "$prio" = "$hv" && prio=
+ check ovn-nbctl lrp-set-gateway-chassis $port ${hv%%:*} $prio
+ done
+ check ovn-nbctl --bfd lr-route-add lr0 10.$n.0.0/16 172.16.$n.100 $port
+}
+
+# bfd_exit_wait_claimed PORT HV
+#
+# Waits until the chassisredirect binding of PORT is bound to HV.
+bfd_exit_wait_claimed() {
+ wait_row_count Chassis 1 name=${2}
+ wait_column "$(fetch_column Chassis _uuid name=${2})" Port_Binding \
+ chassis logical_port=cr-${1}
+}
+
+# bfd_exit_set_status DST_IP STATUS
+#
+# Sets the SB status of the BFD row for DST_IP and waits until ovn-northd
+# has copied it to NB.
+bfd_exit_set_status() {
+ check ovn-sbctl set BFD $(fetch_column BFD _uuid dst_ip=${1}) status=${2}
+ wait_column ${2} nb:BFD status dst_ip=${1}
+}
+
+# bfd_exit_sb_txns
+#
+# Stores the transactions that ovn-controllers sent to the SB database in
+# sb-txns.log, one per line, in the order the database received them.
+bfd_exit_sb_txns() {
+ grep 'received request, method="transact"' ovn-sb/ovsdb-server.log \
+ | grep '"comment":"ovn-controller' > sb-txns.log
+}
+
+# bfd_exit_bfd_update DST_IP
+#
+# Prints a regular expression that matches the update of the BFD row for
+# DST_IP in a transaction.
+bfd_exit_bfd_update() {
+ echo
"\"table\":\"BFD\",\"where\":..\"_uuid\",\"==\",.\"uuid\",\"$(fetch_column BFD
_uuid dst_ip=${1})\""
+}
+
+# bfd_exit_check_no_write DST_IP...
+#
+# Checks that no ovn-controller updated the BFD rows for DST_IP....
+bfd_exit_check_no_write() {
+ local ip
+ bfd_exit_sb_txns
+ for ip in "${@}"; do
+ check test "$(grep -c "$(bfd_exit_bfd_update $ip)" sb-txns.log)" = 0
+ done
+}
+
+# bfd_exit_check_order DST_IP...
+#
+# Checks that the BFD rows for DST_IP... were set down by one transaction
+# that updated nothing but BFD rows, and that the SB database received it
+# before the transaction that released the port bindings.
+bfd_exit_check_order() {
+ local ip down release
+ bfd_exit_sb_txns
+ down=$(grep -n '"row":{"status":"down"},"table":"BFD"' sb-txns.log \
+ | tail -1 | cut -d: -f1)
+ release=$(grep -n '"row":{"chassis":."set",...},"table":"Port_Binding"' \
+ sb-txns.log | tail -1 | cut -d: -f1)
+ echo "BFD down in transaction $down, release in transaction $release"
+ check test -n "$down"
+ check test -n "$release"
+ check test "$down" -lt "$release"
+ sed -n "${down}p" sb-txns.log > sb-txn-down.log
+ check test "$(grep -o '"table":"' sb-txn-down.log | wc -l)" = \
+ "$(grep -o '"table":"BFD"' sb-txn-down.log | wc -l)"
+ for ip in "${@}"; do
+ check grep -q "\"row\":{\"status\":\"down\"},$(bfd_exit_bfd_update
$ip)" \
+ sb-txn-down.log
+ done
+}
+
+# bfd_exit_log_count HV PATTERN
+#
+# Prints how many lines of HV's ovn-controller log match PATTERN.
+bfd_exit_log_count() {
+ grep -c "${2}" ${1}/ovn-controller.log
+}
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - cleanup exit marks the only gateway's sessions down])
+AT_KEYWORDS([bfd bfd-exit])
+BFD_EXIT_HELPERS
+ovn_start
+net_add n1
+bfd_exit_add_hv hv1 192.168.0.1
+
+# The session to 172.16.1.101 stays down. No route uses 172.16.1.200, so
+# ovn-northd keeps that row admin_down.
+check ovn-nbctl lr-add lr0
+bfd_exit_add_gw lr0-pub 1 hv1
+check ovn-nbctl --bfd lr-route-add lr0 10.11.0.0/16 172.16.1.101 lr0-pub
+check_uuid ovn-nbctl create BFD logical_port=lr0-pub dst_ip=172.16.1.200
+check ovn-nbctl lr-add lr1 -- set Logical_Router lr1 options:chassis=hv1
+check ovn-nbctl lrp-add lr1 lr1-pub 00:00:00:00:ff:07 172.16.7.1/24
+check ovn-nbctl ls-add ls-lr1-pub
+check ovn-nbctl lsp-add-router-port ls-lr1-pub ls-lr1-pub-lr1 lr1-pub
+check ovn-nbctl --bfd lr-route-add lr1 10.7.0.0/16 172.16.7.100 lr1-pub
+bfd_exit_wait_claimed lr0-pub hv1
+wait_column "$(fetch_column Chassis _uuid name=hv1)" Port_Binding chassis \
+ logical_port=lr1-pub
+wait_column admin_down BFD status dst_ip=172.16.1.200
+wait_row_count BFD 3 status=down
+bfd_exit_set_status 172.16.1.100 up
+bfd_exit_set_status 172.16.7.100 up
+check ovn-nbctl --wait=sb sync
+AT_CHECK([ovn-sbctl lflow-list lr0 | grep -c 'reg0 = 172.16.1.100'], [0],
+ [ignore])
+
+OVN_CONTROLLER_EXIT([hv1])
+
+# The sessions are down, and the route via 172.16.1.100 is gone.
+check_column down BFD status dst_ip=172.16.1.100
+check_column down BFD status dst_ip=172.16.7.100
+wait_column down nb:BFD status dst_ip=172.16.1.100
+check ovn-nbctl --wait=sb sync
+AT_CHECK([ovn-sbctl lflow-list lr0 | grep -c 'reg0 = 172.16.1.100'], [1], [0
+])
+AT_CHECK([bfd_exit_log_count hv1 'Marked 2 BFD session(s) down'], [0], [1
+])
+check_column admin_down BFD status dst_ip=172.16.1.200
+bfd_exit_check_no_write 172.16.1.101 172.16.1.200
+
+# The cleanup still ran, and only after the BFD update was committed.
+check_row_count Chassis 0 name=hv1
+check_column "" Port_Binding chassis logical_port=cr-lr0-pub
+check_column "" Port_Binding chassis logical_port=lr1-pub
+bfd_exit_check_order 172.16.1.100 172.16.7.100
+
+# The chassis of this ovn-controller was deleted just before the exit.
+# While paused, ovn-controller does not see the deletion. It also acts on
+# "exit" only after something else wakes it up, hence the second request.
+as hv1 start_daemon ovn-controller
+bfd_exit_wait_claimed lr0-pub hv1
+bfd_exit_set_status 172.16.1.101 up
+pid=$(cat hv1/ovn-controller.pid)
+check as hv1 ovn-appctl -t ovn-controller vlog/set unixctl:file:dbg
+check as hv1 ovn-appctl -t ovn-controller debug/pause
+check ovn-sbctl chassis-del hv1
+check_column "" Port_Binding chassis logical_port=cr-lr0-pub
+as hv1 ovn-appctl -t ovn-controller exit > exit.out 2>&1 &
+exit_pid=$!
+OVS_WAIT_UNTIL([grep -q 'received request exit' hv1/ovn-controller.log])
+as hv1 ovn-appctl -t ovn-controller debug/status > /dev/null 2>&1 &
+exit_status=0
+wait $exit_pid || exit_status=$?
+check test $exit_status = 0
+OVS_WAIT_WHILE([kill -0 $pid 2>/dev/null])
+check_row_count Chassis 0 name=hv1
+check_column up BFD status dst_ip=172.16.1.101
+bfd_exit_check_no_write 172.16.1.101
+AT_CHECK([bfd_exit_log_count hv1 'BFD session(s) down'], [0], [1
+])
+AT_CHECK([grep -i 'assert\|backtrace' hv1/ovn-controller.log], [1])
+
+as hv1 start_daemon ovn-controller
+bfd_exit_wait_claimed lr0-pub hv1
+OVN_CLEANUP([hv1])
+AT_CLEANUP
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - cleanup exit with another registered gateway chassis])
+AT_KEYWORDS([bfd bfd-exit])
+BFD_EXIT_HELPERS
+ovn_start
+net_add n1
+bfd_exit_add_hv hv1 192.168.0.1
+bfd_exit_add_hv hv2 192.168.0.2
+
+check ovn-nbctl lr-add lr0
+bfd_exit_add_gw lr0-pub 1 hv1:20 hv2:10
+ovn_wait_for_bfd_up hv1 hv2
+bfd_exit_wait_claimed lr0-pub hv1
+wait_column down BFD status dst_ip=172.16.1.100
+bfd_exit_set_status 172.16.1.100 up
+
+# hv2 is registered and can take over, so hv1 writes nothing.
+OVN_CONTROLLER_EXIT([hv1])
+check_column up BFD status dst_ip=172.16.1.100
+bfd_exit_check_no_write 172.16.1.100
+AT_CHECK([bfd_exit_log_count hv1 'BFD session(s) down'], [1], [0
+])
+bfd_exit_wait_claimed lr0-pub hv2
+
+# Now hv2 is the last registered member.
+check_row_count HA_Chassis 2
+OVN_CONTROLLER_EXIT([hv2])
+check_column down BFD status dst_ip=172.16.1.100
+AT_CHECK([bfd_exit_log_count hv2 'Marked 1 BFD session(s) down'], [0], [1
+])
+bfd_exit_check_order 172.16.1.100
+
+as hv1 start_daemon ovn-controller
+as hv2 start_daemon ovn-controller
+ovn_wait_for_bfd_up hv1 hv2
+bfd_exit_wait_claimed lr0-pub hv1
+OVN_CLEANUP([hv1
+/Incorrect your_disc/d
+/No tunnel endpoint found for HA chassis in HA chassis group/d
+], [hv2
+/Incorrect your_disc/d
+/No tunnel endpoint found for HA chassis in HA chassis group/d
+])
+AT_CLEANUP
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - no BFD update on exit without cleanup])
+AT_KEYWORDS([bfd bfd-exit])
+BFD_EXIT_HELPERS
+ovn_start
+net_add n1
+bfd_exit_add_hv hv1 192.168.0.1
+
+check ovn-nbctl lr-add lr0
+bfd_exit_add_gw lr0-pub 1 hv1
+bfd_exit_wait_claimed lr0-pub hv1
+wait_column down BFD status dst_ip=172.16.1.100
+bfd_exit_set_status 172.16.1.100 up
+
+OVN_CONTROLLER_EXIT([hv1], [--restart])
+check_column up BFD status dst_ip=172.16.1.100
+check_row_count Chassis 1 name=hv1
+
+as hv1 start_daemon ovn-controller
+check ovn-nbctl --wait=hv sync
+check as hv1 ovs-vsctl set open . external-ids:ovn-cleanup-on-exit=false
+OVN_CONTROLLER_EXIT([hv1])
+AT_CHECK([bfd_exit_log_count hv1 'resource cleanup: False'], [0], [2
+])
+check_column up BFD status dst_ip=172.16.1.100
+check_row_count Chassis 1 name=hv1
+bfd_exit_wait_claimed lr0-pub hv1
+bfd_exit_check_no_write 172.16.1.100
+
+# A later cleanup exit does mark the session down.
+check as hv1 ovs-vsctl remove open . external-ids ovn-cleanup-on-exit
+as hv1 start_daemon ovn-controller
+check ovn-nbctl --wait=hv sync
+OVN_CONTROLLER_EXIT([hv1])
+check_column down BFD status dst_ip=172.16.1.100
+
+as hv1 start_daemon ovn-controller
+bfd_exit_wait_claimed lr0-pub hv1
+OVN_CLEANUP([hv1])
+AT_CLEANUP
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - cleanup exit marks down only the sessions it may])
+AT_KEYWORDS([bfd bfd-exit])
+BFD_EXIT_HELPERS
+ovn_start
+net_add n1
+bfd_exit_add_hv hv1 192.168.0.1
+bfd_exit_add_hv hv2 192.168.0.2
+
+# lr0-pub1 and lr0-pub4: only hv1. lr0-pub2: only hv2. lr0-pub3: hv1,
+# with hv2 as backup.
+check ovn-nbctl lr-add lr0
+bfd_exit_add_gw lr0-pub1 1 hv1
+check ovn-nbctl --bfd lr-route-add lr0 10.11.0.0/16 172.16.1.101 lr0-pub1
+bfd_exit_add_gw lr0-pub2 2 hv2
+bfd_exit_add_gw lr0-pub3 3 hv1:20 hv2:10
+bfd_exit_add_gw lr0-pub4 4 hv1
+ovn_wait_for_bfd_up hv1 hv2
+bfd_exit_wait_claimed lr0-pub1 hv1
+bfd_exit_wait_claimed lr0-pub2 hv2
+bfd_exit_wait_claimed lr0-pub3 hv1
+bfd_exit_wait_claimed lr0-pub4 hv1
+wait_row_count BFD 5 status=down
+for ip in 172.16.1.100 172.16.1.101 172.16.2.100 172.16.3.100; do
+ bfd_exit_set_status $ip up
+done
+bfd_exit_set_status 172.16.4.100 init
+
+OVN_CONTROLLER_EXIT([hv1])
+check_column down BFD status dst_ip=172.16.1.100
+check_column down BFD status dst_ip=172.16.1.101
+check_column down BFD status dst_ip=172.16.4.100
+check_column up BFD status dst_ip=172.16.2.100
+check_column up BFD status dst_ip=172.16.3.100
+AT_CHECK([bfd_exit_log_count hv1 'Marked 3 BFD session(s) down'], [0], [1
+])
+bfd_exit_check_order 172.16.1.100 172.16.1.101 172.16.4.100
+bfd_exit_check_no_write 172.16.2.100 172.16.3.100
+
+as hv1 start_daemon ovn-controller
+ovn_wait_for_bfd_up hv1 hv2
+bfd_exit_wait_claimed lr0-pub1 hv1
+bfd_exit_wait_claimed lr0-pub3 hv1
+bfd_exit_wait_claimed lr0-pub4 hv1
+OVN_CLEANUP([hv1
+/Incorrect your_disc/d
+/No tunnel endpoint found for HA chassis in HA chassis group/d
+], [hv2
+/Incorrect your_disc/d
+/No tunnel endpoint found for HA chassis in HA chassis group/d
+])
+AT_CLEANUP
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - cleanup exit when the SB database does not answer])
+AT_KEYWORDS([bfd bfd-exit])
+BFD_EXIT_HELPERS
+ovn_start
+net_add n1
+bfd_exit_add_hv hv1 192.168.0.1
+
+check ovn-nbctl lr-add lr0
+bfd_exit_add_gw lr0-pub 1 hv1
+bfd_exit_wait_claimed lr0-pub hv1
+wait_column down BFD status dst_ip=172.16.1.100
+bfd_exit_set_status 172.16.1.100 up
+
+# The exit command is answered only after the cleanup, which cannot
+# finish while the SB database sleeps.
+pid=$(cat hv1/ovn-controller.pid)
+sleep_sb
+as hv1 ovn-appctl -t ovn-controller exit > exit.out 2>&1 &
+exit_pid=$!
+OVS_WAIT_UNTIL([grep -q 'did not answer the BFD status update' \
+ hv1/ovn-controller.log])
+AT_CHECK([bfd_exit_log_count hv1 'Could not mark 1 BFD session(s) down'],
+ [0], [1
+])
+# ovn-controller gave up at the 1.5 s cap, timed by its own clock. The
+# 0.5 s margin is for a slow wake-up on a busy machine.
+elapsed=$(sed -n 's/.*Could not mark.* \([[0-9]]*\) ms).*/\1/p' \
+ hv1/ovn-controller.log)
+check test "$elapsed" -ge 1500
+check test "$elapsed" -lt 2000
+
+# Then the cleanup runs as before.
+wake_up_sb
+exit_status=0
+wait $exit_pid || exit_status=$?
+check test $exit_status = 0
+OVS_WAIT_WHILE([kill -0 $pid 2>/dev/null])
+check_row_count Chassis 0 name=hv1
+check_column "" Port_Binding chassis logical_port=cr-lr0-pub
+AT_CHECK([grep -i 'assert\|backtrace' hv1/ovn-controller.log], [1])
+
+as hv1 start_daemon ovn-controller
+bfd_exit_wait_claimed lr0-pub hv1
+OVN_CLEANUP([hv1
+/did not answer the BFD status update/d
+/Could not mark 1 BFD session(s) down/d
+])
+AT_CLEANUP
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - cleanup exit under SB RBAC marks down the rows it may])
+AT_KEYWORDS([bfd bfd-exit])
+AT_SKIP_IF([test "$HAVE_OPENSSL" = no])
+BFD_EXIT_HELPERS
+ovn_start
+net_add n1
+bfd_exit_add_hv hv1 192.168.0.1 rbac
+
+check ovn-nbctl lr-add lr0
+bfd_exit_add_gw lr0-pub1 1 hv1
+bfd_exit_add_gw lr0-pub2 2 hv1
+bfd_exit_wait_claimed lr0-pub1 hv1
+bfd_exit_wait_claimed lr0-pub2 hv1
+wait_row_count BFD 2 status=down
+bfd_exit_set_status 172.16.1.100 up
+bfd_exit_set_status 172.16.2.100 up
+
+# With ovn-northd paused, SB RBAC lets hv1 update only the second row.
+check as northd ovn-appctl -t ovn-northd pause
+check ovn-sbctl set BFD $(fetch_column BFD _uuid dst_ip=172.16.1.100) \
+ chassis_name='""' \
+ -- set BFD $(fetch_column BFD _uuid dst_ip=172.16.2.100) chassis_name=hv1
+
+# The update of both rows is rejected, then the second row is updated
+# alone.
+OVN_CONTROLLER_EXIT([hv1])
+check_column up BFD status dst_ip=172.16.1.100
+check_column down BFD status dst_ip=172.16.2.100
+AT_CHECK([grep 'Could not mark' hv1/ovn-controller.log | \
+ sed 's/.*|WARN|//; s/[[0-9]]* ms)/X ms)/'], [0], [dnl
+Could not mark 1 BFD session(s) down during exit cleanup (1 marked, 1
attempt(s) on all rows, 1 on rows with this chassis_name, X ms); continuing
exit. If SB RBAC is enabled, check BFD chassis_name: logical_port lr0-pub1
dst_ip 172.16.1.100 (status up, chassis_name "")
+])
+bfd_exit_check_order 172.16.2.100
+check_row_count Chassis 0 name=hv1
+check as northd ovn-appctl -t ovn-northd resume
+
+as hv1 start_daemon ovn-controller
+bfd_exit_wait_claimed lr0-pub1 hv1
+bfd_exit_wait_claimed lr0-pub2 hv1
+OVN_CLEANUP([hv1
+/Could not mark 1 BFD session(s) down/d
+/transaction error/d
+])
+AT_CLEANUP
+])
--
2.48.1
_______________________________________________
dev mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-dev