Project homepage Mailing List  Warmcat.com  API Docs  Github Mirror 
    npro  
 Modern all-safe Rust Network Protocol library supporting h1, h2, h3, ws, wt sans-IO and with socket IO + tls
git clone https://npro.rs/repo/npro
 
root / src / builder / b-metrics.c
Author[]Andy Green <andy@warmcat.com> 2026-09-20 11:10 UTC
Committer[]Andy Green <andy@warmcat.com> 2026-09-21 07:08 UTC
Tree1b33096e71d3c580ef210e0f4696d3f3fbedd9bd   Raw Patch
 
server, web: survive vhosts where the comms protocol failed init
server, web: survive vhosts where the comms protocol failed init

lws frees a protocol's vh priv when its PROTOCOL_INIT callback returns
nonzero, but it keeps the protocol bound and serving on that vhost, so
connections there arrive with a NULL vhd from lws_protocol_vh_priv_get().
sai-server hits this in practice: the comms protocol is in the
context-wide protocol list, so any conf.d vhost that binds it without
its pvo set (eg, a leftover unixskt-style vhost) fails init at the
required pvos.

The CLOSED handler unconditionally called sais_builder_disconnected()
with that NULL vhd, segfaulting sai-server (valgrind: invalid read of
size 8 at 0x18, ie vhd->server.builder_owner) when a ws conn on such a
vhost was torn down after ESTABLISHED rejected it.  Reject the upgrade
at FILTER_PROTOCOL_CONNECTION with an explicit error instead, and guard
the other vhd-derefing paths (RECEIVE, the CLOSED teardown and the hook
POST handling) so nothing can follow a NULL vhd.

sai-web lands in the same trap from the other side: its PROTOCOL_INIT
returned nonzero when the initial connect to sai-server's control
socket failed synchronously (eg, sai-server not restarted yet), killing
the whole vhost's protocol until the next sai-web restart.  The websrv
stream is nailed_up, so lws already owns the initial attempt and its
backoff; stop issuing an explicit lws_ss_client_connect() from init and
keep init alive across the link being down, so browsers keep a working
protocol and the link comes up when the server appears.
diff --git a/src/server/s-comms.c b/src/server/s-comms.c index a6c3fbf..29198ba 100644 --- a/src/server/s-comms.c +++ b/src/server/s-comms.c @@ -433,6 +433,9 @@ s_callback_ws(struct lws *wsi, enum lws_callback_reasons reason, void *user, goto passthru; case SHMUT_HOOK: + if (!vhd) + /* no db etc without completed protocol init */ + return -1; pss->our_form = 1; lwsl_notice("LWS_CALLBACK_HTTP: sees hook\n"); return 0; @@ -454,6 +457,9 @@ s_callback_ws(struct lws *wsi, enum lws_callback_reasons reason, void *user, goto passthru; } + if (!vhd) + return -1; + /* create the POST argument parser if not already existing */ if (!pss->spa) { @@ -536,6 +542,9 @@ s_callback_ws(struct lws *wsi, enum lws_callback_reasons reason, void *user, goto passthru; } + if (!vhd) + return -1; + if (pss->spa) { lws_spa_finalize(pss->spa); lws_spa_destroy(pss->spa); @@ -572,6 +581,18 @@ s_callback_ws(struct lws *wsi, enum lws_callback_reasons reason, void *user, */ case LWS_CALLBACK_FILTER_PROTOCOL_CONNECTION: + /* + * If PROTOCOL_INIT failed on this vhost (eg, it has the + * protocol bound but not the pvo set), lws frees the vhd + * yet keeps serving the protocol here: refuse the upgrade + * instead of letting a connection in that can only be torn + * down half-established. + */ + if (!vhd) { + lwsl_wsi_err(wsi, "refusing conn: protocol init failed" + " on this vhost\n"); + return -1; + } return 0; case LWS_CALLBACK_ESTABLISHED: @@ -651,6 +672,14 @@ s_callback_ws(struct lws *wsi, enum lws_callback_reasons reason, void *user, /* remove pss from vhd->builders (active connection list) */ lws_dll2_remove(&pss->same); + /* + * On a vhost where protocol init failed there is no vhd and + * the conn was never established as a builder: there is + * nothing vhd-relative to tear down. + */ + if (!vhd) + break; + sais_builder_disconnected(vhd, wsi); sais_resource_wellknown_remove_pss(&pss->vhd->server, pss); @@ -674,6 +703,9 @@ s_callback_ws(struct lws *wsi, enum lws_callback_reasons reason, void *user, case LWS_CALLBACK_RECEIVE: + if (!vhd) + return -1; + pss->wsi = wsi; ssf = (lws_is_first_fragment(wsi) ? LWSSS_FLAG_SOM : 0) | (lws_is_final_fragment(wsi) ? LWSSS_FLAG_EOM : 0); diff --git a/src/web/w-comms.c b/src/web/w-comms.c index 55a3ef1..06fb559 100644 --- a/src/web/w-comms.c +++ b/src/web/w-comms.c @@ -378,12 +378,19 @@ w_callback_ws(struct lws *wsi, enum lws_callback_reasons reason, void *user, return 1; } - r = lws_ss_client_connect(vhd->h_ss_websrv) ? -1 : 0; - - if (r) - lwsl_wsi_err(wsi, "client connect for web -> srv failed"); + /* + * The streamtype is nailed_up, so lws_ss_create() above + * already tried the connection and owns retrying it via + * the ss backoff policy... a synchronous failure there (eg, + * sai-server not restarted yet) is normal startup racing. + * Don't call lws_ss_client_connect() from init and don't + * fail init over the link: returning nonzero would make + * lws free the vhd but keep serving this vhost, leaving + * browsers on a dead protocol until the next sai-web + * restart. + */ - return r; + return 0; case LWS_CALLBACK_PROTOCOL_DESTROY: saiw_event_db_close_all_now(vhd);
Page fetched 0s ago, creation time: 3ms (vhost etag hits: 0%, cache hits: 0%)