Project homepage Mailing List  Warmcat.com  API Docs  Github Mirror 
    npro  
 Modern all-safe Rust Network Protocol library supporting h1, h2, h3, ws, wt sans-IO and with socket IO + tls
git clone https://npro.rs/repo/npro
 
root / READMEs / README-watchers.md
Author[]Andy Green <andy@warmcat.com> 2026-10-05 10:40 UTC
Committer[]Andy Green <andy@warmcat.com> 2026-10-05 10:40 UTC
Treea3c3cd982eab18cc5c4b7bf59ac4d76625e1d19d   Raw Patch
 
server: a started task's later steps stay with its builder
server: a started task's later steps stay with its builder

When a builder refused an offered step as busy (eg, not enough disk free
yet, or all its instances in use) we unbound the task and made it
WAITING for any builder of the platform, keeping its build_step.  For a
task partway through, that let another builder, eg a second VM instance
of the platform, pick up the next step and run it in a fresh job dir
with no src/ tree:

  cd: /home/sai/jobs/DB3FC983/src: No such file or directory

Only unbind if nothing of the task has run yet.  Otherwise leave it
bound to the builder that has its job dir, in SAIES_STEP_SUCCESS, so it
is offered to that builder again.

And have the task allocation give a builder its own tasks waiting for
their next step before anything new, on any repo.  Their earlier steps'
work is sitting in its job dirs, and starting new tasks instead eats
the disk space they may be waiting for.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
diff --git a/src/server/s-task.c b/src/server/s-task.c index 38ba5d1..98490f1 100644 --- a/src/server/s-task.c +++ b/src/server/s-task.c @@ -484,10 +484,12 @@ bail: /* * Find a task that still needs doing for platform, on any event. * - * Each repo with something startable for the platform puts forward its newest - * event that has some, and the repos take turns: the one the platform was - * least recently given a task from goes first. So a big push on one repo - * doesn't hold up a push on another until every one of its tasks is done. + * A task the builder already started and that is waiting for its next step + * comes first. Otherwise, each repo with something startable for the + * platform puts forward its newest event that has some, and the repos take + * turns: the one the platform was least recently given a task from goes + * first. So a big push on one repo doesn't hold up a push on another until + * every one of its tasks is done. */ static const sai_task_t * sais_task_pending(struct vhd *vhd, struct pss *pss, sai_plat_t *cb, @@ -495,7 +497,7 @@ sais_task_pending(struct vhd *vhd, struct pss *pss, sai_plat_t *cb, { const sai_event_t *cand[SAIS_PENDING_EVENTS]; lws_usec_t cand_served[SAIS_PENDING_EVENTS]; - char esc_plat[96], esc_bname[128], pf[256], query[384]; + char esc_plat[96], esc_bname[128], pf[512], query[384]; struct lwsac *ac = NULL; unsigned int pending_count; int n, m, i, cands = 0; @@ -538,18 +540,61 @@ sais_task_pending(struct vhd *vhd, struct pss *pss, sai_plat_t *cb, lws_start_foreach_dll(struct lws_dll2 *, p, o.head) { sai_event_t *e = lws_container_of(p, sai_event_t, list); + lws_dll2_owner_t owner; sqlite3 *pdb = NULL; + if (sai_event_db_ensure_open(vhd->context, &vhd->sqlite3_cache, + vhd->sqlite3_path_lhs, e->uuid, 0, &pdb)) + continue; + + /* + * A task this builder already started, waiting for its next + * step, comes before anything new whatever repo it is from. + * Its earlier steps' work is sitting in our job dir, which is + * only safe from being reclaimed for a while, and starting new + * tasks instead just eats more of the disk it is waiting for. + */ + + lws_snprintf(pf, sizeof(pf), + " and state=%d and idle=0 and platform='%s' and " + "builder_name='%s' and run = (select max(run) " + "from tasks t2 where tasks.uuid = t2.uuid)", + SAIES_STEP_SUCCESS, esc_plat, esc_bname); + + lwsac_free(&pss->ac_alloc_task); + lws_dll2_owner_clear(&owner); + n = lws_struct_sq3_deserialize(pdb, pf, "uid asc ", + lsm_schema_sq3_map_task, + &owner, &pss->ac_alloc_task, 0, 1); + if (n >= 0 && owner.count) { + sai_task_t *t = lws_container_of(owner.head, + sai_task_t, list); + + if (sais_check_and_fix_stale_task(pdb, t)) { + sai_event_db_close(&vhd->sqlite3_cache, &pdb); + goto bail; + } + + lwsl_info("%s: plat %s: continuing %s\n", __func__, + platform, t->uuid); + + memcpy(&pss->alloc_task, t, sizeof(pss->alloc_task)); + lws_strncpy(pss->alloc_repo, e->repo_name, + sizeof(pss->alloc_repo)); + sai_event_db_close(&vhd->sqlite3_cache, &pdb); + lwsac_free(&ac); + + return &pss->alloc_task; + } + for (i = 0; i < cands; i++) if (!strcmp(cand[i]->repo_name, e->repo_name)) break; - if (i != cands) + if (i != cands) { /* this repo already has a newer candidate */ + sai_event_db_close(&vhd->sqlite3_cache, &pdb); continue; - - if (sai_event_db_ensure_open(vhd->context, &vhd->sqlite3_cache, - vhd->sqlite3_path_lhs, e->uuid, 0, &pdb)) - continue; + } /* * Find out how many tasks in startable state for this platform, diff --git a/src/server/s-ws-builder.c b/src/server/s-ws-builder.c index 190c363..a4ba8fb 100644 --- a/src/server/s-ws-builder.c +++ b/src/server/s-ws-builder.c @@ -787,8 +787,31 @@ sais_process_rej(struct vhd *vhd, struct pss *pss, lwsl_notice("%s: SAI_TASK_REASON_BUSY: Set busy: %s\n", __func__, rej->task_uuid); do_remove_uuid = 1; - sais_bind_task_to_builder(vhd, NULL, NULL, rej->task_uuid); - sais_set_task_state(vhd, rej->task_uuid, SAIES_WAITING, 0, 0); + + sai_task_uuid_to_event_uuid(event_uuid, rej->task_uuid); + if (!sai_event_db_ensure_open(vhd->context, &vhd->sqlite3_cache, + vhd->sqlite3_path_lhs, event_uuid, 0, &pdb)) { + sais_rej_is_stale(pdb, rej, &build_step); + sai_event_db_close(&vhd->sqlite3_cache, &pdb); + } + + if (build_step > 0) { + /* + * It refused a later step of a task it already + * started. The earlier steps' work is in its job + * dir, so the task can only go on there: leave it + * bound, waiting for its next step. Making it WAITING + * for anyone let another builder of the platform run + * the next step in a job dir with no src/ tree. + */ + sais_set_task_state(vhd, rej->task_uuid, + SAIES_STEP_SUCCESS, 0, 0); + } else { + sais_bind_task_to_builder(vhd, NULL, NULL, + rej->task_uuid); + sais_set_task_state(vhd, rej->task_uuid, + SAIES_WAITING, 0, 0); + } sais_plat_busy(sp, 1); break;
Page fetched 0s ago, creation time: 12ms (vhost etag hits: 0%, cache hits: 0%)