add_doctor_register_step and add_doctor_teardown_step add the
two halves of the deploy-doctor watch list to a slurmworkflow
workflow. Together they stop the shared deploy doctor
(hpc_doctor_script) automatically once the last campaign it is
watching has finished, so it no longer runs to its walltime or needs a manual
scancel.
Arguments
- wf
a
slurmworkflowworkflow summary (e.g. fromslurmworkflow::create_workflow).- wf_name
the campaign name, used as the marker file name. It must be identical in the register and teardown calls and unique per concurrent campaign; the workflow name is the natural choice.
- watch_dir
the watch-list directory, relative to the HPC repository root (the working directory of a workflow step job). Defaults to
data/run/.doctor_watch, which is gitignored in the standard EpiModel project layout.- sbatch_opts
a named list of sbatch options, merged over the step defaults (1 CPU, 5 minutes, 1G, mail on FAIL).
- job_name
the SLURM job name of the deploy doctor to stop. Defaults to
deploy_doctor, matching the#SBATCH --job-namedefault in the bundleddeploy_doctor.sh. If you rename the doctor's job at submit time to scope it to one project, pass the same name here or the teardown will not find it.
Details
The deploy doctor is one standalone SLURM job, launched once per deploy and
shared across every concurrent campaign whose job name matches its
PATTERN. It therefore must not be stopped when any single workflow
finishes, only when the last one does. The watch list is a directory holding
one empty marker file per live campaign. Add the register step near the start
of a workflow (right after the workflow is created) and the teardown step as
the final step:
the register step creates the marker for this campaign;
the teardown step removes this campaign's marker and, only if the directory is then empty, stops the doctor.
Pass the same directory to the doctor itself as WATCH_DIR when you
submit it. The teardown step is the primary stop signal and is immediate, but
the doctor also self-terminates after a run of empty sweeps, and WATCH_DIR
is what tells it which kind of empty queue it is looking at. With a campaign
still registered it waits the full IDLE_EXIT, treating an empty queue as
a lull between steps; with nothing registered it exits after the shorter
IDLE_FAST. The watch list only ever makes it more patient, and
IDLE_EXIT still caps the wait, so a stale marker left by a workflow that
died before its teardown ran cannot hold the doctor open to its walltime.
Because each campaign removes its own marker before testing emptiness,
whichever campaign finishes last always observes an empty directory. No
concurrent sibling is ever left unmonitored, and there is no interleaving in
which two simultaneous finishes both leave the doctor running. slurmworkflow
chains steps with afterany, so the teardown runs even when an earlier
step fails. Register and teardown are a matched pair: a workflow that registers
but never tears down leaves a stale marker that blocks teardown for other
campaigns, and a teardown with no matching doctor is a harmless no-op.
See also
hpc_doctor_script for the doctor itself.