An infrastructure and operations post-mortem: what happened when the session crashed, why the background jobs died with it, and how the fix keeps them running independently.
What happened (step by step, like a detective story)
Morning, 07:01
I started two training jobs in parallel through my parallel_trains.sh script,
which was supposed to keep them running even if my session died.
Morning, 07:08
Both jobs completed successfully (about 470 seconds each). Model files were saved
as model.zip.
Morning, 07:08–07:21
The evaluation pass ran across several test maps and episode lengths. It saved
the usual metrics_*.json and action_dist.json output files.
CRASH 1 — somewhere after 07:21
My session died. It did not get to write:
- the final experiment report
- the time-tracking close-out
- the entry in
SESSION_LOG.mddescribing what was done
Day, 13:28
A different session of mine (apparently restarted by you or orch) launched four more training jobs. Their PIDs were 638968-638992.
CRASH 2 — almost immediately after 13:28
That session also died, very quickly. All four jobs died at the moment of start. The logs show only the config got printed, then nothing — they never reached model loading. The output directories were left empty.
Now, 15:27
I (a new session) saw the picture:
- six open “START” entries without a matching “FINISH” in time-tracking
- two experiments with finished models but no reports
- four experiments left as empty folders
Why this happened (root cause)
Superficially
In the script I used nohup ... & — the typical “start and forget.” This was
supposed to protect processes from death when the terminal closes.
Deeper
The morning jobs actually lived normally — nohup worked, and the session died
only after they had finished. The problem there was simply that I did not get
time to write the reports before the crash.
The afternoon jobs died together with the session, seconds after start. The
reason: nohup blocks the hang-up signal (SIGHUP, sent when a terminal closes
cleanly) but does not block SIGKILL — what the system sends when it tears a
process down abruptly. And a crashing session dies abruptly, so the kernel kills
all the children in the same process group before they ever detach.
This is the classic shape of a cgroup out-of-memory event: when the session’s
memory cgroup hits its memory.max limit, the kernel SIGKILLs processes in
that cgroup. Anything launched as a child of the crashing shell, and still in its
process group, goes down with it. You can confirm the diagnosis after the fact
with journalctl --user (or the system journal), which records the OOM kill and
the cgroup it fired in.
Analogy
Imagine you started four kettles on one stove and left the house. If the stove is turned off correctly, the kettles stay powered, all OK. If a car crashes into the house and the stove shatters, all four kettles shatter with it. To survive the accident, the kettles need to be on a separate stove in another room.
setsid (what I added) is that “separate stove in another room.” It creates a
new process session, independent of the parent shell, so even if the parent dies
instantly the kernel has no link tying the child processes to the gone shell.
What I fixed
1. scripts/parallel_trains.sh
Replaced:
nohup python train.py ... &
with:
setsid nohup env ... python train.py ... < /dev/null &
disown $!
Decoding:
setsid— creates a new process session (the separate stove)nohup— ignore the “terminal closed” signal< /dev/null— disable stdin, just in casedisown— bash forgets it started this process, so it won’t try to kill it on exit
Plus I added a recover command: at the start of a new session I can run
./scripts/parallel_trains.sh recover and see which jobs got orphaned after a
crash.
2. experiments/$TASK/.train_meta.json
Each launch now writes a meta-file with the PID, start time, and config. If the session dies, the next one sees this file and understands that something was launching here and needs checking.
3. CLAUDE.md (my work rules)
- Added a mandatory item to the start ritual: check for orphaned processes and unclosed STARTs.
- Added a hard rule: backgrounded jobs only through
setsid + nohup + disown. - Added a rule: don’t start new jobs while old unclosed ones still exist.
4. Memory feedback
Saved a separate reminder explaining why this is critical — so in future sessions
I know not to do nohup ... & without setsid.
What valuable was saved
The morning experiments finished cleanly, so their trained models and evaluation metrics were preserved on disk despite the missing reports. One of the two runs came out clearly ahead of the others on the shorter-horizon coverage runs, while the second confirmed that pushing the setting in the opposite direction hurts performance. The takeaway from the surviving data was useful, and nothing of the finished work was lost to the crash.
What’s NOT done
The four afternoon experiments were left as empty directories. Their configs are ready, and they will be launched again via the fixed script as the next task.
Glossary (terms from the text)
- Session crash — the controlling process dies (memory pressure, OOM, parent kill, and so on). From your perspective, I simply disappear.
- nohup — a Unix utility that protects a process from the “terminal closed”
signal (
SIGHUP). It does not protect againstSIGKILL. - setsid — a Unix utility that creates a new process session, independent of the parent shell. This makes a process truly detached.
- SIGHUP / SIGKILL — Unix signals.
SIGHUPmeans “hung up” and can be ignored;SIGKILLmeans “kill immediately” and cannot be ignored, since the kernel enforces it. - process group / session — the Unix process hierarchy. If a parent shell dies, the kernel may kill the whole group with it.
- cgroup
memory.max— the hard memory ceiling for a control group. When a cgroup’s usage hits this limit, the kernel triggers an OOM kill inside it. - OOM (out of memory) — the condition where the kernel must free memory by
killing processes. The kill, and the cgroup it happened in, show up in
journalctl.