claudeDroneteam-docs
documentation · reference
Docs reference

Structured knowledge from collected_doc_media/claudedrone_docs/. Browse the tree on the left; the source of truth is markdown in the repo.

Session crash post-mortem — how background jobs survive a setsid fix

An ops post-mortem: why a crashed session killed its background jobs, OOM diagnosis, and the setsid fix that keeps them alive.

stablerl-labupdated 2026-05-11T00:00:00.000ZClaudeDroneRLDevLog

An infrastructure and operations post-mortem: what happened when the session crashed, why the background jobs died with it, and how the fix keeps them running independently.

What happened (step by step, like a detective story)

Morning, 07:01

I started two training jobs in parallel through my parallel_trains.sh script, which was supposed to keep them running even if my session died.

Morning, 07:08

Both jobs completed successfully (about 470 seconds each). Model files were saved as model.zip.

Morning, 07:08–07:21

The evaluation pass ran across several test maps and episode lengths. It saved the usual metrics_*.json and action_dist.json output files.

CRASH 1 — somewhere after 07:21

My session died. It did not get to write:

  • the final experiment report
  • the time-tracking close-out
  • the entry in SESSION_LOG.md describing what was done

Day, 13:28

A different session of mine (apparently restarted by you or orch) launched four more training jobs. Their PIDs were 638968-638992.

CRASH 2 — almost immediately after 13:28

That session also died, very quickly. All four jobs died at the moment of start. The logs show only the config got printed, then nothing — they never reached model loading. The output directories were left empty.

Now, 15:27

I (a new session) saw the picture:

  • six open “START” entries without a matching “FINISH” in time-tracking
  • two experiments with finished models but no reports
  • four experiments left as empty folders

Why this happened (root cause)

Superficially

In the script I used nohup ... & — the typical “start and forget.” This was supposed to protect processes from death when the terminal closes.

Deeper

The morning jobs actually lived normally — nohup worked, and the session died only after they had finished. The problem there was simply that I did not get time to write the reports before the crash.

The afternoon jobs died together with the session, seconds after start. The reason: nohup blocks the hang-up signal (SIGHUP, sent when a terminal closes cleanly) but does not block SIGKILL — what the system sends when it tears a process down abruptly. And a crashing session dies abruptly, so the kernel kills all the children in the same process group before they ever detach.

This is the classic shape of a cgroup out-of-memory event: when the session’s memory cgroup hits its memory.max limit, the kernel SIGKILLs processes in that cgroup. Anything launched as a child of the crashing shell, and still in its process group, goes down with it. You can confirm the diagnosis after the fact with journalctl --user (or the system journal), which records the OOM kill and the cgroup it fired in.

Analogy

Imagine you started four kettles on one stove and left the house. If the stove is turned off correctly, the kettles stay powered, all OK. If a car crashes into the house and the stove shatters, all four kettles shatter with it. To survive the accident, the kettles need to be on a separate stove in another room.

setsid (what I added) is that “separate stove in another room.” It creates a new process session, independent of the parent shell, so even if the parent dies instantly the kernel has no link tying the child processes to the gone shell.

What I fixed

1. scripts/parallel_trains.sh

Replaced:

nohup python train.py ... &

with:

setsid nohup env ... python train.py ... < /dev/null &
disown $!

Decoding:

  • setsid — creates a new process session (the separate stove)
  • nohup — ignore the “terminal closed” signal
  • < /dev/null — disable stdin, just in case
  • disown — bash forgets it started this process, so it won’t try to kill it on exit

Plus I added a recover command: at the start of a new session I can run ./scripts/parallel_trains.sh recover and see which jobs got orphaned after a crash.

2. experiments/$TASK/.train_meta.json

Each launch now writes a meta-file with the PID, start time, and config. If the session dies, the next one sees this file and understands that something was launching here and needs checking.

3. CLAUDE.md (my work rules)

  • Added a mandatory item to the start ritual: check for orphaned processes and unclosed STARTs.
  • Added a hard rule: backgrounded jobs only through setsid + nohup + disown.
  • Added a rule: don’t start new jobs while old unclosed ones still exist.

4. Memory feedback

Saved a separate reminder explaining why this is critical — so in future sessions I know not to do nohup ... & without setsid.

What valuable was saved

The morning experiments finished cleanly, so their trained models and evaluation metrics were preserved on disk despite the missing reports. One of the two runs came out clearly ahead of the others on the shorter-horizon coverage runs, while the second confirmed that pushing the setting in the opposite direction hurts performance. The takeaway from the surviving data was useful, and nothing of the finished work was lost to the crash.

What’s NOT done

The four afternoon experiments were left as empty directories. Their configs are ready, and they will be launched again via the fixed script as the next task.

Glossary (terms from the text)

  • Session crash — the controlling process dies (memory pressure, OOM, parent kill, and so on). From your perspective, I simply disappear.
  • nohup — a Unix utility that protects a process from the “terminal closed” signal (SIGHUP). It does not protect against SIGKILL.
  • setsid — a Unix utility that creates a new process session, independent of the parent shell. This makes a process truly detached.
  • SIGHUP / SIGKILL — Unix signals. SIGHUP means “hung up” and can be ignored; SIGKILL means “kill immediately” and cannot be ignored, since the kernel enforces it.
  • process group / session — the Unix process hierarchy. If a parent shell dies, the kernel may kill the whole group with it.
  • cgroup memory.max — the hard memory ceiling for a control group. When a cgroup’s usage hits this limit, the kernel triggers an OOM kill inside it.
  • OOM (out of memory) — the condition where the kernel must free memory by killing processes. The kill, and the cgroup it happened in, show up in journalctl.
© 2026 claudeDrone Team · auto-pipeline · Nuxt 3 SSR