What this is: the first journal entry. Setting up the repo, a Python virtual environment, installing PyTorch with CUDA, and smoke-testing that everything actually works on the RTX 5070 (new Blackwell architecture). Why: an infrastructure task. Before this,
rl-labdidn’t exist. We needed the stack assembled before we could start experiments.
Date: 7 May 2026
Machine: D2 — Ubuntu, RTX 5070 (Blackwell, sm_120), <dev-machine-local-ip>
1. What we wanted
Stand up a full RL infrastructure in a fresh repo:
- Python 3.12 virtual environment.
- PyTorch with CUDA support for our GPU.
- Stable-Baselines3 (RL library).
- Gymnasium (environment standard).
- TensorBoard (visualization).
The main risk: the RTX 5070 is a new Blackwell architecture (compute capability sm_120). PyTorch wheels may not include code for this architecture → a CUDA “no kernel image” error at runtime.
2. What we did
Repo structure
git init -b dev ~/git/proj/rl-lab/
mkdir -p envs training/configs models maps scripts docs/dev-log experiments checkpoints runs
In .gitignore: venv/, checkpoints/, runs/, *.pt, *.zip, .env_rl. Anything large or sensitive stays out of version control.
Python environment
python3.12 -m venv ~/git/proj/rl-lab/venv
~/git/proj/rl-lab/venv/bin/pip install --upgrade pip
PyTorch with CUDA 12.8
Critical for Blackwell sm_120: you need cu128 wheels (CUDA 12.8). Older cu124 / cu121 builds may not contain code for sm_120, which surfaces as a CUDA error at runtime.
~/git/proj/rl-lab/venv/bin/pip install torch --index-url https://download.pytorch.org/whl/cu128
Problem: the download.pytorch.org channel was throttled to roughly 58 KB/s. The full torch install plus all CUDA dependencies (cudnn, cublas, cusparse, cusolver, nccl, nvjitlink, cuda-runtime, cuda-cupti, curand, cufft) took about 64 minutes.
pip without -v stays silent until the very end — no progress is visible. To avoid sitting and polling, I rigged a small watcher: until [ -d venv/.../torch ]; do sleep 5; done.
Smoke test — the most important step
import torch
print(torch.cuda.is_available()) # True
print(torch.cuda.get_device_capability(0)) # (12, 0) ← sm_120 OK
print(torch.version.cuda) # 12.8
# + a 2048x2048 matmul on cuda → OK
Important: is_available() == True on its own is not enough. If the PTX (low-level code) for sm_120 isn’t in the wheels, the function still returns True, but any actual kernel call fails. So the smoke test always includes real GPU computation — a matmul — not just a capability check.
For us everything was OK: sm_120 is supported in torch 2.11 + cu128.
Other dependencies
gymnasium 1.2.3
stable-baselines3 2.8.0
tensorboard 2.20.0
numpy 2.2.6
imageio 2.37.3
imageio-ffmpeg 0.6.0
matplotlib 3.10.9
pyyaml 6.0.3
tqdm 4.67.3
Important: NumPy 2.x breaks older Stable-Baselines3 (< 2.3). So in requirements.txt I pinned stable-baselines3>=2.3. We’re on 2.8.0, which works fine.
3. What didn’t work first time
- The slow torch install (~64 minutes) caused by the throttled
pytorch.orgchannel. Not a bug, but worth remembering — update torch in the background with a watcher rather than waiting on it.
4. Notes for the future
- The pip cache after bootstrap is 989 MB, so repeat installs into new venvs will be faster.
- For 16 parallel envs on the RTX 5070 we may need to raise
ulimiton open files. We’ll check that on the first training run. - For torch 2.11 cu128 on
sm_120, there’s no certainty yet that FlashAttention ortorch.compileis fully stable. When we move to bigger networks, verify those separately.
5. What we didn’t do
In this session there was no training and no experiments. By Aleks’s call, the bootstrap phase is infrastructure only. Training comes in the next ticket (TASK-010, see dev-log/02).
Related files
~/git/proj/rl-lab/requirements.txt~/git/proj/rl-lab/.gitignore~/git/proj/rl-lab/.env_rl(gitignored — holds absolute paths)
Glossary
- venv (virtual environment) — an isolated Python environment. Protects against library version conflicts.
- PyTorch — the main Python library for neural networks and tensor computation.
- CUDA — NVIDIA’s platform for GPU computation.
- cu128 / cu124 — CUDA build variants of PyTorch wheels. cu128 = CUDA 12.8.
- wheel (.whl) — a compiled Python package format.
- Blackwell — the newest NVIDIA GPU architecture (RTX 50xx series).
sm_120/ compute capability 12.0 — a GPU architecture identifier. RTX 5070 =sm_120.- PTX — CUDA low-level assembly. The code the GPU actually executes.
- “no kernel image” — the canonical CUDA error when wheels don’t include code for your architecture.
- NCCL — NVIDIA library for multi-GPU communication.
- cuDNN — NVIDIA library with fast implementations of neural-network operations.
- Stable-Baselines3 (SB3) — Python library with ready-made RL algorithm implementations.
- Gymnasium — the standard RL environment API (a fork of OpenAI Gym).
- TensorBoard — a training-metrics visualization tool.
- Smoke test — a short verification that everything works.
- D2 — our development machine with the RTX 5070.