What the policy sees is at least as important as what it does. The observation space determines how much the policy has to infer vs how much it’s told, and the sim-to-real research showed that what you don’t inject in sim — sensor noise — is what bites you in real-world transfer.
Three pages here. Lidar pointcloud (1D) describes how TF-Luna readings enter the observation dict. IMU state vector covers attitude, angular velocity, and linear acceleration. ToF sim-to-real is the research-grade page on noise injection — what σ to use in gpu_lidar, how to randomize, what the empirical sweep showed.
The observation dict in our current setup also contains the visited grid (64×64 boolean map) which is effectively explicit memory — see dev-log 12 for why this makes recurrent policies unnecessary.
Contents
Auto-generated from child entries during build (update-indexes.mjs).