This section is where claims become numbers. The business cases describe what we’re trying to do; the architecture and hardware sections describe how the system is wired; this section answers does it actually work better than what’s already out there. The format is comparison-driven — RL policy vs classical baseline on the same map, with the same hardware parameters, in controlled conditions — and the numbers are what they are.
The current flagship is RL vs lawnmower, which compares the PPO policy against the classical A* + Frontier exploration baseline on indoor coverage. The headline numbers (coverage, time-to-95%, energy per square meter) are honest about where RL wins and where it doesn’t — there’s at least one map where the classical baseline still beats us, and we say so. Future benchmarks will cover sensor-failure robustness, sim-to-real transfer accuracy, and multi-agent coordination — each gets its own page when it’s ready.
A note on methodology: every benchmark here is reproducible from the dev-log entry that generated it. The numbers in rl-vs-lawnmower come from TASK-RL-EXP-7, whose run logs, configurations, and raw output are in /rl/dev_log. If a number in this section looks suspicious, it can be re-derived from the original experiment. No marketing-shaped numbers; if a result didn’t survive a re-run, it doesn’t get published here.
Contents
Auto-generated from child entries during build (update-indexes.mjs).