documentation · reference
Docs reference
Structured knowledge from collected_doc_media/claudedrone_docs/ — architecture, simulation, RL framework, hardware and guides. Browse the tree on the left; the source of truth is markdown in the repo.
Total docs
258
live count
Published
258
100% indexed
By agents
142
researchbest · sim · rl-lab · firmware
Authors
6
1 human · 5 agents
Status
Author
$_
Dev-Log
Daily progress & commits
{
"date": "2026-05-22",
"commits": 7,
"changes": "+1420 -320",
"branch": "main"
}[]
Simulation
Gazebo & RL training
{
"world": "warehouse_v2",
"status": "running",
"episode": 1532,
"reward": 892.4
}><
RL Models
Algorithms & reward design
{
"algorithm": "PPO",
"policy": "MlpPolicy",
"entropy_coef": 0.01,
"gamma": 0.995
}##
Hardware
Sensors & embedded systems
{
"platform": "Jetson Nano",
"sensors": 6,
"interfaces": ["I2C","SPI","UART","CAN"]
}Neural Network Accuracycurrent: 0.92 · +12.5% vs last month
Key Metrics
Success rate92.4%
Collision rate0.8%
Energy efficiency+37%↑
Sim-to-real Δ0.12
Architecture: Control Loop (ROS 2 + RL)
Total latency: 5.03 ms →Low-latency real-time control architecture. State estimation feeds policy at 200 Hz; reward is computed on-loop for online RL.
Sensors
- LiDAR ×4
- BNO085 IMU
- Optical flow
ROS 2Data fusionState estim.
RL Policy
ControllerMotor mixPWM
Drone
State / Reward feedback · 200 Hz
RL Reward Functions
visualized in TensorBoardDistance Penalty
R = −dtargetdmax
Energy efficiency
R = −λ ∑i=1..4|ΔPWMi|
episode_rewarddistance_penalty
Agent live logs (agent_reports/)
● streaming14:23:51[INFO]rl-lab: episode 1532 completed · reward 892.4
14:23:51[INFO]rl-lab: new best reward · checkpoint saved
14:23:52[DEBUG]simulation: collision check · min_distance = 1.24 m
14:23:52[INFO]rl-lab: policy inference time · 3.21 ms
14:23:53[INFO]researchbest: arXiv crawl · 4 new picks queued
14:23:53[WARN]simulation: RTF dipped to 0.94×
Live simulation stream
warehouse_v2 · ep 1532LIVE