-- Views
August 26, 26
スライド概要
東北大学 大学院工学研究科 ロボティクス専攻
Safe and Efficient Navigation Considering Texting Pedestrians via Attention-Aware Cost-map Tuning Shunya Tadano, Yusuke Tamura, Ankit A. Ravankar, Yasuhisa Hirata
Research Overview Motivation Robots should consider pedestrian attention Distracted Attentive ⚫ Trajectory Prediction ⚫ Attention classification ⚫ Adaptive costmap tuning ⚫ Evaluation with 12 participants across 252 trials 1
Background Social robot navigation in shared human environments Safety to avoid collision Legibility to make robot acceptable Robots need to adapt to surrounding pedestrians 2
Background Distracted pedestrians are hard for robots to handle ⚫ Smartphone use can change gait and slow responses [Sajewicz+, 2023; Murakami+, 2021] ⚫ 12 - 45% of pedestrians distracted [Simmons+, 2020] Robots should understand not only pedestrian motion, but also pedestrian attention 3
Related Works Existing approaches do not fully connect attention recognition to navigation Trajectory Prediction Gaze / Head Pose [Gupta+, 2018] [Tu+, 2025] Attention state ignored Requires facial detail Smartphone-Use Recognition [Wu+, 2020] Limited attention awareness Attention recognition → Attention-aware cost-map tuning 4
Objective Make robot navigation responsive to pedestrian attention This work proposes: ⚫ Attention classification using 3D skeletal features ⚫ Adaptive costmap tuning based on attention state ⚫ Evaluation with 12 participants across 252 trials Distracted Attentive 5
Method Overview Data Point Cloud Adaptive CostMap Generation 3D Pose Estimation Position Image Depth Trajectory Prediction Future Trajectory Skelton Point Environment Map Attention Classification CostMap Motion Planning Cost & Margin Assignment where to avoid how cautiously to avoid 6
Method: Trajectory Prediction Trajectory Prediction Future Trajectory ● Human Scene Transformer predicts future pedestrian trajectories [Salzmann+, 2023] ● Input: past pedestrian pos & 3D skeleton (2 s / 6 steps) ● Output: future trajectory distribution (4 s / 12 steps ) 7
Method: Attention Classification Attention Classification Cost & Margin Assignment ● 3D skeletal features are extracted from RGB images [Bazarevsky+, 2020] ● A lightweight Random Forest enables real-time processing 8
Classifier Method: Skeletal Feature Extraction Walking 3D pose est. Phone classification ⚫ 3D keypoints 𝒑 = {𝑝1 , … , 𝑝33 } MediaPipe BlazePose ⚫ Skeleton point Normalization with the hip (𝑝23 , 𝑝24 ) 𝒑′𝒊 = 𝒑𝒊 − (𝒑𝟐𝟒 + 𝒑𝟐𝟒 ) / 2 9
Attention Classification Results ● Classification accuracy: 95.1% ● Wrist and upper-body keypoints contributed most ● Consistent with characteristic smartphone-use postures 10
Method: Adaptive Cost-Map Tuning Adaptive CostMap Generation Environment Map CostMap ⚫ Add predicted pedestrian trajectories to the navigation cost map [Lu+, 2014; Rösmann+, 2012] ⚫ The robot changes its avoidance behavior according to pedestrian attention 11
Method: Attention-Aware Cost-Map Tuning Trajectory predictions → map as navigation costs Cost Low High ⚫ Traj. Prediction ⚫ Dynamic Obstacles ⚫ Static Obstacles Attention Classification ⚫ Attentive: safety-margin radius = 0.3 m ⚫ Distracted: safety-margin radius = 0.5 m 12
Experimental Setup ⚫ T-shaped indoor corridor ⚫ 12 participants, aged 22–28 ⚫ 21 trials per participants (7 trials × 3 methods) Straight Scenario Intersection Scenario 13
Four Pedestrian Behaviors Normal Texting → Normal Texting Normal → Texting 14
Experimental Setup Method LiDAR-only Prediction Attention ✗ ✗ ⚫ Dynamic Obstacles ⚫ Static Obstacles 15
Experimental Setup Method Prediction Attention LiDAR-only ✗ ✗ Traj. Prediction ✓ ✗ ⚫ Traj. Prediction ⚫ Dynamic Obstacles ⚫ Static Obstacles 16
Experimental Setup Method Prediction Attention LiDAR-only ✗ ✗ Traj. Prediction ✓ ✗ Attention-Aware (Ours) ✓ ✓ ⚫ Traj. Prediction ⚫ Dynamic Obstacles ⚫ Static Obstacles Attention Classification ⚫ Attentive: safety-margin radius = 0.3 m ⚫ Distracted: safety-margin radius = 0.5 m 17
Quantitative Results: Motion Smoothness Metric LiDAR-only Jerk RMS (m/s³) 37.45±17.46 Traj. Prediction Attention-Aware (Ours) 11.01±8.62 11.42±9.19 ● Prediction-based methods significantly reduced jerk RMS compared with the LiDAR-only (p < 0.001) ● No significant difference in jerk RMS between Traj. Prediction and Ours Trajectory prediction improves motion smoothness; attention-aware tuning does not further reduce jerk. 18
Quantitative Results Metric LiDAR-only Traj. Prediction Attention-Aware (Ours) Min distance (m) 1.74±0.68 0.94±0.28 1.01±0.31 Max lateral. dev. (m) 0.30±0.21 0.31±0.19 0.36±0.23 Travel time (s) 16.8±2.1 17.0±1.9 16.9±2.0 ⚫ Ours produced larger lateral deviations than Prediction ⚫ 8 of 12 participants noticed the behavioral difference while texting (p = 0.050) 19
Results: Subjective Evaluation Item () LiDAR-only Traj. Prediction Attention-Aware (Ours) Q1 Safety 4.6 4.69 5.07 Q2 Comfort 4.24 4.42 4.82 Q3 Smoothness 3.96 4.14 4.89 Q4 Naturalness 4.13 4.22 4.67 Q5 Predictability 3.69 4.01 4.48 Q6 Trust 4.49 4.64 4.96 Q7 Efficiency 4.38 4.57 4.93 Our method achieved the highest score at all metrics 20
Results: Subjective Evaluation * 𝑝 < .05 ** 𝑝 > .01 Significant improvement in Smoothness / Predictability 21
Takeaway Motivation ⚫ Robots should also consider pedestrian attention Method ⚫ 3D skeletal attention classification with adaptive cost-map tuning Key Finding ⚫ It improved perceived smoothness and predictability without increasing travel time Future Work ⚫ Temporal attention modeling ⚫ More complex scenarios ⚫ 360° perception ⚫ Broader participant population 22