Spring 2026 · Research prototype
VISION
CART.
A camera. A cart.
A learned way forward.
An autonomous mobile cart designed to lead a person through an indoor space using rear-facing vision and behavior cloning.
Explore the code ↗
- My role
- Hardware lead & embedded systems
- Team
- 3 people · Spring 2026
- Built with
- Raspberry Pi · Python · YOLO / ArUco
- Outcome
- Integrated prototype · 5 Hz inference
From hardware to behavior
The system in action.

Live perception and wheel-command telemetry from the prototype.

Raspberry Pi, camera, servo-driven chassis and separated power rails.

Camera streaming, manual driving, diagnostics and recording in one dashboard.

ArUco observations provide geometric measurements for the controller.

Recorded comparison of behavior-cloned and AWR policy responses.
Constraints & decisions
A Raspberry Pi 3B with 1 GB RAM and no GPU had to run perception, browser diagnostics, and wheel control. Camera-only sensing kept cost, power, and integration effort manageable for the low-speed indoor prototype. LiDAR and radar were considered, but did not address a V1 failure mode that justified their added complexity.
Camera placement traded reduced chassis occlusion against keeping a nearby person large in the frame. The 640 × 480 calibration gave a 65.9° horizontal field of view and 0.34 px reprojection error; placement was then evaluated physically.
Failure → redesign
The initial shared supply path reset the Pi when the servos started, turned, or reversed. I separated the 5 V compute rail from dual approximately 6 V servo rails, retaining a common ground through WAGO distribution. The final platform operated stably after this change.
My debugging work also covered PWM jitter, overheating, camera issues, and repeated Pi failures. The dashboard supported camera streaming, manual control, diagnostics, and data collection.
Control architecture
The backward-facing camera estimates distance and bearing. A compact behavior-cloned policy maps that state to left/right wheel commands. Perception, velocity estimation, deadband logic, and wheel actuation remain geometric or rule-based stages; only the policy was learned.
Recorded results
The documented data collection contains 6,748 frames over 449.5 seconds, with 93.5% ArUco lock. Training used 1,296 balanced rows and mirror augmentation for 2,592 samples. Policy inference ran at 5 Hz on the Pi.
These figures describe the collected dataset, calibration, and prototype inference. They do not establish reliable operation across unfamiliar environments.
Limitations & next tests
This is a research prototype for controlled indoor use. The next measurements are servo current, rail voltage, Pi temperature, wheel speed, command-to-motion latency, and tracking loss versus distance and lighting. Those measurements would determine whether to prioritize sensing, state estimation, or hardware reliability.
Explore more of the work.
← Back to Selected Builds