All of this work has been happening on my workbench, in the dark, narrated to nobody. So I did the obvious thing and turned it into a show. My robot now has its own video series on YouTube, called ROBOTICS ADVENTURES, and the whole back catalog is up.
What an episode looks like
Each one is cut from three cameras at once. His own eye, with the detection boxes drawn on it so you see exactly what he is locking onto, sits next to two outside cameras watching him from the room. His voice runs underneath, calling out what he sees. A title card opens it, credits close it, and there is a little music bed under the whole thing.
https://www.youtube.com/watch?v=FIexRfdnJ7k
The part that was genuinely hard
Getting his narration to land at the right moment nearly beat me. When he speaks live, there is a real delay between the instant he recognizes something and the instant the words actually come out of the speaker, because the recognition, the language model, and the voice synthesis all take time. If you sync the audio to when it played, every call-out lands late, and in a multi-camera cut it looks like he is talking about something before he looks at it.
The fix was to stamp the exact millisecond his head settles on each object during the scan, and anchor the clean voice clip to that recognition moment instead of to playback. When two objects share one glance, the call-outs play back to back and his head does not turn until he has named both. It is a small thing that took a long time, and it is the difference between the cut feeling alive and feeling broken.
The catalog
The series runs from the early days of him learning to walk and throw, through him falling and getting back up and renaming himself, through learning to see, all the way to the show-and-learn episodes where he recognizes a room from memory. Twelve episodes, in order, in one place.
Watch the full ROBOTICS ADVENTURES playlist.
If you have been reading these logs, the videos are the same story with the sound turned on. Go watch him work.