Last time I introduced you to AgentAdam, the AI assistant living in my pocket. Today I want to tell you about the other member of the family: AiNex, my Hiwonder humanoid robot, two feet tall, twenty-four servos, one purple ball, and as of this week, a fully autonomous grab-and-throw. Not scripted. Not remote-controlled. He sees the ball, walks to it, lines himself up, decides for himself that everything measures right, picks it up, and throws it. Then he does a victory dance, because we are not savages.

This session was enormous. Hereās what we built.
Watch It
Donāt take my word for it, hereās the film. Attempt 36, the very next run after the flawless one: title card, the road-to-autonomy attempt log, all three synced views (two room cameras plus his own eyes with the detection overlay), his live narration and jokes, and the full credits roll. Everything you hear him say, he decided to say himself.
Attempt 36, start to finish: intro, approach, grab gate, throw, victory dance, credits.
He Learns What Things Look Like, By Example
Thereās no color threshold or hand-tuned blob detector under the hood. AiNex runs MediaPipe object detection on his Raspberry Pi, embeds every candidate object into a 1024-dimensional vector, and matches it against a Postgres + pgvector visual memory of things heās been taught. Show him the ball a dozen times and he knows it, by resemblance, like youād recognize your own coffee mug. Even better: when he gets a confident sighting, he quietly teaches himself that example too, so his memory tracks the lighting as the day changes. No training runs. No GPU. A $100 robot brain, learning by looking.
He Walks Like He Means It (Because Falling Hurts)
A two-foot humanoid is a pendulum with opinions. He walks in balanced segments, stride, settle, re-check, with his eyes locked on the ball the whole time, feeding corrections continuously. His turns are closed against his own gyroscope: he commands sixteen degrees, measures what actually happened, and learns his own glide per direction, because his right side and left side donāt behave the same (whose do?). The iron rule that emerged from all the bruises: the body never moves blind. If his eyes lose the ball, his feet stop. Every failure this week traced back to some corner of the code that violated that rule, and every fix was making the rule absolute.

The Gold Pose and the Grab Gate
Hereās my favorite engineering trick of the whole project. I physically stood him in the perfect grab stance, and we measured what his own senses read from there: head panned dead center, tilted fully down at 290 ticks, ball radius 76 pixels. That measured stance became the āgold poseā, his ground truth. Distance-from-size turned out to lie badly at close range, but his own neck angle doesnāt: if his head has to point at his toes to see the ball, heās there.
Before heās allowed to grab, three independent conditions must agree: head-down tilt says heās at the spot, visual memory confirms itās really the ball, and the fused center error, head bearing plus image offset, says heās square. Miss any one and he makes a correction and re-measures. He is not allowed to grab a guess. Twice this session he stood in a perfect stance and refused to grab because one number was two hundredths under the bar, infuriating and also exactly what I want from a robot: heād rather do nothing than do something dumb.

He Talks Now. He Has Jokes.
A local Kokoro TTS engine gives him a voice, everything on-device, nothing cloud. He announces every decision as he makes it (āTurning left twelve.ā āStepping right.ā āAt the checkpoint. Centering.ā), and a small local LLM gives him a personality: eager, health-aware, and relentlessly self-deprecating. He opens every attempt with a wave and an introduction, āAttempt number 35. I am AiNex. My therapist says the ball cannot hurt me. My therapist is a text file.ā, and if he repeats an action he swaps the announcement for commentary instead of parroting himself. When a throw lands he says, and I quote, āThatās why they built me,ā and does the twist. Every spoken line is journaled with its audio clip, so I have a replayable decision log of his entire run.
Attempt 35
And then it all came together. Intro wave. āAttempt number 35.ā He found the ball, walked in, hit the checkpoint, centered to a fused error of ten, docked himself forward, and stopped, at ball radius 77 pixels versus the hand-measured gold of 76, one pixel from the stance I had placed him in by hand. The grab gate passed on its first check: tilt 318, similarity 0.68, center error negative three. Deep squat, two-handed clamp, lift, overhead throw. Ball gone. Victory dance. Flawless, end to end, on his own.



I have watched the film an unreasonable number of times.
Mission Control
All of it feeds a one-screen dashboard on my Linux box: three live camera views (two external angles plus the robotās own annotated vision), a little robot figure with all twenty-four servos colored by temperature, a head-to-toe servo list with live positions and voltages, his brainās thought stream, and a voice toggle. Ankle servos run hot when he fights an unlevel stance, so we calibrated his standing lean and put thermal gates on every run, heāll refuse to start a marathon on cooked ankles.
Then I Gave Him a Mind of His Own
After the throws, the project turned inward: I wanted to talk to him, and I wanted him to actually know things about himself when he answered. So the newest layer is a full AI brain, and a Mission Control cockpit to watch it think.

The Conversation Layer
A local Llama 3 8B (served by Ollama on my desktop, nothing leaves the house) is his inner voice. When I type into the Comms panel, the prompt he receives isnāt just my message: itās stitched together with what his camera currently sees, the objects in his learned memory, and his live body state, battery voltage, roll and pitch, and his hottest servo temperature, pulled from his telemetry the moment I hit send. The persona contract is baked in: eager, honest about his limits, self-deprecating. His answers come back in character, get spoken aloud through his Kokoro voice, and if I actually ask him to do something, the same reply carries a skill decision and his body executes it.
The screenshot above caught a real exchange. I asked how his memory was holding up: āMy memory is working okay, I think! Iām still getting used to recognizing things.ā Then I asked if he wanted to learn to shoot a basket next, and this small plastic person, with two ankle servos at 64 degrees and a 92% full disk, replied: āIām not sure Iām ready for that level of coordination.ā Self-awareness: achieved.
A Memory With a Timeline
His visual memory got real bookkeeping. Every example he learns is timestamped in Postgres, so his entire education is queryable history, first memory July 17th, newest one minutes ago. Every time a recognition fires in the wild, a recall counter ticks. The Mind panel on the dashboard reads it live: at screenshot time he held 84 taught examples of the purple ball, had recalled them 258 times, and had self-learned 36 new examples that day, because whenever he gets a confident look at the ball, he quietly files the sighting back into memory. He is, in the most literal sense, studying while he works. The panel calls itself āMemory & Educationā and it is not being cute; that is genuinely what the table does.
The Cockpit
The rest of Mission Control watches everything else: a glowing schematic of his exact body where all 24 servos live as clickable nodes colored by temperature, quiet when healthy, and a warning list that names names when an ankle runs hot. His Raspberry Piās vitals stream in as an EKG-style pulse graph (CPU spikes when the vision pipeline leans in), with memory, that poor 85%-full swap, and disk gauges below. An activity console logs every decision, link event, servo warning, chat line, and each new thing he learns. And a webcam mode lets me control him with my own body, my fingertip steers his head, my arms mirror onto his, MediaPipe tracking me from the browser the same way he tracks the ball.
Every piece of it runs local: his Pi, my Linux box, a desktop GPU for the LLM. The whole stack could keep working if the internet disappeared. The robot, meanwhile, keeps learning either way.
The Film Crew Problem
Recording this properly turned into its own project: three video streams (two cameras plus his eye view) recorded raw with millisecond wall-clock stamps, combined afterward into one synced three-view film, with his clean TTS audio composed onto the timeline at the exact moment each line was spoken, instead of whatever the room mics caught. We even measure his vision pipelineās latency at record start (about half a second) so his eye view lines up with reality. Lesson learned the hard way: an MJPEG stream has no native timing, and if you forget to stamp it, your robotās point-of-view plays at 3.6Ć speed like a silent film.
The Complete Build Log
For the fellow nerds: everything heās actually running, layer by layer. All of it lives on the robotās Raspberry Pi and my Linux box ā nothing phones home to a cloud.
The Full Stack
| Layer | Tech | What it does for him |
|---|---|---|
| Body | Hiwonder AiNex Ā· Raspberry Pi 5 Ā· Docker + ROS Noetic | 24-servo humanoid; all control code runs in a container on his chest-mounted Pi |
| Object detection | MediaPipe EfficientDet-Lite0 | Proposes candidate objects in every camera frame, on-device |
| Embeddings | MediaPipe MobileNetV3 embedder | Turns each candidate into a 1024-d vector, his āwhat does this look likeā signature |
| Visual memory | PostgreSQL + pgvector | k-NN cosine recognition over taught examples; timestamped education history + recall stats |
| Lock-on tracking | OpenCV CSRT tracker | Follows the actual pixels at 4 Hz between recognitions so his head never loses the ball |
| Mind | Llama 3 8B via Ollama | Persona, chat, and action planning, prompted with what he sees and his live vitals |
| Voice | Kokoro TTS (kokoro-onnx) | Local speech for every decision; each line journaled with its audio clip |
| Balance & turns | Onboard IMU | Gyro-closed turns with learned glide, fall detection with auto get-up, stance self-leveling |
| Somato control | MediaPipe Tasks JS (hand + pose) | Browser webcam control: fingertip steers his head, my arms mirror onto his |
| Cameras | Chest cam + 2 external cams via MediaMTX | His eye (MJPEG with overlays) plus two room angles (RTSP/WebRTC) for the films |
| Mission Control | Hand-rolled HTML/JS + Python HTTP APIs | The cockpit: servo schematic, EKG vitals, memory panel, chat, activity console, no framework |
| Film pipeline | FFmpeg | Three streams stamped with wall-clock time, synced offline, clean TTS voice composited in |
| Infrastructure | Pi 5 + Linux desktop + GPU box | Every layer runs in the house, zero cloud, works with the internet unplugged |
The Visual Memory (pgvector + embeddings)
- MediaPipe EfficientDet-Lite0 proposes object candidates from his chest camera; every candidate is embedded on-device into a 1024-dimensional vector (MobileNetV3, L2-normalized).
- A Postgres database with pgvector stores embeddings of objects heās been shown. Recognition is nearest-neighbor cosine similarity ā the ball scores ~0.77-0.97, random clutter stays under 0.42.
- He teaches himself: any confident sighting (similarity > 0.72) is added back to memory, rate-limited, so his recollection of the ball tracks the roomās changing light. The single biggest vision bug all week was stale memory vs. current lighting ā auto-learning killed it.
- Close-range truth: a partially-cropped, frame-filling ball embeds differently than a distant one. His grab-range gate accepts 0.55+ because identity also rides track continuity ā the lock has followed the same pixels the whole walk.
The Gaze (how he keeps the lock)
- A 4 Hz fast-track loop owns his head: a CSRT pixel tracker follows the ball between recognitions, recognition re-confirms identity and re-aims it.
- Two-tier search when he blinks: quick local glances near the last known position first, then slow, level, side-to-side sweeps at three tilt tiers. He never starts a search staring at his own feet (a hard-won rule).
- Vestibulo-ocular reflex: when his body turns, the head counter-pans by the commanded amount so the ball stays in frame ā same trick your eyeballs pull.
The Legs (walking without dying)
- All locomotion goes through the proven joystick gait pathway ā the same parameters a human drives him with, because those are the ones that never fall.
- Side-steps must complete whole gait cycles: halting a strafe mid-cycle rocks him back and cancels the step (a bug that masqueraded as a reversed direction for a full day).
- Turns are closed against his gyroscope: command 16 degrees, integrate the IMU, stop early by his measured per-direction glide ā which he learns and updates every turn. His right side is weaker, so right turns are gentle walking arcs, capped near the ball because arcs also advance.
- The iron rule: the body never moves blind. Eye loses the ball, feet stop within a second.
The Ruler (knowing where he is)
- Monocular distance-from-size works at range but lies up close (it read 31-44 cm at the same physical spot). At close range the ruler is his own neck angle: head fully down = he has arrived. Bracketed on hardware to the tick.
- The gold pose: I placed him in the perfect grab stance by hand and recorded what his senses read. That measured stance ā not theory ā is the target for every approach.
The Grab Gate (the lawyers)
- Three conditions must agree simultaneously: head-down tilt (position), fresh visual identity (recognition), and fused center error under 18 ticks (alignment = head bearing + image offset).
- Each failing condition triggers its own correction: nudge forward, wait for recognition, or turn/strafe. Overshoot flips trigger minimum-size steps so corrections converge instead of ping-ponging.
- Every reading is a settled double-read ā judging while his head was still slewing once produced a phantom 218-tick error and a 30-degree turn into the ball.
The Mind (a local LLM with a body)
- An Ollama-served Llama 3 8B on my desktop plans actions from plain English, seeded with what he currently sees (from the vision service) and his real body state: battery voltage, roll/pitch, hottest servo temperature.
- The persona is contractual: eager, honest about limits, self-deprecating, and the āreasonā field of every plan is spoken aloud. When his battery is low, he declines strenuous work ā and jokes about why.
The Voice (Kokoro TTS + a decision journal)
- Kokoro runs locally on my machine; every action he takes is announced as he takes it. Repeats get swapped for rotating commentary instead of parroting.
- Every spoken line is journaled with its exact audio clip ā a replayable log of every decision in every run. The filmās voice track is composed from those clean clips at their true timestamps, not room audio.
The Body-Sense
- IMU fall detection with automatic get-up (front and back variants). Falls are rare now; the get-up still gets its reps in.
- All 24 servos report position, voltage, and temperature on an 8-second cycle (after patching a vendor bug that read the wrong register). Thermal gates refuse a run on cooked ankles; a self-leveling routine trims his stance so the left leg stops fighting a lean and running hot.
Mission Control + The Film Crew
- A one-screen dashboard: three live views, a robot figure with temperature-colored joints, head-to-toe servo list, brain thoughts, voice toggle.
- Recording is three raw streams stamped with millisecond wall-clock sidecars, combined offline by real time ā including measuring his vision pipelineās latency at record start (~0.5 s) so his eye view lines up with the room cameras.
The Attempt Ledger
- 31 ā First ever fully autonomous walk, grab, and throw. The milestone.
- 32 ā Stood in a perfect stance and refused to grab: recognition read 0.58 against a 0.60 bar. Infuriating. Correct.
- 33 ā His eye lost the ball at dock start and he marched on blind. Birth of the never-advance-blind rule.
- 34 ā Walked into the ball chasing a phantom reading. Birth of the eye-settle double-read and the right-arc cap.
- 35 ā Flawless. First-check gate pass, docked one pixel from the hand-measured gold pose.
- 36 ā Flawless again, back to back, on a fresh boot ā with new jokes.
Whatās Next
A RAM upgrade for his poor swapping Pi, a custom-trained detector distilled from his own recordings, more objects in his visual memory, and teaching him to do something useful with all this coordination, fetch, maybe. But mostly: more attempts, more numbers, more jokes at his own expense. Heās earned the reps.
Built with a Hiwonder AiNex, a Raspberry Pi 5, MediaPipe, pgvector, ROS, Kokoro TTS, a local LLM, and a truly unreasonable number of measured failures. Every fix came from a number he read himself.