In July 2024 a human-centred design professor emailed me about his lab's NAO, a two-foot humanoid, wanting the AI wired into it and holding a real conversation. The next afternoon he walked the robot to my office, gave me the credentials, and left it with me.
NAO, two feet tall.
Out of the box, NAO converses the way a phone tree does, listening for keywords and playing back scripted replies. Research data meant the model could not sit on a public endpoint, so the work stayed on the university's own AI infrastructure.
Getting it onto the campus network took the first stretch, because nothing gets past DNS until its hardware address is registered and the only way to read the robot's was a wire into the back of its head and into my Mac. Register it, press the chest button for an IP, then ssh in with credentials that ship in the docs. One cat /etc/os-release later it was Gentoo1, close enough to the Arch I run daily that nothing about the filesystem slowed me down. The SDK documentation is barely indexed, sometimes in French and sometimes gone, so the afternoon went to poking at something odd and hunting the docs for it. Then the qi command line set the order of the work by itself, speech first, then motion, then the audio APIs. By the end of day one it could talk and walk on command.
NAOqi, the robot's operating layer, publishes every part of the machine as a network service, the microphones, speakers, motors and the gesture library2. So the robot keeps being a body while a laptop on the same Wi-Fi does the thinking. Between them it is a few hundred bytes of text per turn.
A body and a brain
The conversation program is about 150 lines of Python that connect, load a persona, then loop through listening, transcribing, asking the model, and speaking the answer with gestures.
session = qi.Session()
session.connect(f"tcp://{ip}:9559")
history = [{"role": "system", "content": open("prompt.txt").read().strip()}]
while True:
user_input = listen_and_transcribe() # the laptop listens
history.append({"role": "user", "content": user_input})
response = get_gpt_response(history)
history.append({"role": "assistant", "content": response})
nao_speak_with_animations(session, response) # the robot answers
NAOqi's animated-speech mode picks the gestures to match what the robot is saying, and it takes stage directions inline with the words:
nao_speak_with_animations(
session,
"^start(animations/Stand/Gestures/Hey_1) Hello! I am NAO, and I'm "
"ready for a conversation. ^wait(animations/Stand/Gestures/Hey_1)",
)
NAO does not balance itself out of the box, so one careless "walk" and you are lunging to catch it, and those inline stage directions mean certain words can trip a gesture. Python versioning on a fourteen-year-old SDK is archaeology. Knowing when a person has finished talking is the hard one, since cutting the mic early interrupts them and cutting it late makes the robot feel broken, so the listening runs through the laptop and a speech SDK with that judgment built in while the robot carries the voice and the body.
The first real conversation, I asked the robot for the longest word in the English language. NAO tracks your face while you talk, head tracking being one of its built-in background processes, and a few gestures run on top of that.
What the professor wanted next was to run it himself, and that is what the five days were for. The deliverable was a setup script that builds the whole environment on a fresh Mac and tests the AI connection and the robot connection separately, so he could tell which half was unhappy without me in the room. By mid-September he was running it from his own desktop and writing his own personas, and he kept it going for the next ten months.
Current project status.
-
OpenNAO, the robot's operating system, is a Gentoo-based Linux distribution. doc.aldebaran.com/2-0/dev/tools/opennao. ↩
-
NAOqi exposes the robot's subsystems (audio, motion, speech, memory) as services callable over TCP on port 9559, with Python bindings. SDK reference: doc.aldebaran.com. ↩