← Back to list

What Amazon’s Astro Taught Me About Giving AI a Soul

Character is the difference between a machine people tolerate and a product people trust.

Mike Forst · 2026-04-27 19:10 · 12 claps · 10.9 min read
#human-robot-interaction #robotics #ai #ai-voice-agent #robots
Open on Medium ↗
Wiki topics: AGT · AI Agents AI · AI · General 🧘 · Spirituality

What Amazon’s Astro Taught Me About Giving AI a Soul

Character is the difference between a machine people tolerate and a product people trust.

In 2018, Amazon brought me in as the lead UX Sound Designer for Astro, their first consumer home robot. Astro used cameras and other sensors to map and navigate your home and workplace, and could proactively patrol, check up on loved ones, detect people, and transport small items using its built-in cargo bin. While there was a well-defined feature set and form factor, initially there was no character direction. In fact, even before Astro had a name, there were two main questions — was it simply Alexa on wheels, or was it a robot with its own character?

The Astro team was divided. One option was to focus on Alexa, and treat the mobile robot simply as an added utility. I argued for Astro to not focus on Alexa, along with the majority of the UX team. Our belief was that a thing that moves through your home and turns toward you with intent can never be just an appliance. People would ascribe character to whether we wanted them to or not, and so the only question was whether we shaped that character or let it happen by accident.

Ultimately, Astro was Astro rather than Alexa, and our user testing backed up our decision. People didn’t see the robot as Alexa. They saw it as its own character, and that’s what they wanted it to be. Alexa on the device felt somewhat strange and creepy, but building Astro its own voice was off the table in 2018, being too expensive and too slow. So, we settled on Alexa as a supporting character that handled any actual talking, while Astro was the main character, communicating as much as it could without words, through sound, motion, and facial expressions.

I had been brought on to the Astro team to define the robot’s sound design language and voice. But there was no one to flesh out the robot’s actual character. You cannot make a single real decision about a character without defining it first. Every choice about how Astro moved, sounded, paused, or reacted was a character choice, and those choices required all disciplines working together. As Sound Lead, I was weaving together sound, motion, and character, and how they played together inside each story moment. The animators, who programmed Astro’s motion and facial expressions, were extraordinary at what they did, but the emotional arc they were animating came from the sound (and therefore character) work first. So I stepped into that role, which is where my real work started. What I learned about building character for robots applies to nearly everything being built in embodied AI right now.

Character Is Not Just an Output. It’s a Design System.

Developing a character for Astro meant answering questions never asked about a product at Amazon: What is the emotional range of this robot’s baseline state? How does this robot communicate uncertainty without eroding trust? Where is the line between being expressive and annoying? What are the vulnerabilities of this device’s character?

These are design questions. They have real answers, and every team working on the product has to build from them. For example, Astro’s emotional range was designed to be relatively small at first. We never wanted Astro to get too sad or too angry. It could play sad, but would snap out of it quickly and end the reaction on a high note to keep things positive.

[embed]- YouTube Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on…www.youtube.com

Character leaks out of every seam and can create a disjointed experience if not defined correctly. Even if it’s just animation timing that’s slightly off or a response that’s technically correct but contextually tone-deaf, users feel every one of these inconsistencies, even if they can’t name them. Watch what happens at the beginning and end of this Sing sequence:

[embed]- YouTube Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on…www.youtube.com

Astro goes from nothing, into the emotional moment, and then lands back on nothing. No build up, no cool down, no sense that the feeling came from somewhere or had anywhere to go. I pushed hard for better character stitching, the transitions in and out of expressive moments that make a performance feel continuous rather than assembled, but it never got implemented. The moment itself works. But without the stitching, it reads as a clip playing on a robot rather than coming from within the robot character itself.

Story and Sound at the Beginning of the Character Pipeline

In much of modern feature animation, especially at studios like Pixar and Disney, animators build on storyboards as the foundation of the film. Then the voice actors bring life to all the characters in pre-production. The animators are not just syncing to the sound, but translating and enhancing it all through physical performance. While the process is iterative and can vary, scenes tend to feel most authentic when voice, timing, and animation are developed closely together rather than the voice being imposed after the fact.

We had decided that Astro would have no spoken dialogue, but it had something that functioned the same way: a vocabulary of sounds, tones, and rhythms that acted as a voice. This vocabulary became the leading output of the character’s personality. The robot’s motion and facial expressions were built around it.

The wake-up sequence is a great example. Astro waking wasn’t just a boot animation on the screen. It was a performance. Slow and humble at first, it oriented itself quietly, then stretched its screen, checked its wheels, and finally, with an upward gesture toward its telescoping mast, it popped it up slightly, and did a little dance of joy. Sound, motion, and eyes hit every beat together in full choreography.

The character’s output in that sequence was first written as a story. Astro is waking up in its new home for the first time. Its main aspiration is to be part of a family, so this is the moment it has been waiting for, this is its purpose. Being the responsible character that it is, it wants to make sure everything is good to go before it introduces itself and starts learning its new home.

This narrative came first because it drove every other decision that we made. After the story was written, sound gave that story a metaphorical voice: the excited tones, the pacing as it checked its wheels, and the bright melodic phrase as Astro looked up at its new family for the first time and introduced itself. Once the sound was laid down, animation did their thing with motion and facial expressions, taking cues from the emotional arc the sound had established. Motion didn’t lead — it followed the feeling of the story and the sounds, the same way an animator follows a recorded vocal take.

That wake up sequence became one of the most-discussed moments in early user testing. People described it as “alive”. What they were responding to wasn’t any single element. It was all three channels (sound, motion, and facial expressions) expressing the same defined character in harmony.

Context Is Where Character Becomes Real

The most compelling characters are defined not by a fixed disposition but by how they respond to their environments and the people in them. They’re still recognizably themselves even as they adapt. This is what I call contextual character. A robot living in a home doesn’t occupy a single emotional state. It moves through rooms with different energy, encounters people in different moods, operates at different times of day, and navigates an endless range of social situations it was never explicitly designed for. A character that doesn’t respond to any of that isn’t present. It’s just operating. Getting that right requires the outputs to be alive in the same way: not a fixed library triggered by discrete events, but a responsive system that adjusts based on what is actually happening. Quieter in the evening. More tentative with unfamiliar people. A different quality of attention when someone is approaching versus when moving through an empty hallway.

We got close to a contextual character output with Astro’s sound. We built toward it using a video game audio middleware tool called Wwise, authoring states, switches, and real-time parameters so the sound could shift with context instead of firing fixed audio assets. When the context was fed in, the system worked beautifully. Astro’s driving sound is a good example. We had a real-time signal for how fast the robot was moving, so we mapped that context to the pitch, multiple layers of sound that play at different speeds, and the volume of the sounds. The faster Astro went, the more the sound evolved. Because it was contextual, it felt completely alive. But every state like this was still a prediction we made by hand — a situation we had to imagine in advance and design a response for. And a designed response only comes alive if the system feeds it the right context. Astro’s driving sound worked so well because we had the speed signal. Plenty of other contexts we designed never got connected (e.g. time of day, time since last interaction, who we’re talking to, etc.), because routing that kind of live context into the audio system sat further down the priority list. A random home throws more situations at a robot than anyone can possibly predict, so there was always a longer tail of moments the system was never prepared for.

The difference between a product people describe as “smart” and one they describe as “aware” can often come down to this. Smartness is capability. Awareness is context. Presence is character. And character is always in reaction to the people around it, to its environment, to its own evolving state. That’s what makes it feel like something is emotionally present with you.

From Customization to Adaptation

There is a huge difference between a product that can be customized and one that adapts. Customization puts the burden on the user by asking them to configure it, to choose a preference, and then maintain it. It treats character as a setting to be dialed in rather than a relationship to be developed. Most products today, if they offer any personalization at all, only offer customization. While I agree it is better than nothing, and users will still want to be able to do some direct customization, it is not the same thing as adaptation.

Adaptation means the character evolves because it is paying attention. It learns that this household doesn’t wind down at night, it comes alive. That the person who seems cold in the morning is actually just not a morning person, and by 10am is the warmest one in the room. That the kids get home at 3:15 and the whole energy of the house shifts in an instant. That one family member is going through something, and has been quieter than usual for three days. None of this requires the user to configure anything. No designer could have authored it in advance. The relationship gets deeper the more time you spend with it, the way human relationships do.

This is where AI changes the game for character design in ways that go well beyond what was possible with Astro. The contextual responsiveness we had with Wwise was powerful, but it was authored, designed in advance by people making predictions about what situations the robot would encounter. AI-driven adaptation doesn’t require those predictions. It learns the specific rhythms, preferences, and emotional context of the people it lives and works with. The character doesn’t just respond to context. It grows into it.

What the Industry Is Missing

The character and soul of the impending wave of embodied AI products appears to almost always be an afterthought (if that). And character defined late is character defined by default. It becomes the sum of a thousand small decisions made by different people thinking about function, schedule, and cost, and not about character. People project character onto devices whether you plan for it or not, especially if those devices move — a robot that moves is already a character. If nobody has designed this character, the result will be products that feel like nothing, or worse, feel confusing and not trustworthy. Technically impressive but lifeless.

We did not get this fully right on Astro. So many things were moving fast in parallel that character was an uphill climb even after we agreed it deserved to be its own thing. We did have some real wins, including the wake-up sequence, the sound vocabulary, and contextual sounds. But character was rarely treated as a utility, and I understood why. When you are building a first-of-its-kind product, the things that are the loudest are the ones that break, the deadlines, the costs, the features a customer can point to on a box. Character is quieter than all of that. It is easy to assume it can come later. On a team as large as the Amazon Astro team, it’s lucky to get any idea onto the roadmap when it is competing with a hundred others that all feel more urgent in the moment. None of this comes from people not caring. It came from character being the kind of thing that is hard to prioritize until you see what its absence costs you.

My Asks to Product Leaders

If you are building a product that will share physical or conversational space with people, three things are worth considering earlier than most teams consider them now:

Define character before you define interactions. Simple adjectives won’t cut it. You need a defensible character with enough emotional logic to answer hard questions consistently. What does this product do when it’s confused, and why? How does delight feel different from satisfaction? Where is the line between attentive and unsettling? These questions have answers. Find them early and have every discipline build from the same foundation.

Build story and sound into the character pipeline, not the production pipeline. Story and sound developed alongside character definition has the chance to inform motion, expression, and interaction logic the same way a storyboard and voice performance informs the loveable characters in animated films. This is a different kind of contribution than producing a sound library at the end. It requires a different kind of collaboration, and a different kind of hire.

Design for adaptation, not just consistency. A consistent character is necessary, but the products that will matter most in people’s lives are the ones that deepen through use. The infrastructure to support that is more and more accessible. The design thinking to take advantage of it is still rare.

S.O.U.L

These asks come from one underlying model, the thing I keep coming back to after every project. The AI products that will matter most in people’s lives won’t just be technically impressive, they’ll have real character. That comes down to four things. I call the system S.O.U.L.

Systemize character. Principles that drive decisions, not just adjectives

“Friendly” is not a spec. Real character design produces answers to hard questions: what does this product do when it’s confused? Where is the line between attentive and unsettling? What are the vulnerabilities of this device’s character?

One voice across all outputs. All output channels work together toward one intent

Sound, motion, and expression are not separate deliverables. When they hit the same beats, users feel something coherent. They feel something that’s “alive”. When they don’t, users feel something’s wrong, even if they can’t name it.

Understand and use context. Character changes with environment

A fixed disposition is a setting. Character responds to what’s around it. Quieter in the evening, more tentative with strangers, different approaches when someone is upset versus when everyone is relaxed.

Learn and adapt. Not customization but evolution over time

Customization puts the burden on the user. Adaptation means the character pays attention and remembers. It learns the rhythms of its environment, reads the energy of a room, deepens through lived experience.

This is what I’m after. Character that gets better the longer it lives with you, until the product stops feeling like a product and starts feeling like family.

Note: I worked with AI to sharpen this piece. The Astro work, the conclusions, and the opinions are my own.

Mike Forst is a Los Angeles-based character director and sound lead with 15+ years of experience shaping how technology feels, behaves, and comes alive. He served as Character and Sound Lead on Astro, Amazon’s first consumer robot, and has spent years designing voices and characters for AI. He has worked with companies such as Google, Microsoft, Meta, and Zoox, and currently designs and consults with teams building the next generation of intelligent machines.


메타데이터
post_id
989fcd9c45f4
slug
what-amazons-astro-taught-me-about-giving-ai-a-soul-989fcd9c45f4
url
https://medium.com/@mikeforstmusic/what-amazons-astro-taught-me-about-giving-ai-a-soul-989fcd9c45f4
canonical_url
https://medium.com/@mikeforstmusic/what-amazons-astro-taught-me-about-giving-ai-a-soul-989fcd9c45f4
author_url
https://medium.com/@mikeforstmusic
status
ok
fetched_at
2026-08-05 11:16:34