← Back to list

Do AI Actors Dream of Meryl Streep?

(With apologies to Philip K. Dick*)

Andrew Leitch · 2026-05-08 18:31 · 0 claps · 10.4 min read
#immersive-video #ai #apple-vision-pro #turing-test #uncanny-valley
Open on Medium ↗
Wiki topics: AI · AI · General

Do AI Actors Dream of Meryl Streep?

(With apologies to Philip K. Dick)*

Nancy Burson — First and Second Male Movie Star Composites (Left: Cary Grant, Jimmy Stewart, Gary Cooper, Clark Gable, Humphrey Bogart. Right: Richard Gere, Christopher Reeve, Mel Gibson, Warren Beatty, Robert Redford.) 1984

Nancy Burson — First and Second Male Movie Star Composites (Left: Cary Grant, Jimmy Stewart, Gary Cooper, Clark Gable, Humphrey Bogart. Right: Richard Gere, Christopher Reeve, Mel Gibson, Warren Beatty, Robert Redford.) 1984

The Oscars recently published the rules¹ that will govern the judging of their 99th awards ceremony. The contentious question of using AI in film-making was only briefly addressed, and the Academy was largely neutral with regard to the use of generative AI in the making of a film. The rules simply stated that “the tools neither help nor harm the chances of achieving a nomination.”

On the subject of acting though, the Academy was very clear. Only roles that were “demonstrably performed by humans with their consent will be considered eligible”.

So only humans allowed, for now. But how do we actually verify that an actor is human? Let’s start with some history.

There are two famous conceptual benchmarks to grapple with in the realm of computers trying to impersonate humans—how to pass the Turing Test and how to cross the Uncanny Valley.

The Turing Test was a thought experiment proposed by British mathematician Alan Turing² in 1950. Turing originally called his test the Imitation Game. (The test was subsequently named after him.)

The test Turing outlined involved a human interrogator sitting on one side of a partition. Hidden from view on the other side were a human subject and a computer. The interrogator and the two participants started having a conversation, communicating solely through text-based interactions.

Turing proposed that a computer would pass his test if the interrogator couldn’t tell whether they were interacting with the human or the computer, which he predicted would be achievable by the year 2000.

And his original benchmark was actually quite modest: the interrogator only had to be fooled by the computer 30% of the time, for 5 minutes. Using that standard, it has been claimed that large language models have already passed the Turing Test (with some caveats³).

So we can consider that computers have already passed the Turing Test, at least according to the lenient standards proposed by Alan Turing himself.

Where the Turing Test measures language, the Uncanny Valley measures something more visceral: the unsettling feeling we get when some entity (an avatar, a robot, an NPC) looks almost human, but not quite, and so falls into the gap between an obvious approximation and a completely convincing facsimile. This triggers a feeling of unease and eeriness.

For an AI actor to clear the Uncanny Valley, it must look, sound and move convincingly enough that a viewer believes that they are watching a real person. There is always a lot of talk about the “look” benchmark, particularly with regard to the eyes. And the “sound” benchmark often gets called out, especially when an AI actor has to express some emotional depth with their voice.

But there is not as much focus on an AI actor’s ability to mimic human movement, both the larger movements of our bodies in space, and, more crucially, the subtle movements of our facial muscles.

It is worth noting here that the originator of the term Uncanny Valley, Masahiro Mori,⁴ was a roboticist. He coined the term Uncanny Valley in 1970, long before video games and CGI were a thing. He was certainly concerned that a robot should look as close to a human as possible to avoid that uncanniness. But he was also concerned that it should move like a human as well.

And for our AI actor, facial movements present an even bigger challenge, as they have to mimic the “micro-expressions⁵” that we humans all make, all the time: small fleeting emotional cues that flit across our faces in milliseconds, without our even being aware of them.

So has the Uncanny Valley been crossed? For still images, almost certainly. There are any number of quizzes we’ve all taken where you have to choose which face is AI and which one is real. And probably like many of you, I started failing those tests with some regularity a couple of years ago.

For moving images, though, the picture gets considerably blurrier. The human face, with its expressive eyes and fleeting micro-expressions, remains stubbornly hard to replicate convincingly, and once the length of a shot exceeds a few seconds, that challenge increases exponentially. The Academy is not going to give any awards for the best 8-second monologue.

So the current state of play for our AI actor is that it can probably pass a simple version of the Turing Test, but is not quite ready to cross the Uncanny Valley.

A foundational component of any film or theatrical performance is the collaboration between an actor and a director. So our AI actor has to be able to take direction, even if it’s in the form of prompts. The challenge becomes even more acute if there are contradictory directions. “Play it strong, but let us see how frightened you are underneath.”

Anyone who has done any experimenting with generative AI knows that it’s always a roll of the dice what you’re going to get back from your prompt. And then, to refine that with some subtle nuance and zero in on the performance you’re going for, sustained for a ~5 minute scene? At this point, it’s not happening.

But, I hear you say, AI is disrupting itself every other week, and it’s only a matter of time before these benchmarks are hit. Perhaps that’s true, but as the mountaineers say, don’t mistake a clear view for a short distance.

That said, to better fend off the synthetic thespians, at least for a software cycle or two, let’s raise the benchmarks by several notches.

Allow me to introduce a new thought experiment, which I’m calling the Uncanny Turing Test.

The Uncanny Turing Test is a two-part evaluation that combines a more rigorous version of the Turing Test with the challenge of fully crossing the Uncanny Valley. This test would allow us to unambiguously determine whether an actor is a human or a replicant.

For the more rigorous Turing Test part of this exercise, the interrogator now becomes a director, and the human and the machine become a human actor and an AI actor. And rather than just a casual conversation, the interrogator/director is discussing how they want a role to be played. Things like the character’s motivation, their backstory, what the scene is trying to accomplish in the arc of the story and so on. Pretty standard stuff when a director and actor prepare to shoot a scene.

But here’s the hard part. Both the AI actor and the human actor are still only interacting with the director via a text-based exchange. And the director, who has a separate discussion with the AI actor and the human actor, one after the other (the order isn’t important), doesn’t know which actor is which.

What this first part of our Uncanny Turing Test is asking is: can an AI actor not only impersonate a human, but also follow the (text-based) direction and deliver a performance based on that direction? And more crucially, when having the discussion with the director, can the AI actor convince them that it is actually a human?

Then, on the Uncanny Valley side of the house, there has to be a sustained dramatic performance that looks, sounds and moves like a human.

Based on the earlier discussion with the director, the AI actor will now generate a performance of the scene.

And for the test I’m proposing, the scene should be rendered in a medium close-up, to allow for a full examination of the face while still being able to see some of the body (head and shoulders).

We’ll then shoot the exact same scene, same script, same angle, with a human actor. The end result of this exercise will be two separate artifacts: an AI-generated version of an actor’s performance, and a human actor’s version of the identical scene.

Both actors will have “collaborated” with the director (via text) and then “performed” the scene based on that direction. Now it’s time to judge the two performances.

To make things even more challenging for our AI actor, we are requiring that the performance be viewable in Apple’s Immersive Video format.

Apple Immersive Video is, at the time of writing⁶, the highest-fidelity stereoscopic format available for representing the real world, and real humans. So a medium close-up at this resolution will be extremely unforgiving.

The audio should also be optimized for Apple’s spatial audio format, capturing not just the dialogue, but the ambience and the room tone, in a full 360º soundscape.

(I’ve written at length about the power of Apple’s Immersive Video on my Substack *Watch This Space *and also here on Medium.)

Here are the attributes that will make it especially challenging for our AI actor to convince us that it’s a human.

Apple Immersive Video shoots at 90 frames per second, which removes the motion blur that softens the boundary between frames. At 24 frames per second on a flat screen, micro-expression inconsistencies tend to “hide” in the gaps. At 90 frames per second, there are no gaps. There is nowhere to hide.

Apple Immersive Video shoots at 8K per eye, with two lenses replicating the depth of human vision. The result doesn’t feel like a recording of a face, it feels like a face. Our AI actor’s skin, their pores, the moist surface of their eyes, the subtle movement of the pulse in their neck, all of these attributes will have to be present.

There are 16 stops of dynamic range. Human faces are extraordinarily dynamic when it comes to absorbing and reflecting light. CGI artists have long struggled with fully replicating this play of light across a face, with its attendant emotions (think of blood rushing to the surface of the skin). With 16 stops, it becomes that much harder to get away with.

Spatial audio adds a second perceptual channel, which our brain processes on a pre-conscious level. The voice isn’t coming from a screen, it’s coming from a specific three-dimensional position of a mouth. A small mismatch between where the sound appears to originate from and where the mouth actually is will be immediately jarring.

Finally, the choice of a medium close-up makes it especially challenging for our AI actor. The judges will be completely engulfed in a stereoscopic hemisphere, allowing them to see the whole face in extraordinary detail.

Passing the Uncanny Turing Test means that there is a genuine lack of certainty about which actor is real and which is AI-generated.

My proposed panel to judge these two scenes would consist of a Director, a Director of Photography, a Sound Designer, a Casting Director and a Regular Viewer.

The Director would be laser-focused on the performances, and how well the actors interpret the prompts and turn them into convincing portrayals.

TheDirector of Photography would be paying close attention to the lighting and how it plays across the face, as well as elements like skin tone and the eyes (as ever, one of the most challenging aspects to get right).

The Sound Designer would be listening to the voice and how it reverberates in space, and more subtle cues like the sound of the breath and the ambience of the room.

The Casting Director would be looking for the most convincing performance, channeling their deep experience from watching multiple actors perform the same scene in audition tapes.

The Regular Viewer would experience the test without the trained eye of a professional, but with something equally valuable: an unmediated gut response. They won’t be able to articulate why something feels off, but they’ll feel it immediately.

That’s a formidable jury our AI actor has to convince.

The AI actor will be judged to have passed the Uncanny Turing Test if it a) fools the director in the initial “collaboration” stage, and b) fools the jury after they have viewed both the AI actor and the human actor back to back.

My proposed rules of engagement for this test are spelled out in more detail in the footnotes.⁸

A couple of practical considerations, if anyone wants to actually run the Uncanny Turing Test (and the Academy is also most welcome to borrow any criteria here that might be useful for future award ceremonies).

First off, no-one has generated any AI videos in the Apple Immersive Video format in all its “90 frames per second, 8K per eye, 16 stops of dynamic range” glory, at least as far as I’m aware (and I follow the immersive cinema space pretty closely).

And secondly, to expect an AI actor to be able to take direction, internalize it, and then adjust its performance is another way of saying it would have to get a lot closer to general intelligence. The benchmark to end all benchmarks.

The bar is deliberately set very high.

So I’d like to think that Apple’s Immersive Video format is a “moat” (as the VCs like to say) to provide some protection and shelter for our endangered human actors, from Meryl Streep on down.

And when we do pass the Uncanny Turing Test, and hand the best actor statuette to Claude, we might also be ready to move into a full simulation⁹. Unless of course we’re already in one.

Let’s watch this space.

For those who don’t get the Philip K. Dick reference, his novel “ Do Androids Dream of Electric Sheep” was the basis for the movie Blade Runner.

Originally published at https://andrewleitch.substack.com.

Notes:

  1. See the full Academy of Motion Picture Arts and Sciences guidelines here.
  2. More information on the origins of the Turing Test can be found here.
  3. Gary Marcus has a very good analysis of the LLMs taking on the Turing Test. He also proposes his own modified Turing Test, which would be very hard to pass!
  4. See more on Masahiro Mori here. He’s a fascinating figure, and worthy of a separate article, particularly around his fusion of robotics and Buddhism.
  5. See Cinematic Immersive for Professionals (Ben Allen ACS CSI and Clara Chong) for a detailed analysis of the power and the importance of micro-expressions in immersive cinema. Their book is highly recommended reading for anyone looking for a comprehensive overview of this emerging cinematic language.
  6. Apple Immersive Video, as viewed on the Vision Pro, is the highest resolution stereoscopic format currently available. But the Uncanny Turing Test should work with any future formats coming to the market that are of a similar, or even higher quality.
  7. The rules of engagement are:
  • There will be a single monologue, chosen for its dramatic range, and it should be around five minutes long (giving a nod to Turing’s original length for his imitation game).
  • There will be an AI actor and a human actor both performing the exact same scene.
  • Both actors will be given text-based prompts to perform the scene from a director, and there will be an exchange between each actor and the director to arrive at an agreed-upon approach on how to perform the scene.
  • The AI actor and the human actor will be of the same gender, race and approximate age.
  • All five judges will view the performances on an Apple Vision Pro, and listen to the dialogue and sound design wearing AirPods Max headphones, to allow for the full immersive experience.
  • For the AI actor to pass the Uncanny Turing Test, there are two separate components. First, the director giving the text-based prompts should be unable to tell if they are collaborating with an AI actor or a human actor. Second, all five judges viewing the performances should be unable to tell the difference between the two.
  • Adding up all the judges in both components of the test, there are six votes to identify whether the actor is AI or human. The AI actor has to fool the judges 50% of the time (or three out of six votes) to win.
  • If any individual judge is unable to definitively choose which actor is which, that also counts as a win for the AI actor.
  • The test can be run multiple times with different sets of judges to get a more statistically-meaningful sample.
  1. See more on the Simulation Hypothesis here.

메타데이터
post_id
8b5828b77f56
slug
do-ai-actors-dream-of-meryl-streep-8b5828b77f56
url
https://medium.com/@betavi11e/do-ai-actors-dream-of-meryl-streep-8b5828b77f56
canonical_url
https://medium.com/@betavi11e/do-ai-actors-dream-of-meryl-streep-8b5828b77f56
author_url
https://medium.com/@betavi11e
status
ok
fetched_at
2026-06-09 15:37:30