← Back to list

Lost in Transcription: The Week the Machine Started Lying

Experimenting with Whisper AI — Part 3

Pej Canlas · 2026-07-01 04:46 · 0 claps · 9.6 min read
#artificial-intelligence #whisper-ai #google-translate
Open on Medium ↗
Wiki topics: MM · Multimodal & Generative Media AI · AI · General 🔬 · Science · General

Lost in Transcription: The Week the Machine Started Lying

Experimenting with Whisper AI — Part 3

I lined up my AI against the other machines that translate for a living. I expected them all to be the same flavor of cold. One of them turned out to be a liar.

Part 3 of a four-week experiment in AI subtitles, nuance, and everything that gets lost in between.

Cover image made by the author with AI (Claude). One word, harot, translated three ways by three machines. Google’s is the one that got it wrong.

Cover image made by the author with AI (Claude). One word, harot, translated three ways by three machines. Google’s is the one that got it wrong.

I went into this week planning to do one thing and ended up doing another, so let me just be honest about that from the start.

If you read last week, I signed off promising to throw my AI straight at Netflix. The professionals. The people who do this for a living. But when I actually sat down to start, I stopped myself, because it felt like skipping a step. Before I put my little pipeline up against human experts, I wanted to know how it did against its own kind first. The other machines. The free auto-translators everybody already uses. If my carefully steered AI couldn’t even beat a plain machine, there was no point dragging Netflix into it at all. So the pros can wait for the finale. This week was machine against machine.

And before I go any further I have to clear up something I’d honestly muddled in my own head, because if I was confused about it, you probably are too. This whole series is named after Whisper, and yet this week Whisper barely shows up. That’s not me losing the plot. It’s just how the thing is actually built. My pipeline is really two different machines doing two totally different jobs. Whisper is the first one, and all it does is listen. It takes Bambi’s audio and writes down the Tagalog she’s saying, sound in, text out, same language start to finish. It never touches English. It doesn’t translate a thing. It’s the ears. Everything after that, the actual carrying of Tagalog into English, is a completely separate machine. Google Translate, or SubtitleCat, or the AI I steer myself. That’s the second stage. The mind. Weeks one and two were really about the ears, about whether the machine could even hear a fast, sobbing, code-switched monologue, and what happened when it couldn’t, like that hallucinated “baby.” But by this week the ears had already done their part. My Tagalog transcript was sitting there, cleaned up and correct. This week was never about hearing. It was about meaning, and about what happens when three different machines try to carry those already-heard words across.

So that’s the setup. I took my own translation, the one I’d spent last week coaxing an AI into doing with real care, and I lined it up against two others. SubtitleCat, one of those services that auto-translates subtitle files. And plain old Google Translate, fed my own hand-corrected transcript. Three versions of the same scene, side by side. I figured they’d all come out roughly the same, a bit flat, a bit mechanical, and that my coached one would edge ahead on warmth.

That is not what happened.

I almost hadn’t bothered with Google Translate. I threw it in mostly to have a “worst case” on the table, the cold baseline nobody expects anything from. And then it did something I wasn’t ready for. It didn’t just translate the scene flatly. It translated it wrong, in ways that quietly change who Bambi even is. There’s a line where she says she never had the time to feel kilig, or to be harot. That fizzy, innocent, fluttery stuff, the butterflies of a crush, the harmless teasing, the small youthful joys she gave up so her family could eat. SubtitleCat softened it to “fall in love and flirt,” a little flat but basically fair. Google looked at harot and gave me “horny.”

I actually stopped and stared. Horny. It took a tender, aching line about a girl who never once got to feel butterflies and made it crude. That is not a cold translation. That is a translation that swaps out her soul and hands you a stranger wearing her name. And it kept doing it. Her curse, putang-ina, the kind you breathe out in pure grief at nobody in particular, came back as a calm, polite “Mother,” like she’d paused her breakdown to gently address her mom. Then the worst one, right at the emotional peak of the whole scene, the line where she says that once you break her open and scoop out her laman, her insides, she’s worthless. Google heard laman and gave me “you took my sword.” A sword. The single most important image in the monologue, the thing the entire ending is built on, turned into gibberish.

That was the moment this stopped being a tidy comparison and became something I actually had to sit with. I’d been so busy asking whether a translation was warm or cold that I’d missed a worse thing hiding underneath. A translation can be perfectly fluent, perfectly confident, grammatically spotless, and still be lying straight to your face. “Horny.” “Mother.” “Sword.” Three clean little sentences, every one of them wrong, every one of them completely sure of itself.

And here’s the strange part. Cleaning up that mess, I found myself feeling almost tender toward the thing. Because not that long ago, this was the only option there was. If you wanted to translate something yourself, you pasted it into Google Translate, or something just like it, and you took whatever came back, “sword” and all. There was no second tool quietly weighing a better word. No model you could nudge and tell, keep the tone, she’s grieving, don’t make her crude. You got the literal output and you were alone with it, which is exactly why doing your own subtitles used to be so hard unless you already knew the language cold. You had to catch every one of those lies yourself, because nothing was helping you catch them. That’s the part that gets buried under all the noise about AI lately. The thing I have now, a model I can actually steer, that weighs and filters words instead of grabbing the first literal match, is a real kind of help that just did not exist for someone like me a few years ago. Google handed me “horny.” The newer tool, asked properly, handed me back the innocence of harot. That gap isn’t small. It’s the difference between a tool that translates at you and one you can translate with.

Then the other machine embarrassed me

Help isn’t the same as winning, though, and this is where my little David-and-Goliath story fell apart. Goliath, it turned out, was just another machine, and it was doing fine without me.

I was so sure SubtitleCat would butcher the alkansya. You know by now how much I’ve made of that line, the sealed coin bank you can only empty by smashing it, the one that becomes basag na ako, “I’m shattered,” a few beats later. I was certain a plain auto-translated subtitle, with no room for one of my clever notes, would have no choice but to let it die. It didn’t. SubtitleCat gave me “I’m just a savings jar,” then carried the whole thing through, clean. “If you break me and take everything I have, I’ll be worthless.” Then, a beat later, “Shattered. I’m broken into pieces.” The metaphor lived. Fully intact, right there in the line, no asterisk, none of the machinery I’d been so proud of.

That stung. And it stung twice as hard once it hit me that SubtitleCat runs on the same family of technology as the Google Translate that had just handed me “sword.” Same lineage. Wildly different results. Which nagged at me for a while, actually, because if they’re both machines working from the same Tagalog, why did they come out so far apart? The answer, once I dug into it, was that I’d been picturing them as the same kind of thing when they’re really not. Google Translate is the old guard, built to map words from one language to another as directly as it can, hunting for the most statistically likely equivalent. That’s literally why it reached for “sword.” It grabbed a probable match and never once stopped to feel the scene. It’s a fast, literal dictionary that guesses. SubtitleCat runs on that same Google-family engine but tuned for subtitles, so it comes out a little cleaner. And the thing I steer, a large language model, is a genuinely newer kind of machine that wasn’t built just to swap words. You can talk to it. You can say she’s grieving, keep the tone, and it can actually weigh that. So they’re not three versions of one machine. They’re three different generations of the whole idea of what a translating machine even is, and the distance between “sword” and “savings jar” is really just a few years of that idea growing up.

The thing that quietly rearranged the whole project for me, though, came when I went looking for why SubtitleCat’s version worked so well. Because it turns out the goal of subtitling was never to be exact, and I’d been grading all three of these like it was a spelling test. One thing I read put it bluntly: being literal doesn’t protect accuracy, it wrecks the experience, because it forces the viewer to read a line that makes no sense where they’re sitting. What you’re actually chasing is the same effect. If the original made you ache, the translation has to make you ache too, even if the words come out completely different. The thought has to land. The person watching has to feel it. That’s the whole job. Which means “savings jar” was never a compromise at all. It was the craft doing exactly what it’s meant to do, folding the feeling into a line that hits the way the original hits. I’d been measuring for the wrong thing the entire time, and the plain little machine had quietly been playing the right game.

So were my notes pointless, then? For about an hour, sitting there, it honestly felt like yes. But not quite. Even SubtitleCat, doing the job well, let things slip. “Savings jar” holds the breakable idea but tells a non-Filipino nothing about what an alkansya means in a house, the slow saving, the small daily sacrifice it stands for. It softened harot to “flirt” and lost the innocence. It censored the curse instead of carrying the grief. None of those are mistakes exactly. They’re residue. The fine cultural dust even a good, feeling-first translation can’t hold, because one line of screen text is only so big. And that residue is the one place my translator’s notes still earned their keep. Not as a replacement for good translation, which is the arrogant thing I walked in believing, but as a quiet second layer underneath, a little shelf for the things no single line has room to carry. A far smaller claim than the one I was making last week, standing over that alkansya line like I’d personally invented the idea of caring about it. Also the first version of the claim that’s actually true.

Counting it, because a feeling isn’t proof

I didn’t want to just vibe my way to all this, so I made myself score it. Ten of the hardest nuance points in the scene, three versions each, every one marked preserved, partial, lost, or distorted. The numbers were blunt. Google distorted five of the ten. Not softened, distorted, actually wrong. SubtitleCat distorted zero and preserved six. My own version sat in between, never wrong but often only halfway there.

Seeing it counted out finally made the whole thing click. This was never a clean wall with warm humans on one side and cold machines on the other. There were no humans in it at all. Three machines, spread along a slope. The oldest approach distorts, a newer one pulls most of it back, a steerable one with notes catches even more, and that whole spectrum from “sword” to “savings jar” is just the same technology caught at different points in its own growing up. And the most dangerous spot on that slope, I realized, isn’t the cold end at all. It’s the confident-but-wrong end, where “sword” and “horny” slide right through wearing perfectly good English.

The cocky little story I walked in with is gone, and honestly I’m glad, because what replaced it is more interesting than a David-and-Goliath rerun. It was never me versus the machine. It was one machine against another against another, and me off to the side watching which of them could carry a grieving woman’s words across a gap that’s so much wider than it looks. The real difference between them was never warm versus cold. It was old versus new. Which one lied, and which one had learned better. Google, the oldest, lied to my face three separate times in a single scene. The newer ones, including the one I can now actually steer, mostly didn’t.

So, the question this whole series keeps circling back to. Can AI hear hugot? Three weeks in, here’s where I’ve landed for now. The newer AI can hear it well enough to genuinely help, in a way the old translate-and-pray tools never could. It can be steered toward the feeling instead of the bare literal words. But exactness was never the point anyway. The point was making a stranger feel what Bambi felt, and measured against that, the real bar, the machine is a genuinely useful partner now, as long as somebody who knows the language is sitting right beside it, deciding when it’s reaching for the truth and when it’s just handed you a sword.

There’s one contender I still haven’t put in the ring, though. The professionals. The official Netflix subtitles, made by actual human translators who do this for a living. That’s the fight I really want to see, and it’s the one I’ve been saving for last.

One week left.

Part 3 of 4. Next time: the finale. My AI, and the machines, finally up against the professionals at Netflix.

A few notes and sources

  • Both English versions compared here are machine translation: SubtitleCat (a free subtitle service whose engine runs on Google-Translate-family technology) and Google Translate of my own hand-corrected Tagalog transcript. SubtitleCat’s reliance on machine translation is described in its own documentation and in third-party guides to the tool.
  • The principle that subtitling aims for equivalence of effect rather than literal accuracy, and that literalism can wreck the viewing experience, draws on professional guidance from Translated.com and on subtitle-translation scholarship grounded in relevance theory and Newmark’s communicative translation.
  • The film is And the Breadwinner Is… (MMFF 2024); the scene is Vice Ganda’s monologue as Bambi. All scoring and analysis are my own.

메타데이터
post_id
a767daa13ee0
slug
lost-in-transcription-the-week-the-machine-started-lying-a767daa13ee0
url
https://medium.com/@pejcanlas25/lost-in-transcription-the-week-the-machine-started-lying-a767daa13ee0
canonical_url
https://medium.com/@pejcanlas25/lost-in-transcription-the-week-the-machine-started-lying-a767daa13ee0
author_url
https://medium.com/@pejcanlas25
status
ok
fetched_at
2026-07-11 13:11:35