When AI Meets Your Messy Data
What We Found in Past the Pilot Episode 8. Watch it here on youtube
When AI Meets Your Messy Data
What We Found in Past the Pilot Episode 8. Watch it here on youtube
Every business has two filing systems. One is the official record. The other is reality.
The official record lives in your CRM, your accounting system, your ERP. Rows, columns, clean fields. Auditable. Queryable. Trustworthy — in theory.
Reality is the invoice sitting in someone’s email. The PDF contract buried three folders deep in a shared drive. The engagement letter that got signed, scanned, and never quite made it into the database.
Most businesses spend enormous human energy trying to close the gap between these two worlds. Manually. Slowly. Expensively.
That’s exactly what Jonny McFadden and I tackled in Episode 8 of Past the Pilot. And what we found — live, on camera — was more than a workflow improvement. It was a different way of thinking about what AI can actually do with your data.
The Problem No One Wants to Admit
Real businesses don’t get to choose between structured and unstructured data. They get both. Always. The invoice arrives as a PDF. Someone keys it into the accounting system. Later, someone else verifies that the keying was accurate — that the numbers match, nothing was missed, that the document and the record are telling the same story.
Multiply that by hundreds of documents. Thousands.
That’s not a workflow. That’s a tax on your most expensive resource: human attention.

The traditional answer was OCR. Better OCR. Smarter OCR. But even the best optical character recognition of the past decade could read a document. It couldn’t understand one. That gap — between reading and understanding — has always been where the process breaks down.
That gap is now closable.
What Happened When Jonny Uploaded a PDF
Here’s the moment that stopped me mid-sentence.
Jonny uploads a single invoice PDF. No schema defined. No agent configured. No instructions given. He just drags it in. Within seconds, the platform surfaces structured metadata alongside the document — field names, extracted values, context — all of it inferred directly from the file’s contents. No prompting. No configuration. No human in the loop.
It just did.
But the part that genuinely made me pause came earlier in the story. When Jonny first uploaded those invoices to a completely empty account — no schemas, no templates, nothing predefined — the platform didn’t wait to be told what mattered. It read the documents, inferred the structure, named the fields, and built a working schema from scratch. The kind of schema a developer would normally spend days designing. Done automatically. Done well.
Jonny put it plainly: the schema the platform created was better than anything he would have built himself.

That’s not a small claim. Schema design is one of the most time-consuming phases of any data pipeline project — the part where smart people argue about field names and edge cases for weeks before a single line of code gets written. Here, it happened before Jonny even thought to ask for it.
Building an Agent by Talking to One
Vertesia Studio is where the reconciliation agent actually came to life — and the way it got built is as interesting as what it built.
No code. No pipeline configuration. No flowchart. Jonny described what he wanted in plain language, and the platform’s built-in agent walked him through the rest — suggesting tools, generating logic, producing test cases. The whole workflow, assembled through conversation.
What made this tangible wasn’t the speed. It was watching Jonny’s relationship with the interface change in real time. He stopped worrying about precision. Stopped correcting himself. At some point he noted, almost offhandedly, that he’d stopped caring whether he even spelled things correctly — because the system was understanding what he meant, not parsing what he typed.

That’s the tell. When you stop fighting the interface, you start solving the actual problem.
I pointed it out to him directly: what he’d done wasn’t prompting. He’d asked a question. The agent answered it by building something. That distinction — between instructing a tool and collaborating with one — is the thing most people haven’t fully absorbed yet.
Traditional automation requires you to specify how to do something, step by step. Vertesia Studio lets you describe what you want to achieve and works out the how. It’s the difference between giving someone a recipe and asking them to cook dinner.
12 Mismatches. Zero Missing Documents.
The reconciliation agent ran against the dataset and found 12 discrepancies between the invoice PDFs and the database records. Zero missing documents. Jonny had already done his own review, so he could verify the accuracy live. It matched.
The speed improvement is almost embarrassing to say out loud. Ten times faster than his previous workflow — which was already fast by any normal measure. A complex agent that would have taken weeks to build now takes a few hours at most. A production-ready, tested version: maybe a day.
The testing itself, which used to mean hours of manual work, now runs automatically across multiple models in parallel.
But the speed isn’t even the real point. The real point is what happens at scale.
At 50 invoices, a human can check the work. At 500, it’s painful but possible. At 50,000, you’re not choosing between fast and slow. You’re choosing between automated and broken.

The reality? This isn’t just a faster version of the old process. The old version of “automation” was an OCR system that maybe caught a field here or there — expensive, brittle, and still requiring a team of people to clean up after it. What we’re talking about now is different in kind, not just in degree.
A Scalable Team of One
This is the part I keep coming back to — not the speed numbers, not the accuracy rate, but what it means for the person sitting in front of the screen.
Jonny built something that most individuals simply couldn’t have built before. Not because the knowledge wasn’t available, but because the tools weren’t. And the tools that did exist required specialists to operate them. The gap between “I have an idea for a workflow” and “I have a working, tested, production-ready agent” used to be measured in weeks and headcount.
That gap is now measured in hours. For one person.
“As a single person, I have a scalable team that has a ton of knowledge in all domains of life — and can help me build out use cases that I typically wouldn’t be able to build.”
I’d add one thing. The team doesn’t just have knowledge — it has tools. Serious tools. And it knows how to use them. Knowledge without tools is a consultant with no budget. Tools without knowledge is software no one can configure. The combination is what changes the equation.
What’s Coming Next
Episode 8 showed what’s possible when the data pipeline works.
Episode 9 is about making sure it does. Before any agent can reconcile, analyze, or extract value from unstructured data, that data needs to be prepared in a way that makes it useful for AI reasoning. Not just readable. Useful. That’s what we call Semantic DocPrep — Vertesia’s patent-pending approach to converting complex documents into structured, semantic XML that LLMs can understand with high fidelity.
It’s one of the most important — and least discussed — challenges in enterprise AI deployment. We’re going to dig into it properly.
Past the Pilot is a series where Stefan Born and Jonny McFadden go beyond the proof-of-concept to explore what agentic AI looks like in real-world business contexts. New episodes drop regularly — subscribe so you don’t miss what’s coming next.
메타데이터
- post_id
- 28c8ae944b0f
- slug
- when-ai-meets-your-messy-data-28c8ae944b0f
- url
- https://medium.com/@sborn_63068/when-ai-meets-your-messy-data-28c8ae944b0f
- canonical_url
- https://medium.com/@sborn_63068/when-ai-meets-your-messy-data-28c8ae944b0f
- author_url
- https://medium.com/@sborn_63068
- status
- ok
- fetched_at
- 2026-07-08 02:40:31