Nobody Reads Eleven Million Files. They Query a Graph.
In early 2015, an anonymous message reached Bastian Obermayer, a reporter at the German newspaper Süddeutsche Zeitung: “Hello. This is John…
Nobody Reads Eleven Million Files. They Query a Graph.
In early 2015, an anonymous message reached Bastian Obermayer, a reporter at the German newspaper Süddeutsche Zeitung: “Hello. This is John Doe. Interested in data?”
Obermayer said yes. The source insisted on encrypted communication, refused every request for a face to face meeting, and warned that being identified could cost him his life. What arrived, over the following months, was 2.6 terabytes of internal records from a Panamanian law firm called Mossack Fonseca: 11.5 million files. Incorporation papers, emails, memos, contracts, spreadsheets, scanned passports, shareholder registers, decades of shell companies built in the British Virgin Islands, in Panama itself, in a couple of dozen other jurisdictions built for exactly this purpose.

On 4 March 2015, Obermayer and his colleague Frederik Obermaier brought the leak to Gerard Ryle at the International Consortium of Investigative Journalists, ICIJ. Ryle understood something immediately: no single newsroom could read eleven and a half million files. So he did not build a newsroom. He built a network. More than 370 journalists, in more than 100 media organisations, across 76 countries, working in secret for just over a year.
Marina Walker Guevara, ICIJ’s deputy director at the time, told everyone involved exactly how fragile that arrangement was. “You have no business talking about this work with your spouses, with your friends, with your parents,” she told them. Reporters shared notes, interviews and documents with rivals from other countries, sometimes chasing a lead that would end up published under someone else’s byline entirely. It would have taken one person breaking confidentiality for the whole collaboration to crumble.
I have read a lot about how that investigation actually worked, because the quieter version of the same problem is the one I have spent my career on. Almost nobody talks about that part of the story.
Eleven million documents is not eleven million facts
Mar Cabra ran ICIJ’s Data and Research Unit during the investigation. In a talk she gave a few months after the story broke, she described the real problem facing her small team of developers. It was not the volume. Reporters had handled large leaks before. It was that any single document, read alone, told you almost nothing.
One scanned file might show a company incorporated in the British Virgin Islands. Another might show that company’s shareholder, a name typed slightly differently by a clerk in Panama City. A third might show that shareholder’s home address, an address shared with somebody else entirely. Read one at a time, these are three unrelated scans. Read together, they are a person hiding behind a company hiding behind another company.
Cabra’s team had already learned this lesson once. In 2013, ICIJ built something called the Offshore Leaks Database, a public search tool for an earlier, smaller leak, running on a plain MySQL database with a graph visualisation library called Sigma.js bolted on top. It became the most visited thing the organisation had ever published, precisely because ordinary readers could finally see a name connect to a company rather than sift through a folder of scans on their own. The lesson stuck: a pile of documents is an archive. The same documents, connected, are an investigation.
What the graph actually was
For the Panama Papers, Cabra’s team rebuilt Mossack Fonseca’s own internal records, more than three million files on their own, into a graph database, Neo4j, with a visualisation layer called Linkurious sitting on top. Every company, person, address, shareholder and intermediary became a node, kept as it appeared in the source document. Every connection between them, a shared address, a shared signature, a matching identifier, became an edge.
The finished graph held around 950,000 nodes and 1.2 million edges, at roughly four gigabytes. Cabra pointed out how small that sounds next to 2.6 terabytes of raw files, and that was rather the point. The graph was not the data. It was what the data meant once somebody had connected it.
Reporters who had never written a line of query code could click a dot and watch it expand into everything attached to it. Cabra described asking whether her own name connected to the then British prime minister, David Cameron, and getting an answer back in one click, through a feature the team used for finding the shortest path between two nodes. A Finnish reporter found a cluster of her country’s Mossack Fonseca clients that she had missed entirely in the raw documents, surfaced only because the matching could handle names that sounded alike without being spelled alike. Cabra called this “fuzzy searching”. I would call it probabilistic matching running alongside deterministic matching, which happens to be close to the discipline I have built a working life around. Some connections are exact: a shared tax identifier, a verified account number. Some are close enough that only a method built for real world variation should be trusted to draw the line, and a system that only trusts exact matches will miss most of them.
The graph is why the story broke
On 3 April 2016, ICIJ and its partners published at once. The leak named 140 politicians and public officials around the world, including 12 current or former heads of state: among them Iceland’s prime minister, Ukraine’s president Petro Poroshenko, the king of Saudi Arabia, and the family of Pakistan’s prime minister, Nawaz Sharif. Public figures ranging from the footballer Lionel Messi to sitting officials appeared too, all surfaced through the same law firm’s records.
Two days later, Iceland’s prime minister, Sigmundur Davíð Gunnlaugsson, resigned. Years earlier, he had co-owned an offshore company, Wintris Inc., registered in the British Virgin Islands, with his wife, Anna Sigurlaug Pálsdóttir, before selling her his half at the end of 2009. The company held bonds worth millions of dollars in Iceland’s own collapsed banks. Failing to disclose that when he entered parliament violated the country’s parliamentary ethics rules and left him sitting on an undisclosed conflict of interest for years. Nobody found any of that by reading a random file out of eleven million. Somebody found it by expanding a dot.
That is the part of this story I keep coming back to. The leak was the material. The graph was the investigation. Without it, eleven million documents just sit there, technically searchable and practically illegible, because the question a reporter actually needs answered is never “what does this one file say”. It is “who is this person connected to, and how do I know”.
The part everyone underrates is the querying
Building the graph was the hard, expensive, unglamorous part of this story. But a graph nobody can question is only a more elegant archive. ICIJ’s reporters did not find Gunnlaugsson’s company by staring at a database schema. They found it because the system let them ask, in effect, who is connected to this person, and it returned an answer built from traceable evidence rather than a hunch.
That distinction is the one I think about most in my own work, on a far less dramatic scale. A support agent trying to work out if two accounts belong to the same household. A fraud team trying to tell whether two applications share a real connection or a coincidence. The stakes are smaller than a leaked law firm’s client list, but the underlying question is identical: is what I am looking at one thing, recorded incorrectly in several places, or genuinely several things that merely resemble each other.
That question only stays answerable if somebody, or something, can ask it and get back a reason, not just a match. Every fragment arriving from a messy source system has to be kept exactly as it was recorded, never quietly overwritten. Every connection drawn between fragments, whether from an exact identifier or from a method reading through the ordinary noise of real names and real addresses, has to carry the reason it was drawn. Only then do enough of those connections resolve into one entity that a person, or a system acting on a person’s behalf, can actually interrogate, rather than merely store.
A plain keyword search over the same eleven million files would have returned plenty of hits too. Search Gunnlaugsson’s name and you would find him. What search cannot do is start from him and walk outward, one honest edge at a time, until a company he had quietly sold his half of years earlier, still tied to him through a spouse and a shared address, comes into view on its own. That walk is the entire difference between finding a document and understanding a person, and it is only possible once the fragments have already been resolved into something you can ask questions of.
The eleven million files were never the story. The graph was, and only because somebody could ask it a question.
Steven Renwick is the co-founder and CEO of Tilores, which provides real-time entity resolution through an API for AI and data teams, combining deterministic and probabilistic matching so that scattered records resolve into one entity worth trusting.
Sources
-
Mar Cabra, “How the ICIJ Used Neo4j to Unravel the Panama Papers,” Neo4j Blog, May 12, 2016: https://neo4j.com/blog/cypher-and-gql/icij-neo4j-unravel-panama-papers/
-
Roberto V. Zicari, “How the 11.5 Million Panama Papers Were Analysed: Interview with Mar Cabra,” ODBMS Industry Watch, October 11, 2016: https://www.odbms.org/blog/2016/10/how-the-11-5-million-panama-papers-were-analysed-interview-with-mar-cabra/
-
ICIJ, “The Story That Rocked the World: Ten Years of the Panama Papers, Part 1”: https://www.icij.org/investigations/panama-papers/the-story-that-rocked-the-world-ten-years-of-the-panama-papers-part-1/
-
ICIJ, “About the Investigation: The Panama Papers”: https://www.icij.org/investigations/panama-papers/about-the-investigation/
-
ICIJ, “Giant Leak of Offshore Financial Records Exposes Global Array of Crime and Corruption,” April 3, 2016: https://www.icij.org/investigations/panama-papers/20160403-panama-papers-global-overview/
-
ICIJ, “Iceland Prime Minister Tenders Resignation Following Panama Papers Revelations,” April 5, 2016: https://www.icij.org/investigations/panama-papers/20160405-iceland-pm-resignation/
-
“What an Entity Resolution Graph Looks Like and How It Is Queried”: https://tilores.io/content/what-entity-resolution-graph-looks-like-how-queried
메타데이터
- post_id
- f54daeafe46f
- slug
- nobody-reads-eleven-million-files-they-query-a-graph-f54daeafe46f
- url
- https://medium.com/@major-grooves/nobody-reads-eleven-million-files-they-query-a-graph-f54daeafe46f
- canonical_url
- https://medium.com/@major-grooves/nobody-reads-eleven-million-files-they-query-a-graph-f54daeafe46f
- author_url
- https://medium.com/@major-grooves
- status
- ok
- fetched_at
- 2026-07-27 22:17:44