The Voice Layer: The Channel Nobody Declares
Zero Data Retention covers your documents. It does not cover what happens before them.
The Voice Layer: The Channel Nobody Declares
Zero Data Retention covers your documents. It does not cover what happens before them.
By José Luis Caresani · Technology & AI Executive, Rational Core LLC
A lawyer finishes reviewing an AI agent’s output and dictates her observations into her phone. Or she verbally summarizes the markup to a colleague, and the meeting tool transcribes it.
That audio never touches the legal AI platform. It travels through whatever speech pipeline the device uses — often a cloud-connected one — before any of the platform’s privacy controls apply.
A reader of my Harvey implementation guide, Ari Nakos, pointed this out in the comments last week, and it is exactly the right observation: the voice layer sits upstream of everything the platform governs. The Zero Data Retention guarantee that firms negotiate so carefully in Phase 1 covers documents inside the platform’s perimeter. The dictation that happens before ingestion is outside it — and in most firms, nobody has made a single governance decision about it.
An undeclared channel is still a channel
In UCIF terms — the Unified Conversational Intelligence Framework I publish in open access — every data channel an organization uses should carry a declared Owner Schema: its legal basis, its visibility scope, its retention regime, its authorized purpose.
Matter records have one. Client communications have one. The document management system has one.
The voice layer usually has none. Which legal basis covers a lawyer’s dictated case notes processed by a third-party speech API? What retention applies to the audio? Was any purpose ever authorized?
Nobody decided — and that is itself a decision. An undeclared channel doesn’t stop existing. It just stops being governed.
When voice becomes evidence, the stakes change
There is a context where this stops being a hygiene issue and becomes a structural one: audio as judicial evidence.
The UCIF Conceptual Whitepaper devotes a specific section (15.5) to transcription in legal and judicial contexts, because the cost of error there is qualitatively different. Speech-to-text quality varies with language, accent, capture conditions, and background noise — and imperfect transcripts propagate upward as factual substrate for every inference built on top of them.
For evidence destined for a judicial or administrative body, the framework’s position is explicit: automated transcription — Whisper or any equivalent — must never run unsupervised. The process requires escalation to a qualified human transcriber who can review the original audio segments, equalize the sample when capture quality demands it, annotate the relevant metadata, and certify the integrity of the produced text without distorting the chain of custody.
The productive pattern is a hybrid: automated transcription as a first pass, professional human review focused on the critical segments, original audio preserved alongside the transcript, and full traceability — the system must be able to reconstruct, before a legitimate authority, exactly what it did with each piece of evidence, on what basis, under whose supervision, and with what certified result.
That slows down the cycle the sales deck describes as instantaneous. It is also a legal and ethical condition that cannot be waived: a judicial decision with serious consequences cannot rest on an unsupervised transcript whose imprecision, in any sensitive segment, could alter the substantive meaning of the evidence.
The general lesson
The voice layer is one instance of a broader principle. Governing an organization’s conversations includes the ones that travel through a microphone before touching any platform — the dictation apps, the meeting transcribers, the voice notes, the assistants listening in the room.
Platform guarantees govern the platform. Everything upstream needs its own declaration.
Every channel exists. The only question is whether it has been declared.

Channels Observations Governance Schemas

Extended Human in the loop
Resources
- UCIF Conceptual Whitepaper v1.0 (section 15.5 covers transcription in judicial contexts): doi.org/10.5281/zenodo.20579364
- UCIF Executive Whitepaper v1.0: doi.org/10.5281/zenodo.20702214
- How Harvey Is Implemented in a Law Firm (the guide that prompted the reader’s comment): doi.org/10.5281/zenodo.21043650
- Technical Note — Governing Legal AI Agents: Harvey as a Use Case for UCIF: doi.org/10.5281/zenodo.21041793
José Luis Caresani is an electronic engineer and Technology & AI Executive at Rational Core LLC. Co-editor of the ITU-T Z.410 standard, author of the Unified Conversational Intelligence Framework (UCIF), published open access on Zenodo. Independent consultant in Madrid, specialising in conversational governance and AI-First transformation.
Contact: joseluis.caresani@rationalcore.com · linkedin.com/in/jose-luis-caresani-127172
메타데이터
- post_id
- be24d9b2b6c2
- slug
- the-voice-layer-the-channel-nobody-declares-be24d9b2b6c2
- url
- https://medium.com/@caresanijose/the-voice-layer-the-channel-nobody-declares-be24d9b2b6c2
- canonical_url
- https://medium.com/@caresanijose/the-voice-layer-the-channel-nobody-declares-be24d9b2b6c2
- author_url
- https://medium.com/@caresanijose
- status
- ok
- fetched_at
- 2026-07-19 00:45:20