← Back to list

Semantic Data Engineering: People First, Platform Last

How to Make Business Definitions Survive Platform Change, Enterprise Complexity, and Agentic BI

Srini Rao S · 2026-10-01 16:05 · 12 claps · 12.5 min read
#data-strategy #semantic-layer #agentic-ai-architecture #data-governance #enterprise-ai
Open on Medium ↗
Wiki topics: AGT · AI Agents CRY · Crypto & Web3 🔧 · Data Engineering 🏛️ · Architecture

Semantic Data Engineering: People First, Platform Last

How to Make Business Definitions Survive Platform Change, Enterprise Complexity, and Agentic BI

Why Agentic AI Needs Agentic BI First — and Neither Survives on a Definition That Only Lives in One Person’s Head

PART 1 · Quick Problem Statement

Every enterprise chasing agentic AI is already running Agentic BI experiments — agents that investigate, query governed data, and recommend an action instead of just answering a question (1). It’s the cheapest place to find out whether the semantic foundation actually holds, before anyone hands an agent a P&L decision.

It fails on the same problem BI has always had, just faster. Two business units call the same figure “revenue” and calculate it differently, because the rule for why a “Payer” differs from a “Bill-To” lives in one controller’s head, not a system.

That’s almost never a data problem first. It’s a people problem wearing a data costume — and BCG’s own 10–20–70 rule backs it up: 70% people and process, 20% technology and data, just 10% algorithms (2).

An agent doesn’t fix that split. It just fails at inference speed, confidently, before anyone notices.

When the sanctioned path is too slow, the business builds its own: 71% of employees already use unapproved AI tools at work, more than a quarter of them because no sanctioned alternative was ever offered (3) — which is why the sequence below starts with the business problem, not the platform.

PART 2 · The Roadmap — Current State to Target State

Most organizations open with a platform decision. That’s backwards, and it’s the single most common reason semantic layer and Agentic BI pilots stall after the demo: technology is the last of four readiness dimensions, not the first. What follows assesses today’s state and the target state, in the order they actually need solving — people, then process, then definitions, then technology.

DIMENSION 1 · People Readiness

Situation

– Ownership defaults to IT. No named business owner exists per domain, so a data engineer ends up guessing at intent for entities like Customer, Product, Cost Object, or GL Account.

– Knowledge is tribal. The reason a “Payer” differs from a “Bill-To” lives in one controller’s head — and leaves when that person does.

– Shadow IT fills the vacuum. With no named owner and no fast, sanctioned path to a new report, the business builds its own — spreadsheets, unapproved SaaS, now unapproved AI tools — not out of malice, but because central IT can’t keep pace.

– Conflicts resolve by default, not by design. Whichever system loaded last, or whichever integration ran first, quietly wins — and no one signed off on it.

Task

Give every domain a named owner and a documented decision trail, so institutional knowledge outlives the person who holds it.

Action

– Name domain-aligned owners, not just IT — every certified entity needs a business owner who can arbitrate disputes, plus a steward accountable for day-to-day quality and documentation.

– Govern where people already work — a Community of Practice, not a portal. OpenMetadata’s community shows the ceiling: 400-plus contributors, 11,600-plus users, inside a shared Slack — and Carrefour Brazil crowdsourced 300-plus glossary terms the same way (4).

– Let AI draft, let a human approve — an AI data steward proposes a definition for business SMEs to correct, escalate, or approve, cutting manual curation effort by roughly 60% without a dozen-plus full-time steward team (5).

– Treat modeling as its own discipline. Semantic and ontology modeling blends business analysis with data modeling; most data teams haven’t staffed for it and default to whoever knows SQL best.

– Define the data objects, not just the org chart. For Vendor, Contract, People/Security, Customer, Cost Object: decide the attributes, relationships, and properties that actually exist, and whether they’re enough to answer the business questions the semantic layer needs to serve.

– Set survivorship rules, with a business SME as final authority. When the same vendor or contract disagrees across source systems, decide in advance which one wins and why — signed off by the owner who understands the consequence, not whichever system loaded last.

– Let a supervisor agent handle routine arbitration as governance matures — the same orchestration pattern that already resolves conflicting outputs across multiple AI agents runs the routine calls automatically; only what falls outside its confidence threshold escalates (6) (7). The SME’s job shifts from adjudicating every conflict to owning the policy (8).

– Look to regulated-industry leaders for proof this compounds — Forrester points to Bank of New York as one of the furthest along in agentic AI deployment inside heavily regulated banking — not for better technology, but because it invested in workforce readiness to manage highly autonomous systems, a competitive advantage most enterprises still lack (25).

Result

Institutional knowledge becomes documented policy instead of tribal memory — it survives the reorg, the departure, and the platform migration.

DIMENSION 2 · Process Readiness

Situation

– No standing change process. Shadow spreadsheets reappear within a quarter of go-live because nobody owns how a new metric or attribute gets proposed and approved.

– The model only ever grows. Old definitions are never retired, so the glossary becomes as unreliable as the spreadsheets it replaced.

Task

Build a process that can approve a change fast, and let the model evolve with actual usage — not governance reviews alone.

Action

– Stand up a standing definition-change process — who can propose, who approves, and how fast — with a lightweight extension path for fast-moving segments (sales territories, product lines, promotions) instead of a quarterly release train.

– Let AI draft, not just govern — “vibe modeling”: AI drafts and refines a definition while a practitioner states the intent and reviews the result, the same shift “vibe coding” brought to code, compressing days of drafting into hours (9). Every AI-drafted definition still passes the same certification gate as a human-drafted one.

– Let the model evolve with usage — an agent that hits a missing metric mid-conversation can derive it, answer immediately, and queue the addition as a reviewable change a human approves or reverts (22).

– Invest in change management, not just training — a clear vision, honest communication, and peer “Change Champions” instead of a top-down mandate (10).

Result

One rollout that kept the model in step with usage cut managed semantic assets 75% and lifted self-service adoption 3.6x (23).

DIMENSION 3 · Governance Readiness

Situation

– KPIs aren’t actually shared. A long-standing, still-cited MIT Sloan Management Review benchmark found only 26% of senior leaders strongly agreed their KPIs aligned to strategy (11).

– Harmonization is assumed, not done. Two business units calling the same metric “revenue” but calculating it differently is the default state, not the exception, especially after a merger.

– Territory and entity churn outpaces the model. Sales/marketing realign by geography or segment, and divestitures or M&A happen, far more often than most semantic models are built to absorb without a rebuild.

Task

Certify every data product against six checks — as a score that travels wherever the data goes, not a one-time pass/fail at launch.

Action

Every data product feeding the semantic layer — and every agent reading it — is certified, regardless of which cloud, platform, or ecosystem it physically lives in:

✓ Lineage — traced end-to-end from source system to the field an agent or dashboard actually reads.

✓ Quality — measured against a published SLO, not asserted once at launch and never checked again.

✓ Domain-mapped — registered against a business domain in the ontology (Customer, Cost, Product, Asset), not left as an orphaned schema in a catalog.

✓ Domain-aligned owner and steward — the same two-role split People Readiness requires.

✓ Data contract in place — schema, quality thresholds, semantic meaning, freshness SLA, and access policy, in one document — the same one an agent’s access is granted against (12).

✓ Architecture, compliance, and classification — aligned to enterprise data-architecture standards and tagged for sensitivity, PII, and regulatory scope.

Certification travels with the data, not the platform. Every major cloud and data platform already ships an open way to share a certified product with the others — Example 1, below, walks through what that looks like across a real four-platform estate. None of this is new: it’s the decades-old discipline of a canonical data model and a business artifact, just applied to agentic AI instead of BI (13).

– Treat the first deployment as a governance pilot, not an ROI pilot — especially in regulated functions like finance: Gartner’s own guidance to CFOs is that early AI agent pilots fail from unclear controls, not poor technology, so success should be measured by control consistency and full traceability, not autonomy or speed (24).

Result

Governed decisions are five times more trusted and 80% faster than ungoverned ones (14) — and most AI-related breaches trace back to missing access controls or a governance policy nobody finished writing (15).

DIMENSION 4 · Technology Choice

Current State

– The platform conversation starts on day one — before anyone has agreed what a “customer” or a “cost” even is. That’s the classic mistake this roadmap exists to prevent.

Target State

Once people, process, and definitions are ready, the technology decision is mostly about where the semantic layer should physically live relative to the business and the data it serves. Four patterns cover most enterprise choices:

– Fuzzy matching is the mechanism that scales certification — the “domain-mapped” check doesn’t stay a manual tagging exercise once catalog tooling does the first pass: score candidate matches on a similarity threshold and a steward approves rather than authors every mapping by hand (16). That’s how “domain-mapped” holds at enterprise scale instead of only on the first hundred tables.

– Generative BI is how you scale the scarce expert, not just outlaw shadow IT — natural-language analytics lets a business user ask a governed question directly instead of building a workaround; only 25% of employees ever adopted traditional BI tools, and Generative BI frees the scarce expert for strategic work while everyone else self-serves (17). It only holds up on the governance this roadmap builds — without it, natural language just gets you a confidently wrong answer faster.

“Gen BI is only as trustworthy as the data behind it.” (18)

– Agentic BI takes this further, and is where it gets tested first — before an autonomous agent gets a financial or customer-facing decision, it almost always gets an analytics one first. Its promise — planning an investigation and recommending an action — is only as trustworthy as the certified metric layer underneath it. Get this right, and Agentic AI inherits a foundation, not fragmentation.

PART 3 · Examples from Practice

EXAMPLE 1 · One Enterprise, Four Platforms, One Certified Layer

Most enterprises never get a clean-slate platform choice. A more common reality: SAP S/4HANA carries Finance and Supply Chain, an acquired division runs on Oracle ERP, and the operational edge — labs, plant-floor MES, quality, contracts, IT service management — was never going to live inside an ERP. All four platforms from Dimension 4 end up in play at once, each the natural home for a different domain’s system of record:

What changed in the last year: none of these four have to stay behind a copy anymore — all four converged on open, zero-copy sharing as the mechanism (19) (20) (21). Connectivity is solved; trust in what crosses it is not — that’s what certification buys.

EXAMPLE 2 · Customer 360: Ship-To, Bill-To, Payer, and Business Partner

Annotation — this is the textbook business-artifact pattern from Part 2 — one identity, one lifecycle, multiple roles — not an ad hoc modeling problem.

A single commercial relationship routinely spans four partner roles — Sold-To, Ship-To, Bill-To, and Payer — each potentially a different legal entity, all sitting on one underlying Business Partner record. A “Customer 360” entity has to decide, and document, which role answers which question: sales wants Sold-To, collections wants Payer, compliance wants the Business Partner’s legal entity. That’s a call made once, in the open — not a shortcut made by whichever team builds the pipeline first.

EXAMPLE 3 · Cost Object Fragmentation and GL-to-Cost-Element Mapping

Annotation — also a business artifact — an Internal Order or WBS Element carries its own identity and lifecycle from open to settled, independent of which system posts it.

Finance and Controlling see the same spend differently by design: an Internal Order tracks a temporary activity, a WBS Element a capital project, a Cost Center a standing unit — three lifecycles, all valid at once. Each posts back to a GL Account, which Controlling re-expresses as a Cost Element: the same P&L account, maintained in two places on two change cycles. Certify “cost” without reconciling GL-to-Cost-Element first, and Finance and Controlling both distrust the number — for different reasons.

EXAMPLE 4 · The Join the AI Got Wrong

Annotation — an illustrative composite, not a documented incident — but this exact failure mode (plausible on a spot check, wrong at the grain) is common enough to be worth naming.

A finance Community of Practice needed a “Subscription ARR” metric and didn’t want to wait two weeks in the backlog, so an analyst vibe-modeled it herself — describing the calculation in plain language and letting AI draft the SQL. The AI got the grain wrong: it joined at the contract level instead of the subscription level, which would have double-counted every multi-year deal. The totals looked right in a spot check — off by less than 2%, easy to wave through — and it would have shipped clean if the steward hadn’t run it through the same six-point certification gate as any other data product. Lineage tracing caught the bad join before quality scoring even ran; the fix was an hour, not a rewrite. That’s the actual case for certifying AI-drafted definitions the same as human-drafted ones — not that AI gets the logic wrong more often, but that when it does, the mistake is the kind that passes a glance.

PART 4 · Recommendations

– Start with the business problem, not the platform. Write down what breaks today — the definition dispute, the reconciliation, the rebuild — before evaluating a single tool.

– Certify the definition, not the platform. Lineage, quality, domain mapping, ownership, a data contract, and classification travel across clouds; consolidation doesn’t buy anything certification doesn’t already buy.

– Put governance where people already work. Stand up Communities of Practice inside Slack or Teams, not a separate portal, and let AI draft the first version of a definition for a human steward to approve.

– Use AI to scale the boring parts, keep a human at the gate. Vibe-modeled definitions and fuzzy-matched domain mappings compress weeks into hours — but every one of them still passes the same six-point certification check as work done by hand.

– Give the business a sanctioned alternative before you take away the unsanctioned one. Shadow IT and shadow AI exist because official channels were too slow; Generative BI on top of a certified semantic layer removes the reason to build a workaround instead of just banning the workaround.

– Prove it on Agentic BI before you bet on Agentic AI. An analytics agent that plans an investigation and recommends an action is the cheapest, most reversible place to find out whether the semantic foundation actually holds.

– Augment generative AI with human talent, not instead of it. The fastest, most durable gains come from pairing AI’s speed with the judgment your best people already carry — treat this as a talent-acceleration strategy, not a cost play, and the outcomes compound instead of plateauing.

– Sequence it, and don’t skip a step. People, then process, then definitions, then technology — reversing the order is the single most common reason these programs stall after the demo.

Customer 360, cost transparency, and every agentic use case built on top of them are only as trustworthy as the readiness work underneath. Sequence it — people, then process, then definitions, then technology — never the reverse.

Note: This is not a theoretical framework. It is a Value Delivery Roadmap shaped by practical enterprise delivery experience and tested against leading research on AI transformation, semantic governance, and decision intelligence. References below;

  1. Holistics — What Is Agentic BI? (Huy Nguyen), June 2026.

  2. BCG — Agentic AI Strategy for CIOs and CTOs in 2026 (Mark Abraham, Neveen Awad), July 2026.

  3. Microsoft / Censuswide, cited via Certero — shadow AI adoption survey, October 2025.

  4. Collate — 2026 Predictions: Why Semantics Will Determine AI Success (Suresh Srinivas), January 2026.

  5. Enterprise Knowledge — Crowd-Sourcing Data Governance.

  6. Kore.ai — Supervisor Agent documentation (AI Agent Platform).

  7. CluedIn — Agentic MDM vs. Traditional MDM Data Stewardship, July 2026.

  8. eWeek — How Will Agentic AI Change Enterprise Data Management in 2026 and Beyond? (Corey Noles), September 2025.

  9. Domino Data Lab — What Is Vibe Modeling? Code Less, Analyze More (Etan Lightstone), July 2025.

  10. Method — AI Adoption: Data Governance & Change Management (Jared Purcell, Ben Stagg), March 2026.

  11. MIT Sloan Management Review — Only 26% of Leaders Say KPIs Align to Strategy.

  12. Atlan — Data Contracts for AI: Why They Matter More Than Ever, July 2026.

  13. IBM Research RC24282 — Bhattacharya et al., Towards Formal Analysis of Artifact-Centric Business Process Models, 2007.

  14. Gartner — Top Trends for Data and Analytics, 2026 (Carlie Idoine), June 2026.

  15. IBM — Cost of a Data Breach Report 2025 (Ponemon Institute), July 2025.

  16. Pentaho — Manage Metadata Similarity Suggestions (Data Catalog documentation).

  17. IBM — What Is Generative Business Intelligence?, October 2024.

  18. AtScale — What Is Generative BI?, updated January 2026.

  19. Databricks — SAP Business Data Cloud Connect for Databricks (GA).

  20. Snowflake — Polaris Catalog: Open-Source Catalog for Apache Iceberg.

  21. Microsoft Fabric Blog — Zero-Copy Access to OneLake Data in Azure Databricks (Preview).

  22. Lightdash — Your Semantic Layer Can Now Fix Itself.

  23. Databricks — Beyond Tables: How Musinsa Built a Self-Optimizing Semantic Layer (Data + AI Summit).

  24. Gartner — Gartner Says CFOs Must Pilot Governance First Before Scaling AI Agents (Alex Levine), August 2026.

  25. Forrester — The State of Agentic AI in 2026: Companies Are Chasing, Few Are Catching.


메타데이터
post_id
f10fd6c3c736
slug
semantic-data-engineering-people-first-platform-last-f10fd6c3c736
url
https://medium.com/@sree622/semantic-data-engineering-people-first-platform-last-f10fd6c3c736
canonical_url
https://medium.com/@sree622/semantic-data-engineering-people-first-platform-last-f10fd6c3c736
author_url
https://medium.com/@sree622
status
ok
fetched_at
2026-10-03 21:39:07