Data governance
6 min

Semantic Layer, Data Fabric, Data Mesh, Context Layer: Is Any of This New?

The substrate is recycled. The obligation is new.

In this article
Share

Every few years the industry produces a new name for describing what data means, and every few years a portion of the people who have to implement it point out that they have seen this before. They are largely right, and the honest thing to do is start there rather than argue with it.

From a thread on r/dataengineering that ran to several hundred points:

I'm getting really tired of all the buzz words. Semantic layer vs Ontology. imo it's the same thing. you assign meaning to your data.

That is a fair reading of the situation, and anyone writing about this owes an answer to it rather than a reframe.

What in this is genuinely recycled?

A great deal, and it is worth being specific rather than gesturing at continuity.

Modelling meaning separately from storage is old. Describing entities, attributes and relationships in a layer above the physical schema, so that the description survives the migration: that idea predates every term in the title of this article by decades. Conceptual and logical modelling were doing it before the semantic layer had a name.

Federating rather than consolidating is old. The insight that you can leave data where it lives and resolve meaning across it, instead of moving everything into one store first, has been rediscovered at least three times. Data fabric's version of it is not materially different from what distributed query and federated architectures were proposing well before.

Describing something once and reusing the description is old. That is the central claim of the semantic layer, and it is also the central claim of the glossary, the catalog, the dimensional conformed dimension, and the master data programme. The pattern keeps returning because it keeps being right, not because it keeps being new.

Domain ownership of definitions, the part of data mesh that stuck, is a reorganization of who does the work rather than a discovery about what the work is. Which is not a criticism; organizational answers are often the missing piece. But it is not a new capability.

A commenter in that same thread put the relationship between two of these terms more precisely than most published material manages:

It's a running joke amongst ontologists that they desperately need an ontology of ontology. Marketing people have done their terrible work, and now the term can have different meanings depending on context. Your ontology should be a specification and documentation, and your semantic layer is an implementation built from that specification.

That is correct, and it is the kind of distinction that gets flattened every time a new label arrives. So: the substrate is recycled. Anyone telling you the modelling underneath this generation is fundamentally new is selling something, and the reader who suspected as much was right to.

Then what is actually different?

One thing. Not a list: a single change, and everything else follows from it.

Every previous consumer of these layers was a person.

A semantic layer fed dashboards, and the dashboards were read by analysts. A data fabric fed queries, and the queries were written by people who knew what they were asking. A mesh served domain teams who understood their own domain's conventions. In all of those cases, a wrong or ambiguous definition met a human being who could apply judgement to it: who knew their department counted customers differently, who noticed that the number looked off, who asked.

That judgement was load-bearing infrastructure, and nobody counted it as infrastructure because it was free.

An agent applies none of it. Handed a definition, it does not weigh it against experience or notice that it contradicts something it saw yesterday. It takes the definition, acts, and produces a result formatted exactly like a correct one. Then it does that at volume, on behalf of multiple teams at once, answering the same question differently depending on what it happened to retrieve.

So the question the layer has to answer has changed. It used to be what does this mean, which a good description answers. It is now what is this system allowed to treat as true, which a description does not answer at all, because the second question is about authority, and a description has none.

Why that changes what the layer has to carry

Once the consumer cannot apply judgement, the layer has to carry the things judgement was supplying.

Who agreed this. Not who wrote it. A description by an engineer who read the schema is a reading; a definition signed by the person accountable for the term is a decision, and only the second one settles a disagreement.

Who owns it now. Ownership is what gives the definition a change path and an address for the question that follows it.

When it stops being true. Definitions expire, usually silently, usually at an organizational event nobody connected to the data layer.

What happens when it changes. What re-validates, what auto-updates, and what gets routed to a person.

None of those four is a modelling problem, which is why the modelling being recycled is not the objection it first appears to be. They are governance properties, attached to a model whose substrate is indeed decades old. The model is the same shape it has been. What is bolted to it is different.

The objection that deserves a direct answer

The sharpest version of the sceptical case is not about terminology at all, and it showed up in the same discussion:

Nobody cared about Knowledge Graph before, and with modern LLM, they couldn't care less now: just feed the questions to an LLM and it will give them a probabilistic correct answer.

This deserves a straight answer rather than a deflection, because for most questions it is correct. A probabilistic correct answer is a genuinely good product. For search, for summarization, for drafting, for the large majority of what people ask systems to do, probabilistic is the right engineering trade and structure would be wasted effort.

The answer is that a minority of questions cannot accept a probabilistic answer, and it is not a minority defined by importance: it is defined by consequence of error. Whether a customer is eligible for an offer. Whether a safety case starts a seven-day clock or a fifteen-day one. Whether revenue is recognized this quarter. These are not harder questions; some are trivially easy once you know the rule. They are questions where being right 90% of the time is not 90% as good as being right, because the 10% is a finding, a misstatement, or a decision that has to be unwound.

For those, an answer needs to be reproducible and defensible, which means resting on something agreed rather than something inferred. The rest can and should stay probabilistic. Knowing which questions live on which side of that line is most of the work, and it is also the thing none of the previous layers had to ask, because a person was standing between the layer and the consequence.

That distinction is what context management is for, and it is the narrowest claim available: not a new modelling paradigm, not a replacement for anything in the title of this article, but an answer to a question that only becomes urgent when the reader of your definitions is no longer a reader. If you want the full version of the argument, it is here. If you want the specific case for why a derived definition cannot do this job, that is here.

Give your AI the context it's been missing

See how the TQ Data Foundation turns your enterprise knowledge into trusted, Al-ready context.