AI & context
7 min

How to Build an Ontology for AI

Key takeaways

  • An ontology built for agents has a harder job than one built for people, because nothing downstream catches a wrong answer.
  • Build from the questions an agent must answer correctly, not from the data you happen to hold.
  • Constraints are not optional here. An agent will act on a plausible answer as readily as a correct one.
  • The model has to be resolvable at query time, or the agent falls back to inference and the work is wasted.
In this article
Share

Most guides on how to build an ontology for AI are tutorials about modelling. Modelling is rarely where teams get stuck.

The question actually being asked, by someone in the middle of it: "What methods or softwares are there to get non-technical people to engage with ontology design and thrash out agreed definitions? I am beginning to think this is a major roadblock to more reliable AI. Without structured, verified knowledge managed by humans that can be safely inferred from, how is any business going to trust agents with anything important?"

That post drew more than three comments for every point it scored, which is the highest ratio of argument to approval anywhere in this corner of the internet. People are not reading about this. They are fighting about it. So this is a method for resolving a disagreement, and the modelling steps are the easy parts around it.

How to build an ontology for AI, in six steps

1. Start from the questions the system must answer. Write the competency questions down before anyone creates a class. Which accounts were active in Q3 under the revenue definition? If an agent will be asked it, it goes on the list; if nothing on the list traverses an entity, that entity is out of scope. This is the same discipline as scoping an ontology so it ships, and it is doing more work here than it looks like it is.

2. Get entities resolving to one stable identifier, before modelling relationships. This is out of order compared to most guides and it is deliberate. A practitioner who has done it: "the thing that moved the needle was less which ontology tool we picked and more whether the entity resolved to a single stable identifier at all. You can model beautifully, but if the same concept shows up under three slightly different names the model quietly treats it as three things."

The word quietly is the problem. Nothing errors. The mechanism is best named in a comment on one practitioner's write-up: "Apple Inc, Apple, and AAPL as three separate nodes because the LLM extraction never produced the same surface form twice." The same engineer audited his own generated graph and found roughly half of its reported nodes were duplicates: 1,892 reported against 911 actually unique. One person, one graph. Read it as the anecdote it is. The principle underneath it holds anyway. Relationships modelled on top of unresolved entities are relationships between the wrong things.

3. Draft with the model, then correct. Generation is genuinely good at a first pass and refusing it on principle just makes this slower. It is also measurably better at some parts than others, which tells you where to spend the correction.

The best published measurement compared generated ontologies against a human reference across three models. Property F1 landed between 0.12 and 0.27 in every configuration tested. Classes scored better, and far less consistently (the benchmark, CEUR-WS Vol-3953). Qualitatively the output was "a strong and well-structured foundation" but "shallow, flat", with "too many top-level classes and insufficient depth in the hierarchy." That benchmark used 2024-generation models. Nobody has published a rerun. If you think today's models do better on properties, you may well be right. It is simply untested.

Either way the shape of the correction is the same: take the classes, distrust the relationships.

4. Run the consensus step, and treat it as the real work.

This is the step the rest of the method exists to support, and it is the one every competing guide omits.

Put instances and counter-examples in front of people, never class diagrams. A domain expert asked to review a UML-ish diagram will say it looks fine, because the diagram is not in their language. The same expert shown twenty actual records and asked which of these are active customers will argue immediately and productively. Disagreement is the output you want at this stage; it is cheaper now than after an agent has acted on it.

Drive the argument with edge cases. Definitions come out of them. Nobody disagrees about the obvious instances. Bring the account that churned and came back, the trial that was cancelled mid-enrolment, the order billed to one entity and shipped to another. The boundary is where the definition actually lives.

Name who breaks a tie before you need one. A person. A committee cannot be asked why. The absence of a named tiebreaker is why these sessions end in another session.

Record the agreement at the moment it happens. What was decided, who decided it, what it replaced, and when it should be looked at again. A review trail reconstructed next quarter from memory and meeting notes is no review trail at all. Whether it exists is the difference between a model and a shared opinion.

One thread asking how to get experts to consensus landed on an answer worth noting: buy a platform. That is a real answer and sometimes the right one, but it relocates the problem. The same reply contained the more useful sentence: "The upper management took YEARS to understand that this graph would not self organise itself."

5. Constrain the agreement so it is enforceable. An agreement nobody can check is documentation. SHACL shapes turn each decision from step 4 into something a system can refuse to violate, and writing them surfaces the disagreements that were papered over.

6. Wire it to the consumer so retrieval resolves against the definitions. An ontology sitting beside your retrieval stack does nothing on its own. The test is what happens when the agent reaches for active customer. It should resolve against the governed definition. Whatever the index happened to return is not an answer. Otherwise you have built a model and an unrelated AI system that occasionally agree.

What to do at each turn, as standing practice

What the agent doesWhat the model must supplyWhat happens if it does not
Resolve an entity from a user's phrasingIdentity across systems, with agreed aliasesIt picks a plausible candidate and acts on it
Apply a business ruleThe rule as a checkable SHACL constraintIt reasons the rule from examples, differently each time
Traverse to related factsTyped relationships, not inferred joinsIt answers a near-neighbour question instead
Report a numberThe definition that produced it, and who agreed itThe number cannot be defended when challenged
Refuse an unsafe actionConstraints that fail closedThe action is plausible, so it proceeds

One decision at a time. The second contested definition starts after the first is in production. Never alongside it.

Reuse a published model where one fits. Most core domain concepts have been modelled competently already, and cutting an existing model down is faster than authoring from nothing.

Record agreement at the moment of agreement. Repeated because it is the one that gets skipped under deadline, and the only one that cannot be fixed later.

What the model has to stay connected to

  • The agent framework, at query time, over an interface it can call mid-request without a human in the loop.
  • The data, through SHACL shapes, so an output can be validated.
  • The people accountable, so each definition carries a review trail: who proposed it, who approved it, when.
  • The other models in use, because an agent crossing two domains inherits both, and a disagreement between them surfaces as a confidently wrong answer.

Express all of that in RDF, OWL and SHACL. A framework's own format locks the model to that framework. Frameworks in this field change roughly annually. The model is ground truth. The agent stack is this year's way of consuming it. Keeping those straight is the difference between an asset and a dependency. The ongoing practice is ontology management on a platform built to activate it into applications and agents.

What goes wrong

Generated structure is inconsistent across runs. The clearest statement of it: "If you ask a model to make a knowledge graph, it will do a relatively poor job of extracting everything (low recall), then it will make each relation and type on the fly, arbitrarily. Run it again on the same text and you will get different names for everything. If the terms used are unpredictable, then your KG is of little use to anyone else." Academic work corroborates it: models "refer to entities in slightly different ways, even when presented with the same text, prompt, and instructions", returning "New York," "New York City," and "NYC" for the same sentence.

Deterministic retrieval does not give you deterministic answers. Worth knowing before you promise anyone reproducibility: "Ran the same question 3 ways against a knowledge graph. Retrieved the same 90 entities and triples each time. LLM output still varied. That's the finding." The graph did its job. The generation step is still the generation step.

The hardest objection deserves an answer. One version, posted and never answered: "So strange to outright accept ontology first when AI is all about emerging abstractions in embedding space. AlphaGo for chess also showed that training a model from scratch yields better results than training from human notions of what are considered good play."

It is a serious point. The narrow answer is that learned abstractions are excellent at finding structure and have no mechanism for commitment. A system that discovers a better categorisation of customers has not thereby decided the company now uses it. In a regulated report, the question is never which categorisation is best. It is which one the organization is standing behind.

The rest of this is upkeep, and upkeep is ontology management. If you want the Foundry sense of the word, that is a different thing and worth understanding separately.

So what should I take away?

My agent will act on a plausible answer exactly as fast as a correct one. Nothing downstream catches it. A dashboard has a person in front of it who notices when a number looks wrong. An agent has nobody.

So the constraint is the product. An ontology built for agents has to refuse the wrong answer before it reaches anything. That refusal has to be written in a standard my next tool can still read. TopQuadrant builds exactly that: the agreed model, the named owner, and the shapes that enforce both.

About the Author

Give your AI the context it's been missing

See how the TQ Data Foundation turns your enterprise knowledge into trusted, Al-ready context.