AI & context
3 min

What Is Runtime Context?

  • Most of the context your AI runs on is computed while it runs rather than looked up. That is correct design, not a gap.
  • Runtime context is an estimate, and estimates can be excellent. The failure is using one to answer a question that cannot accept one.
  • Every runtime inference sits on top of something that was never inferred, and that layer underneath is where the wrong answers start.
  • Two people with standing agreeing is what makes an estimate binding. Nothing about the estimate itself changes.
In this article
Share

What does runtime context actually mean?

Runtime context is context a system works out while it is running rather than looking it up, computed from what is happening right now: what is trending, what is anomalous, what usually follows what. It is one half of what an AI system is allowed to treat as true about your business, which is not the Kubernetes sense of switching between cluster configurations, not the CRM sense of the record a system keeps about a customer, and not the ITSM sense of the history attached to a ticket.

Most of the context your systems run on today is this kind. It is computed on demand, used, discarded, and computed again from scratch on the next request. It is not written down anywhere as a decision, because it was never a decision.

That is not a weakness, and this is the part that usually gets lost. A trend from six weeks ago is not a trend. Fraud detection running off a fixed list stops detecting fraud. For a large class of problems, being recomputed constantly is the entire point, and something settled and signed would be strictly worse.

Runtime context is an estimate. Estimates can be excellent.

What does runtime context already do in your business?

Four things you almost certainly have running. The pattern is identical in all four: the inference is computed live, and the thing it computes against is not.

Where it runsWhat is inferred at runtimeWhat is never inferred
Recommendation and trendingWhat is trending right now. Nobody maintains a list; it is derived continuously from current behaviour and it is worthless the moment it is stale.Which items are the same item, and what a user is actually entitled to see.
Data lakesDynamic risk weightings computed across the lake and re-scored as new data lands, rather than set once in a policy document.The thresholds the policy actually mandates, and which legal entity a position belongs to.
Operational automationsApproval and exception routing. A transaction is scored live against spend pattern, quarterly budget burn and how unusual the amount is for that team, then passed, held or escalated.Who is authorized to approve what, for which cost centre, to what limit.
Reporting and complianceTransaction monitoring. Behaviour is scored against a baseline the system learns and keeps re-learning, because last year's normal is not this year's normal.Who the customer legally is, which accounts share a beneficial owner, and what the filing obligation is once an alert is confirmed.

Read the right-hand column as a set and something becomes obvious. Every one of those is a decision somebody made and someone is accountable for. None of them is a modelling problem. And every inference in the middle column is only as good as the column on its right, because an approval score computed against the wrong authority limit is a well-calculated wrong answer.

Why is this showing up everywhere now?

Because the work moved and the audience did not follow it.

Statistical refinements like these used to live with data science teams. Those teams understood the confidence interval attached to an estimate, knew which populations the model was thin on, and treated an output as a number with error bars rather than an answer. The inference was produced by people who could evaluate it.

Now the same class of inference is generated continuously, by systems, for consumers who have no way to assess it: an operations manager looking at a routing decision, an analyst reading a variance, an agent acting on a relevance score without pausing.

The inference did not get worse. The reader of it changed. That is the whole shift, and it explains why techniques that were unremarkable inside a data science team are suddenly producing governance problems everywhere else.

How is the market solving this?

A category is forming around exactly this problem, and it is worth naming plainly rather than pretending it does not exist.

ApproachWhat it infers at runtime
GleanBuilds a graph over work content and infers relevance and relationships across tools: org structure, opportunities, tickets, documents.
Databricks Genie OntologyPositions itself as a live context layer and a self-improving knowledge graph that continuously learns the business from data, dashboards, queries and connected apps, claiming better accuracy at lower token cost.
Microsoft 365 Copilot on Microsoft GraphReads work signals across mail, meetings, documents and organizational relationships, and assembles personal and organizational context at request time. The widest deployment of runtime context in the enterprise, and most readers already have it.

These are serious systems and the inference quality is high. But look at how all three describe themselves. Every one of them talks about context that is learning continuously. Not one of them describes a mechanism by which something stops being learned and starts being settled.

They are very good at producing estimates and none of them has a way to retire one. That is not a criticism of any of the products. It is the shape of the category.

What is missing from all of it?

Ground truth, and the gap is narrower than it sounds.

You do not want an inference working out your sales territories. You do not want one estimating customer payment terms. Those facts are definitive, and they are not hard because they are complex. They are hard because somebody has to decide, and then be accountable for the decision.

No amount of continuous learning produces that. A system can observe that a particular rep has closed deals in three states and infer a territory with high confidence, and be describing an accident of history rather than a policy. The inference is a good reading of the evidence. It is still not the answer, because the answer is a decision that has not been made.

In our work with enterprise customers we see roughly 80 to 85% of the accuracy gain, and about 92% of the token-spend reduction, come from getting ground truth right rather than from anything done to the model or the retrieval layer. Both of those are observations from the field rather than benchmarks, and they point the same direction: most of what people try to fix at the model layer was never a model problem.

That is also the argument for runtime inference, not against it. Once the definitive facts are settled and reachable, inference stops burning itself on re-deriving what the company already knows, and gets spent on what it is actually good at: emergent trends, anomalies, and questions nobody could previously answer at all.

How does an inference become ground truth?

Two people with standing agree, and the agreement is recorded where systems can reach it.

That is the whole mechanism, and the important word is the one that is missing. There is no value judgement on the inference. It is not being corrected. It is being ratified.

This matters more than it sounds. A system that treats inference as a defect to be fixed is solving the wrong problem, and it will spend its life arguing with estimates that were right. The inference may have been excellent. It may have been exactly what two people would have written down independently. What changes on ratification is not the content, it is the standing: somebody now owns it, somebody approved the change, and somebody can be asked why.

Which gives a clean division of labour. Inference proposes. People with standing dispose. The set of things that needs to make that trip is much smaller than the set of things your systems infer, and knowing which is which is the actual work.

What does managing runtime context require?

Nine capabilities, and the first four are about connection while the rest are about trust.

CapabilityWhat breaks without it
Integrates with a ground truth systemInference runs against whatever it can reach, which is usually raw data with no notion of which version is current.
Open standardsAn inference cannot be elevated to authoritative without a rewrite, so in practice it never is.
High-performance generationRuntime context arrives after the decision it was supposed to inform.
Can be interrogated for statistical significanceA weak inference and a strong one look identical to everyone downstream.
Provenance on every inferenceYou cannot promote what you cannot trace. What produced it, from what, and when are the minimum.
ExpiryRuntime context is perishable, and a stale inference presents exactly like a fresh one.
Confidence that degrades visiblyA weakening inference goes on looking authoritative right up until it causes a decision.
A promotion path with a named person in itRatification stays a good intention. Nothing ever actually gets settled.
Caller-scoped isolationContext computed for one caller leaks into another's answer, which is a permissions failure wearing a relevance costume.

Most stacks have the first three and none of the last five. That is a reasonable place to have ended up, because until recently the consumer of an inference was a person who supplied the missing scepticism themselves. An agent supplies none of it. It receives the estimate, treats it as settled, and acts.

So the practical move is not to trust runtime context less. It is to work out which of your facts should never have been estimated, settle those, and let inference do the thing it is genuinely excellent at. Context management is the discipline for drawing that line deliberately, a context graph is what the settled side looks like when rules and processes travel with the entities, and why AI gives wrong answers about your own company covers what happens when the line is never drawn at all.

About the Author

Give your AI the context it's been missing

See how the TQ Data Foundation turns your enterprise knowledge into trusted, Al-ready context.