TQ
Content

Your archive is only licensable if you can describe it precisely

The publishers signing AI licensing deals are the ones who can say exactly what is in the corpus, who holds the rights, what is excluded, and how each item was produced. That is a metadata and rights governance capability. TopQuadrant is where that description lives, and it improves your own search and recommendation on the way.

Data products · Corpus definition
Licensing corpus Environment archive, 2015–2026
Rule-defined
Assembled from Concept beneath Environment Published 2015 onward Rights permit model training
84,102 items included
6,318 excluded, with a reason
Third-party stills, no training right3,904
Contributor opted out1,733
AI-generated · excluded by policy681

Every exclusion is attributable to a governed field, so the corpus can be described to a counterparty rather than just delivered.

Two revenue lines now depend on the same metadata

Discovery and licensing pull on the same thing. Recommendation quality is bounded by tag quality and entity resolution. Licensing revenue is bounded by whether you can define a clean, cleared corpus and account for its use. Publishers with organised archives, clear rights, and structured metadata are the ones getting into these programmes at all.

$250M+

Reported value of one major news publisher’s five-year AI licensing agreement, an indication of what a describable archive is worth.

300+

Prebuilt vocabularies to ground search, recommendation, and corpus definition.

1

One governed description of your content, serving discovery, syndication, and licensing.

Taxonomy-grounded discovery

Search that knows what the content is about

Retrieval grounded in governed concepts and resolved entities returns material about the right subject rather than material containing the right words. Hierarchy, synonyms, and superseded labels all work in your favour instead of against you.

  • Concept-based retrieval. Query a concept and get everything beneath it, under any of its labels.
  • Entity-linked content. Find material via the person, place, organisation, or product it discusses.
  • Relationship traversal. Reach content through hierarchies, franchises, series, and events.
EV subsidiesElectric vehicles812 itemsBEV · PHEVnarrowerIncentivesrelated

Provenance & content credentials

Record how each item was produced

Provenance is now a publishing obligation as much as a trust measure. Capture IPTC Digital Source Type values and C2PA content credentials as governed metadata at ingest and edit, across wire copy, commissioned work, user-generated material, and AI-assisted edits.

  • Digital Source Type captured. Governed values distinguishing human, AI-generated, and AI-composited material.
  • C2PA-aligned. Content credentials held and preserved alongside your own metadata.
  • Every path covered. Consistent provenance across each ingestion and editing route.
1Capturedphotographer2Editedcolour3AI assistdisclosed4PublishedsignedContent credentialsa verifiable C2PA chain travels with the item

Corpus definition for licensing

Define exactly what a partner is licensing

A licensing corpus has to be defined by rules and provable by evidence: which titles, which date ranges, which rights status, which authors opted out. Governed metadata makes that a query, and it makes the exclusions auditable.

  • Rule-defined corpora. Assemble a corpus from governed criteria rather than a manual export.
  • Provable exclusions. Show what was withheld and the governed basis for withholding it.
  • Per-asset accounting. Support usage-based licensing models with asset-level records.
Corpusdefinition48,210 incleared3,940 outthird-party

Opt-out & crawler signals

Make your rights position machine-readable

Opt-out only works if the underlying per-asset consent and rights data is current and structured. Govern that data, then generate the robots.txt directives, TDM reservation signals, and rights expressions that carry it outward.

  • Per-asset consent. Training and reuse permissions held as governed fields per item.
  • Signals generated, not hand-maintained. Crawler and reservation directives derived from governed data.
  • Position reportable. Answer what share of the archive is opted out, by title and by date range.
RightsArchive reservedTraining: licenceWire: disallowOpt-outs excluded
What customers say
“The licensing conversation moved fast because we could define the corpus precisely and show what was excluded and why.”
Director, Content Licensing
Global news organisation
1
governed description serving discovery and licensing
100%
of corpus exclusions traceable to a governed rule