NDE alignment
- Linked Open Limburg (LOL) is positioned as a service platform, in line with the NDE report Van data naar dienst.
- LOL is built on the NDE Stack, the ecosystem of NDE-compatible components.
- The search pipeline selects datasets dynamically from the Dataset Register.
- Datasets are only accepted into LOL if they are published according to the NDE Schema.org Application Profile. LOL absorbs data deviations rather than wait for publishers to fix them.
- LOL software components and applications are open source and adhere to the NDE Software Requirements.
- For hosting the software applications, LOL follows NDE’s infrastructural approach.
Deviations from the profile
Relevant datasets do not satisfy the NDE Schema.org application profile in every field, and LOL is deliberately lenient: waiting for a publisher to fix their data would keep relevant datasets out of LOL in the meantime. Each entry below pairs a deviation with the compensation it prompted:
- the deviation: how a value provided by the publisher conflicts with the profile;
- the compensation: how the index reads that value rather than skipping it.
What the index serves is the value the profile meant, even where the input did not state it that way. But a compensation is an interpretation the publisher never stated, which is why each one is written down here.
dateCreated – an interval is collapsed to its start
The profile requires a single date at the highest precision known, and forbids an interval “even when the exact creation date is uncertain”. The reason is typing, not taste:
Unlike
schema:temporalCoverage,schema:dateCreatedranges overschema:Dateandschema:DateTimeonly, neither of which permits an ISO 8601 interval; intervals belong toschema:temporalCoverage, which describes the content’s subject period rather than its creation.
Publishers ship intervals anyway – 466 values across 31 datasets in the register, and in Limburgs Museum the majority: 252 of 387 dated works.
We index the start bound, so a work made somewhere in 1935–1955 is dated, sorted and filtered as 1935. That reads as more certain than it is, and for 244 works the indexed value is our interpretation rather than the publisher’s assertion – the costliest compensation here.
The profile’s text and its worked example disagree, and we follow the example, which resolves 1980/1985 to 1980 – the start bound, not the “lowest common precision” the text prescribes. That operation cannot be written for almost any work needing it: the legal forms are YYYY, YYYY-MM and YYYY-MM-DD, so two bounds share an expressible precision only inside one year (8 of the 252), and neither a decade nor a century can be said at all.
NDE’s own Linked Art mapping makes the opposite choice at the same wall, dropping a production span that “straddles more than one year”; on this collection that discards 244 of 252 works. A consumer comparing the two indexes sees different dates for the same object, and neither is wrong.
Deletable when the profile can express a production period. Intervals are then indexed as the period they state, and the start-bound reading goes. Raised at issue 90; the range argument rules out permitting intervals or EDTF on dateCreated itself, so the answer is a modelling change rather than a syntax one – see schema-profile#177 for the rationale and #192 for EDTF on temporalCoverage, raised from this project.
creator and contributor – a flat value is wrapped into the Role node that publishers in the register use
The profile forbids the wrapper outright. § 4.2.7: multiple creators “MUST each be expressed as separate, flat creator values”, and creator “MUST NOT wrap its values in a Role node, since this would force consumers to handle two value shapes”.
Publishers wrap anyway, and it is the majority shape by a wide margin. Of the register’s SPARQL-served datasets, 29 wrap creators in a Role (1 559 473 works) against 12 that state them flat (18 865) – 98.8% of creator-bearing works take the forbidden form. It is two platforms rather than one publisher: 28 LDMax/Axiell datasets and Beeld & Geluid.
29 datasets publish:
"creator": {
"@type": "Role",
"name": "uitgever",
"creator": {
"@id": "https://n2t.net/ark:/61567/…",
"@type": "Person"
}
}The profile prescribes:
"creator": {
"@id": "https://n2t.net/ark:/61567/…",
"@type": "Person"
}Unrepaired, creator would hold a blank node and the name field a job title: for 159 of Limburgs Museum’s 291 values the Role is a blank node, so there is no facet key, and the Role’s own name is the role – uitgever on 90 works, auteur and fotograaf on 46 each.
We index the Role rather than reading through it, as the work→agent edge it is: one entry per role-and-agent pair, carrying the role the publisher stated and resolving the agent out of the Person collection. A pair rather than a creator, because the entry is what a filter tests: a Role naming two agents, or naming its role once per language, becomes that many entries so that each one can answer this agent in this role. The shape is Schema.org’s own: schema:Role stands in the property’s place and repeats that property to the target – work → creator → Role → creator → agent – so the role rides the edge while the agent stays a node in its own right. Not a container LOL invented to park a role in, but the construction Schema.org defines for the purpose, and the one the 29 wrapping datasets write.
creator { role creator { id name { value } } }
where: { creator: { where: { creator: { in: ["…"] }, role: { in: ["etser"] } } } }Conditions inside one edge’s where are welded to the same entry, so that asks this person in this role rather than this person somewhere and this role somewhere.
The compensation is that the 1.2% is wrapped up to meet it, not that the 98.8% is unwrapped down. One declaration reads every dataset: a publisher who states the creator flat has an edge minted for it at index time, and a consumer handles one value shape rather than two. The claim we make that the publisher did not is an edge with no role on it. (Carrying edge data through a resolved lookup needed ldelements/lde#771.)
The wrapper is found by its second creator hop, not by the node’s type. Publishers vary in everything else: Limburgs Museum names the role on schema:name, Beeld & Geluid on schema:roleName (a Wikidata IRI rather than a string – so its roles reach the index not at all, see below), and Beeld & Geluid’s wrapped contributors are typed schema:PerformanceRole, a Role subclass no store here reasons over. The hop catches every variant and cannot fire on a conforming publisher, because a Person is not the creator of anything.
contributor takes the same edge. The profile leaves contributor unconstrained (see Additions), so the same structure is forbidden on one property and permitted on the other. We apply it to both regardless: a consumer learns one edge, and which property the profile happens to constrain is not a distinction the index should make anyone hold in their head. One dataset wraps it – Beeld & Geluid, across all 140 982 of its contributor works – and its wrapper carries no roleName and no name at all, so those edges simply carry no role. The other 36 datasets that ship a contributor state it flat, and none mixes the two.
The agent’s name is served inside the entry, and creatorName / contributorName are gone. A creator the graph named inline – no URI, which is seven creators in eight at Drapo – is an entry like any other minus its id, so it is displayed and reaches free text; only the facet passes it by, being keyed on identity alone so that two people who share a name are never merged into one bucket. That closes issue 83.
A role stated as a URI is not indexed at all. The field reads literals, so a schema:roleName that names a Wikidata concept rather than spelling one out is passed over silently. That is the whole of Beeld & Geluid: queried against its endpoint, no node carries a literal and an IRI together, and the roleName values are IRIs throughout – wikidata.org/entity/Q…, data.muziekschatten.nl/som/…. Its creator roles are therefore absent from the index rather than degraded in it, and no filter can reach them. Reading them means modelling the role as a term, with the URI as the concept's identity, which is also what would let a role be matched across datasets and faceted on the concept instead of on spellings.
Roles stay literals, and there is no role facet. Sources state the role as a string, through two predicates, so “etser”, “Etser” and “prentmaker” are three values; minting IRIs from those labels would be inventing identity rather than reading it, which is what we decline to do for creators. A facet would in any case be unserviceable: Typesense scopes filter_by to one array element but not facet_by, so a facet over roles counts every role of every matching work, whichever entry satisfied the filter.
Deletable when publishers state creators flat. Unlike the compensations below, this one has a clear end: the profile already forbids the shape, so it is a matter of the two platforms catching up rather than of the profile changing. What ends then is only the minting; the edge itself stays, because a role is still what it carries.
sdDatePublished – filled from the record’s modification date
The profile requires it on every record. Across the register’s CreativeWork partitions, only 7 datasets declare it. Most record the same fact elsewhere: 38 under dcterms:modified and 11 under schema:dateModified, both record-level.
We fall back to both, preferring the profile’s own property wherever a publisher supplies it, then dcterms:modified, then schema:dateModified. Reading only the first fallback would leave those 11 datasets with nothing – and since the field is required, that does not fail loudly: every document is dropped and the run still reports success (ldelements/lde#738). required stays – it is GraphQLNonNull in the API, so relaxing it to admit one publisher would weaken the guarantee for every consumer.
For Limburgs Museum the fallback is total rather than partial: all 412 IRI-identified works carry dcterms:modified and none carries sdDatePublished.
A further 81 datasets record none of the three, and no fallback reaches them. Onboarding one is a decision about that dataset, not a schema change.
Deletable when publishers emit the profile property. The two fallbacks then go, and nothing else changes. Low risk meanwhile, since both properties describe the record, so the substituted value answers the question the field asks.
sameAs on a place – the aligned term becomes the key, and the description
The profile models an external term as an alignment, not as an identity. § 3.4: the DefinedTerm is “a local node, scoped to the publishing dataset”, the sameAs URI is “the canonical identifier”, and consumers MUST compare on it rather than on the nodes’ own identities, “which are local to each publishing dataset and therefore not shared across datasets”. § 4.2.14 adds that a referenced term used as a location needs Place in its @type, because schema:contentLocation takes Place, not DefinedTerm.
Publishers deviate two ways, counted over every node reached by locationCreated, contentLocation, location, birthPlace or deathPlace – 277 in Drapo, 19 in Limburgs Museum:
- Drapo conforms: 263 of its 277 are
DefinedTerm+Placecarrying asameAsinto GeoNames, and the remaining 14 are custom resources with no term available, which § 4.2.14 allows. - Limburgs Museum uses the canonical IRI as the local node’s own identity – 10 nodes whose
@idishttps://sws.geonames.org/…, with no local node and nosameAs. - One is a GTAA node in the same shape – its own IRI is the thesaurus term – so it is keyed and described the same way, from GTAA rather than GeoNames.
- Eleven of Limburgs Museum’s 19 are also typed both
schema:Placeandschema:Person, one IRI carrying a place and the people associated with it:…/2759794/is “Amsterdam” and a printer,…/2751306/is “Maas” and“onbekend”. Reported in #94; the person half is refused a document – a deviation of its own below.
Limburgs Museum publishes:
"contentLocation": {
"@id": "https://sws.geonames.org/2759794/",
"@type": "Place",
"name": "Amsterdam"
}The profile prescribes:
"contentLocation": {
"@type": ["DefinedTerm", "Place"],
"name": "Amsterdam",
"sameAs": "https://sws.geonames.org/2759794/"
}We key a place on the term it aligns to, and take its description from that term. Where a node’s alignment – or its own IRI – lies in a source LOL covers, the document is keyed on that IRI and its language-tagged names, coordinates and country come from the authority, read through the Network of Terms; the publisher’s statements about it are discarded rather than merged. Three reasons, in order of how much they cost:
- Publisher data cannot be trusted for a term the publisher merely referenced. Limburgs Museum’s names for these nodes include a printer’s name and
“onbekend”. Keeping them would put one of those on a place shared with another collection. - One document per place is the point. Six GeoNames IRIs are referenced by both datasets today; keyed locally they are six places twice over, in two facet buckets with two labels.
- Nothing else can supply the description. Neither publisher ships
schema:geoat all, and no address component appears on any place in either dump, so the coordinate a consumer asks for exists only at the authority.
Coverage is GeoNames, GTAA and Wikidata, listed in packages/search-schema/src/covered-sources.ts and decided there rather than against the Network of Terms’ own sources registry. Deciding it statically is an identity constraint, not caution: a registry read would make the document key depend on a request, so one outage would re-key every place, rewrite every reference to them, and let two runs over identical dumps disagree. The Network of Terms is the route to an authority, not the arbiter of which ones LOL keys on. Widening the list is a PIPELINE_VERSION rotation, since no publisher’s dump changes when it does.
GeoNames is preferred where a node aligns to several, because it is the only one carrying coordinates; GTAA and Wikidata carry a name and no coordinates, which is all keying needs.
The name is the authority’s name for the place, not its picker label. GeoNames answers skos:prefLabel with a country code concatenated in – “Venlo (NL)”, “Duitsland (DE)” – to tell homonyms apart in a term picker, and states the plain name on the place it describes. We take the plain one, and the suffix reaches a consumer as addressCountry: the same disambiguation, as a field to render or filter on rather than punctuation to parse back out of a label. A term that denotes no place has no such description, and its prefLabel carries no suffix, so that is what we fall back to.
No spelling is rewritten. Both collections already write every term IRI in its authority's canonical form – all 288 distinct GeoNames targets in Drapo and all 12 in Limburgs Museum as https://sws.geonames.org/N/, and Limburgs Museum's one GTAA term as http://data.beeldengeluid.nl/gtaa/N (measured 2026-08-25). There is nothing to reconcile, and a normalisation that has never run against real data is one nobody can be sure is right. Repetition costs nothing either: Drapo states those 288 targets over 554 triples, one per named graph, and the key rule dedupes. Should a publisher ever spell a term two ways, that is the point to add a rule – for the spelling actually seen.
A place with no alignment keeps the publisher’s identity and the publisher’s description, unchanged. LOL does no entity matching: “Kessel” appears in both datasets as a local place and stays two documents, because neither publisher said they are the same.
Deletable when publishers align their places to a covered vocabulary and describe them as § 3.4 asks – a local node carrying sameAs, rather than the canonical IRI standing in for one. The 10 Limburgs Museum nodes would then key the same way for a conforming reason, and the person conflation would go with them.
DefinedTerm – a term identified by its vocabulary URI is described from it
The same shape as for places: § 3.4 makes a DefinedTerm a local node carrying the canonical URI as sameAs. Both LDMax collections publish the shorter thing instead: the vocabulary URI is the node, with a name and nothing else.
Limburgs Museum publishes:
{
"@id": "http://vocab.getty.edu/aat/300054698",
"@type": "DefinedTerm",
"name": "schrijven"
}The profile prescribes:
{
"@type": "DefinedTerm",
"name": "schrijven",
"sameAs": "http://vocab.getty.edu/aat/300054698"
}Measured 2026-08-26: 185 such nodes in Limburgs Museum and 20 in Discovery Museum, across Getty AAT (175 distinct IRIs), the Cultuurhistorische Thesaurus (26), GTAA subjects, the Archeologisch Basisregister and Wikidata. Drapo publishes the conforming shape – 1022 terms under https://id.drapo.nl/, 291 of them carrying sameAs.
We describe the first kind from its authority, and leave the second alone. A term whose own IRI lies in a covered source takes that source’s language-tagged names, its alternate names, and the authority and fetchedAt stamps. A term with a local IRI keeps everything the publisher said – uncovered, unchanged, and still a separate document from an LDMax term for the same concept until #89 keys it on its sameAs.
The gain is not cosmetic. Publishers state one name, untagged: "schrijven" with no language, which indexes under und and is therefore missing from every locale a visitor searches in. AAT states "kapitelen"@nl and "capitals (column components)"@en, and a median of three alternate labels per term – usually the singular against AAT’s plural (kapiteel, vloer, altaar), which is what a visitor actually types.
Not every reference resolves. Roughly a third of the AAT URIs in the register answer NotFoundError – not because Getty lacks the term, but because a lookup by URI does not route to the AAT sub-vocabularies it lives in (netwerk-digitaal-erfgoed/network-of-terms#1943, filed from this register). Those terms keep the publisher’s untagged label, and say so by carrying no authority.
Deletable when publishers follow § 3.4 and every vocabulary the register uses is reachable by URI lookup. What ends is reading a node’s own IRI as its alignment; describing a term from its authority continues, reached through sameAs instead.
Person – an artist identified by an authority record is described from it
The same shape again, and § 4.2.7 gives our exact case as its example: a creator referenced by an authority record is a local node co-typed DefinedTerm and Person, with the authority’s URI as sameAs. Limburgs Museum publishes neither half of that:
Limburgs Museum publishes:
"contributor": {
"@id": "https://data.rkd.nl/artists/104628",
"@type": "Person",
"name": "Société Céramique"
}The profile prescribes:
"contributor": {
"@type": ["DefinedTerm", "Person"],
"name": "Société Céramique",
"sameAs": "https://data.rkd.nl/artists/104628"
}The RKDartists record is the node, so the canonical URI stands in for a local node rather than being pointed at from one. Drapo publishes the conforming shape, co-typing 140 of its 158 persons.
We key a person on the authority record they align to – or on their own IRI where that record stands in for one – so two collections crediting the same artist credit one person, whichever shape the publisher used. Covered for persons, being every namespace whose Network of Terms genre set carries Personen: RKDartists, the KB thesauri, Wikidata, GTAA and Canon van Limburg.
A person authority, not merely a covered one. Publishers do point person properties at places: Limburgs Museum states schema:contributor → sws.geonames.org eleven times, on nodes typed both schema:Place and schema:Person – …/2803138/ is “Antwerpen” and a printer (#94, finding 1). GeoNames names places and nothing else, so no alignment to it can become a person's identity.
That covers the alignments. It cannot cover a node whose own IRI is the place, because there is nothing to choose between and the node keeps the identity it has – so a second rule decides what may describe a document: an Authority describes a collection only where it names that kind of thing. GeoNames’ name and authority stamp therefore never land on a person. Where the node’s own IRI is one GeoNames reserves for a place, it is not a person at all, and yields no person document – the deviation below.
Which sources name which kinds is the Network of Terms' own answer, from the genres its registry states per source – not a judgement of ours. Recorded in packages/search-schema/src/covered-sources.ts rather than fetched, because a key must not depend on a request.
What an authority adds. RKDartists states the name a record is filed under and the variants it is also known by: “Breetvelt, Henri” beside “Breetvelt, Henri Leonardus August”, and for a firm the names it traded under – “Société Céramique, N.V.” beside “Céramique Maastricht, N.V. Keramische Industrie genaamd”. Spellings a visitor may search for and no publisher records.
Birth and death dates come from the authority too. Since 9 September 2026 the Network of Terms answers a structured person node beside place (network-of-terms#1968), with RKDartists the first source to fill it; before that the dates were prose inside skos:scopeNote, and parsing them would have meant reverse-engineering a display format (#139). Max Verboeket, whom Discovery Museum publishes under his RKD record with a name and nothing else, now answers:
RKDartists states, through the Network of Terms:
"person": {
"birthDate": "1922-07-08",
"deathDate": "2015-01-23",
"birthPlace": [{ "uri": "https://data.rkd.nl/thesaurus/4641", "name": [{ "language": "nl", "value": "Hoensbroek" }] }],
"deathPlace": [{ "uri": "https://data.rkd.nl/thesaurus/46", "name": [{ "language": "nl", "value": "Maastricht (stad)" }] }],
"hasOccupation": [{ "occupation": null, "roleName": [{ "language": "nl", "value": "glasontwerper" }] }, …]
}LOL serves:
{
"id": "https://data.rkd.nl/artists/80083",
"name": [{ "value": "Verboeket, Max" }],
"birthDate": "1922-07-08T00:00:00.000Z",
"deathDate": "2015-01-23T00:00:00.000Z",
"birthPlace": [],
"hasOccupation": [
{ "id": null, "name": [{ "language": "nl", "value": "glasontwerper" }] },
{ "id": null, "name": [{ "language": "en", "value": "glass designer" }] },
…
],
"authority": "https://data.rkd.nl/rkdartists"
}The dates are EDTF, and the field holds one instant. RKD states what it knows at the precision it knows it: Julius Goltzius was born in 1555 and died in 1595-09, and Rembrandt in 1606-07-15/1607 – 15 July 1606 by the earliest source, 1607 by two others. birthDate is a date, so an interval is reduced to its start bound, as dateCreated already is, and a Level 1 qualifier (1620~, “circa”) is stripped – that part is EDTF’s alone, since the profile admits no qualifier on a work’s date. That is a lower bound presented as a date, and it reads as more certain than it is. What cannot be reduced – an unknown start, a decade written 162X – is left absent rather than guessed. The profile-side counterpart, EDTF on a person’s dates, is schema-profile#206.
An occupation stated by name alone is an entry without an id. The profile (§ 4.4.6) has hasOccupation reference a term, “or, if no term is available, a custom Occupation resource” carrying a name – a blank node, in practice. RKD states every occupation that way: Verboeket has fourteen, glasontwerper and emailleur among them, none with an IRI, and each once per language, since nothing in RKD pairs the Dutch role with its English one. hasOccupation is declared as a local lookup, the same construction that serves a creator the graph named inline: the entry carries the name and no id, is displayed and searchable, and only the facet passes it by, being keyed on identity. We do not mint an IRI from the label, for the reason we do not for creators. The facet on hasOccupation is gone with this – a local lookup stores the referent, not an id an engine can facet on – and it counted nothing, since no dataset here and no authority yet identifies an occupation by IRI.
Places are empty on a keyed person, for now. The authority replaces the publisher’s description whole, as it does for a place, so a birthplace a publisher stated on an RKD-keyed person is discarded with the rest – and the authority fills nothing in its stead yet. RKD names a birthplace by its own thesaurus – https://data.rkd.nl/thesaurus/4641 is Hoensbroek – and that IRI keys no Place document here, since the RKD thesaurus is not a source LOL covers; a reference to a document that does not exist resolves no label. RKD’s alignments of its thesaurus to GeoNames live in another endpoint and are a follow-up upstream (network-of-terms#1967); once they reach the Network of Terms the field is filled from there. A person nobody aligned keeps the publisher’s birthplace, as it keeps everything else.
How much this reaches, measured 10 September 2026: four distinct RKD records across the three datasets. Limburgs Museum holds two as schema:Person nodes – Société Céramique, the contributor of 9 works, and Julius Goltzius, referenced by no work at all. Discovery Museum holds three – Société Céramique again, Kristalunie and Max Verboeket – none of them referenced by any statement in the dataset. Drapo references RKD nowhere; its aligned places all point at GeoNames. Two of the four are firms, for which RKD answers person: null, so the dates and occupations land on Goltzius and Verboeket. Thin, and a fact to report rather than a reason to hold back: the field is declared to the profile, and the next dataset that aligns its artists to RKD gets the same treatment without a change here.
Deletable when publishers follow § 3.4 – a local node, co-typed DefinedTerm, carrying the authority URI as sameAs – and stop giving one node two identities. As for DefinedTerm, only the reading of the node’s own IRI ends; the person is still keyed on and described from the authority, through sameAs. Both are reported in #94.
A node typed both Place and Person – the person half yields no document
The profile (§ 3.4) has a referenced place be a local node, co-typed DefinedTerm and Place, pointing at GeoNames through sameAs – and the GeoNames URI “keeps its own typing in its source vocabulary”. A creator or contributor gets a local node of its own, co-typed Person. One IRI, one thing.
Limburgs Museum ships one node for both. Eleven GeoNames IRIs and one GTAA term are typed schema:Place and schema:Person, each carrying the place’s name beside the names of the people associated with it: …/2803138/ is “Antwerpen” and “Vorsterman, Willem”, …/2748000/ is “Roermond (plaats)” and “Meester van Elsloo”. 41 works reach one through locationCreated, 20 through contributor. The cause is the missing local node: with the authority URI standing in for it, the place and the printer have nowhere to be apart. Reported as findings 1 and 7 of #94.
We refuse the person half a document. The Person stage selects the node by its type like any other, and since it carries no sameAs it keeps its own IRI as its key – so until #164 persons answered for a city’s IRI, with a name that was the city’s or the printer’s depending on which value was read last. Now an IRI identifies a document of a kind only where its Authority names that kind at all. GeoNames names places and nothing else, so the indexer withholds the node’s quads from the Person projection and no document is made; the Place projection is untouched. Which sources name which kinds is the Network of Terms’ own classification, recorded in packages/search-schema/src/covered-sources.ts rather than fetched, so a document cannot exist in one run and not the next. The rule is symmetric – an RKDartists record typed as a place would be dropped from places – and refuses nothing it does not cover: a publisher’s own node is whatever the publisher typed it.
What it costs is the printer. “Vorsterman, Willem” has no IRI of their own, so there is no document to move the name to, and minting one would be an identity the publisher never stated. A work’s contributor pointing at the city still serves an entry: the lookup is local, read in the work’s own stage, so the entry keeps the node’s id and whichever name the publisher put on it – only the persons document behind it is gone, and alternateName and authority with it. The GTAA node – “Wales” and “onbekend” – keeps its person half, because GTAA does name persons and nothing in the IRI says which sub-source it lives in; that residue is #146’s.
Deletable when the publisher mints a local node per reference, as finding 7 asks: the place and the printer then have their own IRIs, the Person stage never selects a GeoNames node, and the rule fires on nothing.
iiifManifest – the manifest is read from the property an older profile version put it under
The profile (§ 4.2.6) requires the IIIF Presentation manifest to be an associatedMedia entry of its own, with @id the manifest URI and encodingFormat application/ld+json;profile='http://iiif.io/api/presentation/3/context.json'.
Publishers ship the older shape – a version lag, not a modelling error, which makes this the cheapest deviation here to retire. Before v0.6.0 the profile put both IIIF APIs on the MediaObject under schema:isBasedOn, requiring an encodingFormat on each target and sh:pattern-constraining it to …/api/(image|presentation)/3/context.json. v0.6.0 (2026-03-04) moved the Presentation manifest to the work and narrowed isBasedOn to the Image API alone; v0.11.0 (2026-05-14) dropped it there too. Limburgs Museum and Discovery Museum still emit the pre-v0.6.0 shape, discriminator included.
Limburgs Museum and Discovery Museum publish:
"associatedMedia": {
"@type": "ImageObject",
"contentUrl": "https://iiif.example/3/w1/full/max/0/default.jpg",
"isBasedOn": [
{
"@id": "https://iiif.example/3/w1",
"encodingFormat": "application/ld+json;profile='http://iiif.io/api/image/3/context.json'"
},
{
"@id": "https://example/w1/manifest.json",
"encodingFormat": "application/ld+json;profile='http://iiif.io/api/presentation/3/context.json'"
}
]
}The profile prescribes since v0.11.0:
"associatedMedia": [
{
"@type": "ImageObject",
"contentUrl": "https://iiif.example/3/w1/full/max/0/default.jpg"
},
{
"@id": "https://example/w1/manifest.json",
"encodingFormat": "application/ld+json;profile='http://iiif.io/api/presentation/3/context.json'"
}
]We read the manifest from there, by the discriminator that profile version defined – the same encodingFormat test the conformant route applies, one hop further out. Nothing is inferred: the publisher states which API each target is, and the two collections do so on every reproduction (765 and 37, each with exactly one Presentation target). Without it the field was null for all 463 Limburgs Museum and 17 Discovery Museum works with media, while Drapo – which ships the conformant entry – answered normally and carries no isBasedOn at all.
A conformant entry always wins, so this only ever fills a field that would otherwise be absent. It is the mildest compensation listed here: the value, and the fact that it is a manifest, are both the publisher’s own assertions. What is ours is only the decision to keep reading a retired property.
Deletable when publishers emit the manifest as its own associatedMedia entry with the IIIF encodingFormat. The isBasedOn read then goes, and the conformant route already in place serves every work. Reported as finding 5 on issue 94; raised from issue 134.
size – the number is read out of a structured node
The profile currently types size as a display string. Publishers ship structured values too – 3246 of 3397 size violations across 32 datasets are schema:QuantitativeValue nodes.
32 datasets publish:
"size": {
"@type": "QuantitativeValue",
"name": "diepte",
"value": 12,
"unitText": "cm"
}The profile prescribes:
"size": "diepte 12 cm"We read size/value as an alternative step, because a bare size path frames the node, finds no literal, and empties the field for every work in those datasets without a word. This is repair rather than circumvention: the live proposal keeps the display statement unchanged and additionally allows a QuantitativeValue per measured dimension, so the shape we read is the one the profile is moving towards.
The unit stays out of reach, for a narrower reason than the profile having nowhere to put it: the proposal asks for a unitCode URI and a valueReference naming the dimension, while published nodes carry unitText and an untagged name such as "diepte". Separate width/height/depth properties were proposed and rejected – Schema.org puts them on VisualArtwork, and adoption was below the profile’s bar.
Deletable when publishers emit units the way the profile asks. The reading itself never goes: once the profile admits the node, size/value is the prescribed path rather than a workaround, so what ends is only its standing as a compensation. The proposal landing does not do that by itself, since it sets a target rather than blessing current data.
Additions to the profile
The profile constrains source data: what a publisher puts in their dataset. The index is a property of LOL’s data layer, one level up, and the profile says nothing about what it may serve. So a field that no profile property backs is not a deviation; it is ours to define, and it adds rather than papers over. There are four:
type: the work’s Schema.org class,CreativeWorkplus whatever subclass applies, as the facet that separates paintings from photographs across datasets.hasMedia: whether the work carries a displayable image, as a filter and facet.iiifManifest: the IIIF Presentation manifest, lifted out ofassociatedMediaonto the work.dataset: the dataset a document was drawn from. A fact about the harvest rather than a claim the data makes.
An addition can still acquire a compensation. iiifManifest has one: the field is ours to define, but which property it reads when a publisher follows an older version of the profile is listed above.
contributor is an addition of a second kind – a property the profile names but excludes. § 4.2.7 points publishers who want to qualify a role at schema:contributor, while stating it is “outside this profile and does not affect profile validity”. So indexing it takes nothing away from a publisher who ships none, and 37 datasets in the register do ship it, over 271 814 works. It is deliberately not read as a second creator: at Limburgs Museum the two agent populations barely overlap (58 against 181, 5 in both), and where a work carries both, the contributor is the secondary party – a printer or a distributor. Reading it as creatorship would assert what the publisher declined to. The Role compensation covers it too, on the same terms.