Using the search API
The search API serves GraphQL at http://localhost:4000/graphql: the read side of the stack that Running the search API brings up. Have that running before you try anything here – every example below queries your own machine.
A first query:
{
creativeWorks(query: "Maastricht", perPage: 10) {
items {
name {
value
}
dataset {
id
name {
value
}
}
associatedMedia {
thumbnailUrl
contentUrl
license
}
}
}
}Each reproduction of a work is one associatedMedia entry, carrying its own URLs, license and attribution, so an object with several images keeps each image’s values together. A work that also publishes a IIIF Presentation manifest exposes it as iiifManifest – one field, however the publisher states it, so a IIIF-aware client dereferences that and never has to read associatedMedia to find the manifest or to know which version of the profile a collection follows. hasMedia filters down to works with a displayable image.
dataset is the collection a result came from – a fact about how it was indexed, and the one carried by every entity type that has no containing-collection property of its own: organizations, terms and occupations. places and persons are the exceptions, for a reason worth knowing: both are keyed on a canonical identifier, so one document is shared across the datasets that reference it and a single dataset value would name whichever of them was indexed last – a filter on any of the others would then miss it. See A place is one place and One artist, however many collections credit them. Query the datasets collection by that id for the rest of the register’s description: publisher, license, description and landingPage. (A reference serves the referenced collection’s own fields – and every collection serves its display field as name, so dataset { name { value } } reads the same as a creator or a material reference. A dataset’s name is its dcterms:title in the register; the profile a value came from never changes the word you ask for.) (isPartOf is a different thing – what the publisher claims the work is part of. Where that is just the dataset it was harvested from it is left out, so it names only the other datasets a work belongs to, and is usually absent.)
The playground at http://localhost:4000/graphql documents every collection, field and filter, and is always up-to-date. What follows is the handful of shapes worth knowing before you go exploring. Each is a link that opens the query already loaded.
Eight collections are served: creativeWorks, persons, organizations, places, terms, occupations, datasets and publishers. Every one takes the same five arguments – query, where, orderBy, page, perPage – and returns items, pagination and facets.
Filtering on a date before 1 CE needs the expanded form. A bound is parsed with JavaScript's Date, which reads a leading minus as a UTC offset unless the year has exactly six digits – so dateCreated: {min: "-1100"} silently means 1099 CE and matches nothing, while min: "-001100-01-01T00:00:00.000Z" works. That is also the shape the API returns, so round-tripping a value you read back is always safe. Upstream fix at ldelements/lde#725; until it lands, pad the year.
perPage runs from 0 to 100; page through page for more. Asking for more than 100 currently fails with a bare “Unexpected error.” rather than saying what is wrong, so check this first if a query stops working when you raise the page size (ldelements/lde#730).
Browse without searching
You get pagination.total and every facet bucket, and items comes back empty. Nothing is fetched or hydrated, so it is the cheap query to run on page load, before the user has typed anything, to populate a filter UI.
{
creativeWorks(perPage: 0) {
# 0 is deliberate: it asks for total count and facets without the documents
pagination {
total
}
facets {
material {
value
label {
value
}
count
}
hasMedia {
value
count
}
}
}
}Each bucket carries the value to filter on and the count to show beside it. label is resolved only where the field points at another collection – material here, whose labels come from terms – so a facet over a plain keyword field such as type or license returns label: null, and the interface falls back to the value itself.
Filter
Keys inside where are combined with AND, and each takes a list that is itself an OR. Ask for banners that are made of silk or velvet, and have an image:
{
creativeWorks(
perPage: 10
where: {
hasMedia: true
material: {
in: [
"https://id.drapo.nl/2d744075-7952-5e8a-ad0e-0cd664102d6c"
"https://id.drapo.nl/a472ec0b-957c-57a9-8178-b70fdab1c902"
]
}
}
) {
pagination {
total
}
items {
name {
value
}
}
}
}Combine fields with or
To OR across different fields, list the alternatives under or. Each entry is a criterion: the same field names as where, without further nesting. This asks for banners that either depict a coat of arms or are made of silk:
{
creativeWorks(
perPage: 10
where: {
or: [
{
about: {
in: ["https://id.drapo.nl/a005b7a1-562a-5d71-83c2-b26dd96dfae6"]
}
}
{
material: {
in: ["https://id.drapo.nl/2d744075-7952-5e8a-ad0e-0cd664102d6c"]
}
}
]
}
) {
pagination {
total
}
}
}253 works carry that subject and 123 that material; the union is 341, because 35 have both. or sits alongside ordinary keys, so adding hasMedia: true next to it narrows the union rather than joining it, to 340.
To ask “what refers to this object, through any field?”, generate that list of criteria instead of writing it out: What refers to an object reads every field that can hold an IRI out of the schema itself.
Group with and
Most combinations need no grouping. Ordinary keys already AND, and each field's in already ORs, so “(silk or velvet) and depicting a coat of arms” is just two keys. Reach for and only when both sides of the AND are themselves cross-field ORs, which nothing else can express.
and takes whole where objects rather than criteria, so each entry may carry its own or:
{
creativeWorks(
perPage: 0
where: {
and: [
{
or: [
{
about: {
in: ["https://id.drapo.nl/a005b7a1-562a-5d71-83c2-b26dd96dfae6"]
}
}
{
material: {
in: ["https://id.drapo.nl/2d744075-7952-5e8a-ad0e-0cd664102d6c"]
}
}
]
}
{
or: [
{
about: {
in: ["https://id.drapo.nl/f5ca916a-fab2-5cdd-a0c7-e7105856fcd5"]
}
}
{
material: {
in: ["https://id.drapo.nl/a472ec0b-957c-57a9-8178-b70fdab1c902"]
}
}
]
}
]
}
) {
pagination {
total
}
}
}The first group matches 341 works and the second 554; their intersection is 221.
Dataset details
A detail view usually has to say what dataset an object came from and on what terms: the license to honour, the institution to credit, and somewhere to send a visitor for the full record. A work carries the dataset it was indexed from, but only as a reference: an id and a label.
{
creativeWorks(query: "vaandel", perPage: 1) {
items {
id
name {
value
}
dataset {
id
}
}
}
}The terms live on the dataset, so take that id and ask the datasets collection:
{
datasets(
where: { id: { in: ["https://id.drapo.nl/dataset/drapo-schemaorg"] } }
) {
items {
name {
value
language
}
license
landingPage
publisher {
id
name {
value
}
}
}
}
}Two round trips, because there is no join yet, but only two for a whole page of results, not one per object: collect the distinct dataset ids from the page and pass them together, since a page of results usually shares a handful of datasets between them.
dataset and isPartOf both resolve against datasets, and they are not the same thing. dataset is a fact about the harvest: the dataset this document was drawn from, which is how the index knows it at all. isPartOf is a claim the publisher makes about the work, and a work may assert several. For a work drawn from the dataset it says it belongs to the two coincide – for any other value they do not, and only dataset tells you where the record came from.
Where an authority’s description comes from
A document keyed on an authority’s URI – a GeoNames place, an RKDartists person, a Getty AAT concept – takes its description from that authority, read through the Network of Terms, NDE’s lookup service for terms by URI. authority names the vocabulary, and fetchedAt says when the description was read.
The indexer looks terms up when it builds the index, not when you query it, and stores what it gets. It fetches a stored description again after 30 days, and at most 1000 terms per run. So:
fetchedAtis usually older than the index, and an authority’s correction can take a month to reach LOL.- A new term may keep the publisher’s description for a run or two, and so may every term after a release that changes what the indexer keeps per term.
- An outage changes nothing you can see: the stored description is served. A term that was never fetched keeps the publisher’s.
A missing authority means the publisher did not align the term, the authority has no description of it, or the indexer has not fetched it yet. The response does not tell the last two apart. How the indexer limits its requests is in its README.
The same concept in two datasets
It depends on how the publisher identified the term, and there are two shapes in the register.
Where a publisher writes the vocabulary URI as the term's own id – both LDMax collections do, publishing http://vocab.getty.edu/aat/300054698 as a schema:DefinedTerm – the term is keyed on that URI. One concept is one document with one id, whichever datasets reference it, and it carries its authority's description: language-tagged names, the alternate names that vocabulary knows it by, and authority and fetchedAt saying where they came from and when.
Where a publisher mints a local id and points at the vocabulary with sameAs – Drapo does, using UUIDs under https://id.drapo.nl/ – the term stays local. Two datasets describing the same concept then give you two terms with two ids and two labels, and what they share is sameAs. That is what to compare on, not the term's own id. (Bringing those onto their canonical URI too is #89.)
For that second shape, matching a concept across datasets is two queries. Find the terms first:
{
terms(where: { sameAs: { in: ["https://sws.geonames.org/2751306/"] } }) {
items {
id
name {
value
}
dataset {
id
}
}
}
}Then filter works by the ids you got back – all of them, since each dataset has its own:
{
creativeWorks(
perPage: 0
where: {
about: {
in: ["https://id.drapo.nl/2d744075-7952-5e8a-ad0e-0cd664102d6c"]
}
}
) {
pagination {
total
}
}
}A term with no sameAs and no canonical id of its own cannot be matched at all, and deliberately so: nothing in the data says two similarly named local terms are the same concept, and the search API does not guess.
A term the authority could not describe keeps the publisher's. authority is absent on it, which is how you tell the two apart – and it happens for a reason worth knowing: a lookup by URI does not reach every term a vocabulary holds. Roughly a third of the Getty AAT references in the register resolve to nothing today, because the terms live in AAT sub-vocabularies that a URI lookup does not route to (netwerk-digitaal-erfgoed/network-of-terms#1943). Those keep an untagged publisher label until it is fixed upstream.
Places are the other case of the first shape, and the point of the next section: they are keyed on the term they align to whether the publisher used its URI as an id or stated it as sameAs, so one place is one document and one id.
Place references arrive through their own fields – locationCreated and contentLocation on a work, location on an organization, birthPlace and deathPlace on a person – and resolve against places, while about, material, genre and additionalType resolve against terms. A publisher may describe one place as either, so a location reference can be a Place, a DefinedTerm, or a node typed both.
A place is one place
Where a publisher aligns a place to a term in GeoNames, GTAA or Wikidata – or names it by that term’s own URI – the place is keyed on that URI, and every dataset referencing it points at the same document. Drapo and Limburgs Museum both hold something about Venlo, and it is one place, one id, one facet bucket:
{
places(where: { id: { in: ["https://sws.geonames.org/2745641/"] } }) {
items {
id
name {
value
}
latitude
longitude
addressCountry
authority
fetchedAt
}
}
}name, latitude, longitude and addressCountry come from the authority that owns the identity – authority names it, fetchedAt says when it was read – and not from the publishers, whose descriptions of a term they merely referenced are dropped rather than merged. Neither selected dataset ships coordinates at all, so this is the only place a map gets them.
The name is the one the authority states for the place itself: “Venlo”, not “Venlo (NL)”. GeoNames answers a suffixed label to tell homonyms apart in a picker – “Bergen (NL)” against “Bergen (NO)” – which is a disambiguation, not a name. It reaches you as addressCountry instead, so you can render or filter on it rather than parse it back out of a string.
A place nobody aligned keeps the publisher’s id and the publisher’s description. LOL does no entity matching: “Kessel” appears in both datasets as a local place and stays two documents, because neither publisher said they are the same. Such a place carries no authority.
The facet and the filter differ, deliberately. Only canonical places get a facet bucket, so an unaligned place cannot be reached by faceting locationCreated – but a filter on its id matches, and asking places for it works. Facets are discovery over a result set; filters are exact retrieval, and narrowing the first leaves the second whole.
Who made a work
creator and contributor are edges, not names. Each entry pairs the role a publisher stated with the agent it points at:
{
creativeWorks(query: "Maastricht", perPage: 10) {
items {
name {
value
}
creator {
role
creator {
id
name {
value
}
}
}
}
}
}This is schema:Role: the role rides the work→agent edge because it belongs to neither end: the same person is etser on one work and uitgever on another, and on one work an agent can be both. contributor has the identical shape, with the agent under contributor instead.
Reading an entry:
roleis one of the publisher’s own literals, and is often absent – plenty of collections credit an agent without qualifying how. A Role stated in two languages, or naming two agents, is that many entries, so that each one answers this agent in this role. They are strings, not IRIs: “etser”, “Etser” and “prentmaker” are three values.- the inner
creatoris the agent, resolved out of thepersonscollection, so it serves that document’s own fields:name,alternateName,authority, and the rest of One artist, however many collections credit them. Itsidisnullwhere the graph named the agent inline without a URI. - the entry itself has no
id. The Role node behind it is an artefact of how a publisher modelled the edge – a fragment IRI for one, a blank node for the next – so it is not served, and two statements of one role-and-agent pair are one entry.
Filter in either of two ways:
where: { creator: { in: ["https://data.rkd.nl/artists/104628"] } }
where: { creator: { where: { creator: { in: ["https://data.rkd.nl/artists/104628"] }, role: { in: ["etser"] } } } }The first asks for works crediting that agent in any role. In the second, sibling keys are welded to the same entry, so it asks this person as an etser rather than this person somewhere and an etser somewhere.
Facets on creator and contributor are keyed on the agent’s identity, never on the name, so two people who share one are never merged into a bucket. There is no facet over role, and cannot be: a facet cannot be scoped to the entry a filter matched, so it would count every role of every matching work.
One artist, however many collections credit them
Persons work like places. Where a publisher aligns a person to an authority record – RKDartists, Wikidata, the KB thesauri or GTAA – or names them by that record's own URI, the person is keyed on it, and every dataset crediting them points at one document:
{
persons(where: { id: { in: ["https://data.rkd.nl/artists/80083"] } }) {
items {
id
name {
value
}
alternateName {
value
}
birthDate
deathDate
hasOccupation {
id
name {
value
language
}
}
authority
fetchedAt
}
}
}name, alternateName, birthDate, deathDate and hasOccupation come from the authority, through the Network of Terms, not from the publishers. Discovery Museum publishes Max Verboeket as his RKD record with a name and nothing else; RKDartists adds “Verboeket, Maximiliaan Leo Hubert Maria”, 1922-07-08, 2015-01-23 and fourteen occupations, glasontwerper and glass designer among them. The alternate names matter most: an authority records the variants a person or firm is also known by – “Céramique Maastricht, N.V. Keramische Industrie genaamd” beside “Société Céramique, N.V.” – and those are searchable, so a visitor typing the trading name finds the works.
A date is one instant, where the authority may state a range. RKD gives Rembrandt’s birth as 1606-07-15/1607; birthDate serves the start of it, the same reading a work’s dateCreated gets, so a filter or sort on it treats the earliest possible date as the date. A birth the authority can only place in a decade is served as no date at all.
A person nobody aligned keeps the publisher's id and description, and carries no authority – the same rule as for places and terms.
An authority only describes what it names. A publisher may align a person to something that is not a person; where that happens the person keeps the publisher’s own node and name rather than borrowing a place’s. So a missing authority is not always a missing alignment – sometimes it is the alignment being declined. Where the node’s own id is a place – Limburgs Museum types eleven GeoNames features as schema:Person too – there is no person to keep, and persons answers nothing for that id; a creator or contributor entry pointing at it still carries the id and the publisher’s name for it, with no person document behind it. The reasoning is in the NDE alignment guide.
An occupation is an entry, and its id is usually null. The profile lets a publisher state an occupation as a term or, where none exists, as a resource with only a name, and RKD states every one of Verboeket’s fourteen the second way, once per language. Each is served as an entry with its name and no id, the same shape as a creator named inline on a work: displayed and searchable, filterable only where it has an id, and there is no facet over it. The reasoning is in the NDE alignment guide.
birthPlace and deathPlace are empty on a person the authority describes. The authority’s description replaces the publisher’s whole, and RKD names a birthplace by its own thesaurus, which no places document is keyed on, so nothing fills the field yet. A person nobody aligned keeps the publisher’s.
An institution appears twice
publishers and organizations are separate collections, and the same institution is routinely in both – under two different URIs, with one name.
- A publisher is an agent in a dataset's registration, identified by whatever URI the registrant used.
- An organization is an agent the object data relates to a work.
Neither URI resolves the other, so you cannot join them: to go from an institution to its works, start from publishers, get its datasets, and filter works by those – the two-query pattern under From one institution.
From one institution
There is no join across collections, so this is deliberately two queries: find the publisher’s datasets, then filter works by them.
{
datasets(where: { publisher: { in: ["https://www.tracelimburg.nl/"] } }) {
items {
id
name {
value
}
}
}
}Feed the returned ids into the second:
{
creativeWorks(
perPage: 10
where: { dataset: { in: ["https://id.drapo.nl/dataset/drapo-schemaorg"] } }
) {
pagination {
total
}
}
}Filtering datasets by publisher is the right direction: each dataset carries its own publisher, so the answer is complete. Going the other way – asking a publisher which datasets it has – is not currently reliable, because a document records only one dataset.
Describe one collection
Everything a search UI needs to know about one collection – which fields it can filter on and with what input, which keys it can sort by, and which facets it serves in which bucket shape – comes back from one introspection request, without walking the schema. The type is the unit of declaration, so ask for the type: <Type>Where for the filters, <Type>SortField for the sort keys and <Type>Facets for the facets.
{
where: __type(name: "CreativeWorkWhere") {
inputFields {
name
type {
...Ref
}
}
}
dateRange: __type(name: "DateRange") {
inputFields {
name
type {
...Ref
}
}
}
placeFilter: __type(name: "PlaceFilter") {
inputFields {
name
type {
...Ref
}
}
}
sortFields: __type(name: "CreativeWorkSortField") {
enumValues {
name
}
}
facets: __type(name: "CreativeWorkFacets") {
fields {
name
type {
...Ref
ofType {
ofType {
ofType {
fields {
name
type {
...Ref
}
}
}
}
}
}
}
}
}
fragment Ref on __Type {
kind
name
ofType {
kind
name
ofType {
kind
name
ofType {
kind
name
}
}
}
}For creativeWorks today this yields eighteen filter fields (about, material and genre take a TermFilter, contentLocation and locationCreated a PlaceFilter whose in is [IRI!], dateCreated and sdDatePublished a DateRange with min and max, hasMedia a Boolean, plus or and and), four sort keys (RELEVANCE, NAME, SD_DATE_PUBLISHED, DATE_CREATED) and fourteen facets, each with its bucket: IRIBucket (value: IRI!, count, label) for the reference fields, ValueBucket (value: String!) for type and temporalCoverage, BooleanBucket for hasMedia. Substitute another collection’s type name – PersonWhere, PersonSortField, PersonFacets – for the rest.
Enum values and bucket fields live on the named type, not on its wrappers: [IRIBucket!]! arrives as NON_NULL → LIST → NON_NULL → OBJECT, which is why the facets block digs through ofType three times before asking for fields, and why enumValues is asked of the enum type directly. Asked on a wrapper, both come back empty rather than failing.
What refers to an object
To answer “what refers to the object I am looking at, through any field?” – click “ink” as a term, get every work that has it as material, subject or genre – you need the list of fields that can hold an IRI at all. Don’t write that list down: ask the API, and it stays right when the schema moves.
The IRI scalar makes it decidable. A filter input takes one in list, and its element type says what the field holds:
input TermFilter {
in: [IRI!] # material, about, genre – identity
}
input KeywordFilter {
in: [String!] # name, identifier, license – literals
}A field references IRIs when its filter input’s in takes IRI. Nothing keys on a field name, so this works on any LDE schema, including one you changed.
Ask the schema
Two halves: which filter inputs take IRIs (asked once for the whole schema), and which filter input each collection’s where fields use.
{
filters: __schema {
types {
name
inputFields {
name
type {
kind
name
ofType {
kind
name
ofType {
kind
name
}
}
}
}
}
}
collections: __schema {
queryType {
fields {
name
args {
name
type {
name
inputFields {
name
type {
name
}
}
}
}
}
}
}
}Join the halves on the filter type name:
const namedType = (type) => (type.ofType ? namedType(type.ofType) : type);
// { TermFilter, PersonFilter, PlaceFilter, … } – the inputs that take IRIs
const iriFilters = new Set(
data.filters.types
.filter((type) =>
type.inputFields?.some(
(field) => field.name === 'in' && namedType(field.type).name === 'IRI',
),
)
.map((type) => type.name),
);
// { creativeWorks: ['id', 'dataset', 'creator', 'about', 'material', …], … }
const iriFields = Object.fromEntries(
data.collections.queryType.fields
.map((collection) => [
collection.name,
collection.args
.find((argument) => argument.name === 'where')
?.type.inputFields.filter((field) => iriFilters.has(field.type.name))
.map((field) => field.name),
])
.filter(([, fields]) => fields),
);or and and drop out on their own: they hold criteria, not filters, so they have no in. The collection list comes out of the same query, so a ninth collection needs no code change. Run this at boot and cache it – the surface only moves when the API is redeployed.
Why the shape: introspection has no string form for a type, so [IRI!] arrives as a wrapper chain (LIST → NON_NULL → SCALAR IRI) that you unwrap to the named type. And a query may nest at most two fields/inputFields lists – a third is rejected with Maximum introspection depth exceeded – which is why the IRI filter inputs are collected from __schema.types rather than walked per collection. type, ofType and args are free.
The query it builds
One or criterion per field, all carrying the same IRI. Drop id: that matches the object itself rather than what refers to it.
{
creativeWorks(
perPage: 10
where: {
or: [
{
about: {
in: ["https://id.drapo.nl/a005b7a1-562a-5d71-83c2-b26dd96dfae6"]
}
}
{
material: {
in: ["https://id.drapo.nl/a005b7a1-562a-5d71-83c2-b26dd96dfae6"]
}
}
{
creator: {
in: ["https://id.drapo.nl/a005b7a1-562a-5d71-83c2-b26dd96dfae6"]
}
}
]
}
) {
pagination {
total
}
items {
id
name {
value
}
about {
id
}
material {
id
}
}
}
}You need not know what kind of thing the IRI names. Every filter takes [IRI!], so a term IRI under creator is a valid criterion that matches nothing – a wrong guess costs recall of zero, not an error. Repeat per collection, aliased into one request, to cover objects of every type.
To show through which field the match ran, either ask for the reference fields on the items and compare them to the IRI you filtered on (as above), or give each field its own aliased sub-query with perPage: 0 – then the field is the response key and you get a count per relation.
Caveats
- Which collection a field points at is a naming convention.
material: TermFilterholds term IRIs, but nothing in the schema linksTermFilterto thetermscollection – you need that only to click through to the referenced object’s own record, since a reference already carries itslabel. Strip theFiltersuffix and look the rest up among the collections’ item types –Term→terms, and that half is introspected – which resolves all eight today. Assert it at boot, so a rename fails loudly instead of pointing a click-through at nothing. - Introspection must stay enabled. It is on the published image – the playground runs on it – and there is no other route to the field list.
inis the whole vocabulary for IRIs: set membership, no negation; the disjunction isor.