Skip to main content

Deposits and RO-Crates

Reading — entities, their metadata, files, and search — is mandatory for every implementation. Deposit is the write pathway, and it is optional core: part of the core specification, but a read-only catalog remains fully conformant without it. It is deliberately not entity-level CRUD — all writes flow through deposit sessions against RO-Crates, and catalog entities remain read-only projections that the implementation derives from what was deposited.

Every implementation declares its position in a required deposit capability block.

The deposit surface covers three things as one unit:

  • the deposit session endpoints — open a deposit, stage a metadata document and files, finalise
  • the RO-Crate read surface — list, retrieve, and delete RO-Crates, and fetch a deposited metadata document verbatim
  • the linkage fields on the read schemas (roCrateIds on Entity, roCrateId on File)

Why RO-Crates, Not Entity Writes

The unit a depositor actually holds is the whole RO-Crate, so that is the unit of deposit:

An RO-Crate is a metadata document plus all the files it references, deposited and stored as a unit.

Depositors send RO-Crates; the implementation materialises catalog entities from them by its own rules. There are no entity write endpoints (with one narrow exception for orphaned entities). The announcement post sets out the reasoning.

Materialisation

How an RO-Crate becomes catalog entities is implementation-defined. The contract is only that once a deposit reports complete, the entities materialised from it are readable, and the linkage fields below let clients traverse between the two surfaces. Two real archives illustrate how much the rules can differ:

  • PARADISEC: an RO-Crate is one item's metadata document and its media files. Materialisation yields the item entity plus one file entity per media file — a small, fixed shape.
  • LDaCA: an RO-Crate may describe a whole corpus. Materialisation explodes it into collection, item, file, person, and organisation entities — one deposit, many entities.

Materialisation can also be many-to-one: several RO-Crates may contribute to a single merged entity. If two deposited RO-Crates both describe the same speaker (same @id), an implementation may materialise one Person entity carrying both RO-Crates in its roCrateIds. This merge machinery is also what makes curation RO-Crates work.

Because entities are projections, re-materialisation can change them. A later deposit — of the same RO-Crate or a different one — may add, alter, or remove entities. Whether entities that lose their last contributor are pruned or retained is likewise implementation-defined (see Deletion & Lifecycle).

The Linkage Fields

The two surfaces — deposited RO-Crates and materialised entities — are bidirectionally linked:

  • entityIds on a RO-Crate: the entities materialised from its current version. This is the authoritative answer to "what did my deposit create", and it can change as later deposits alter the materialisation.
  • roCrateIds on an Entity: the RO-Crates whose current versions contribute to the entity. Usually one; more when materialisation merges contributions.
  • roCrateId on a File: the RO-Crate whose deposit supplied the file's bytes. Singular, because bytes arrive in exactly one deposit. Optional, to accommodate files predating any RO-Crate.

GET /entity/{id}/rocrate has implementation-defined provenance: the document may be a stored metadata document, or a view derived from the metadata documents of the entity's contributing RO-Crates. Either way it MUST be a valid RO-Crate whose root data entity describes the entity. To retrieve an original deposited metadata document verbatim, use GET /ro-crate/{id}/metadata.

Declaring the Capability

The deposit block in /capabilities is required of every implementation — read-only catalogs included. A client never has to infer read-only-ness from a missing key; each implementation says where it stands:

{
"apiVersion": "0.3.0",
"deposit": {
"supported": true,
"idMinting": "both",
"fileUpload": ["inline", "presigned"],
"depositTtlSeconds": 604800,
"maxFileSizeBytes": 5368709120
}
}

A read-only catalog declares the block just as plainly:

{
"apiVersion": "0.3.0",
"deposit": { "supported": false }
}

supported is the single flag clients check; when it is false the remaining fields MUST be omitted. The Capabilities guide documents what each field governs.

Deletion behaviour is declared separately, in the top-level tombstonePolicy: it governs entity and file URIs as well as RO-Crate ones, so every implementation declares it, deposit or not.

Deliberately not declared: whether finalise runs synchronously or asynchronously (server's discretion per request — one client code path handles both), and the RO-Crate visibility pattern (observable through behaviour; clients don't branch on it before acting).

Authentication

Deposit and RO-Crate write operations require the coarse OAuth2 write scope (see Authentication). Finer-grained authorisation — who may deposit what — is implementation-defined and expressed through ordinary 403 responses.

RO-Crate Visibility

Whether RO-Crates are readable beyond their depositor is implementation-defined. Implementations should pick one of two patterns and apply it consistently:

  • Public surface: RO-Crates are listable and retrievable by anyone, with the access object governing metadata-document retrieval. Recommended derivation: grant metadata access only if the caller has metadata access to every entity materialised from the RO-Crate, since the deposited metadata document is the union of their metadata.
  • Depositor-only surface: RO-Crate reads require the write scope — a management surface for depositors and curators, not a catalog surface.

The RO-Crate resource carries the access object either way, so client code is identical under both.

The Endpoints

OperationPurpose
POST /depositsOpen a deposit for a new RO-Crate
POST /ro-crate/{id}/depositsOpen an update deposit for an existing RO-Crate
GET /deposit/{id}Deposit state, staged files, recorded errors; the polling resource
PUT /deposit/{id}/metadataStage the metadata document (full replace)
PUT /deposit/{id}/file/{fileId}Stage a file (inline or presigned)
DELETE /deposit/{id}/file/{fileId}Remove a staged file
POST /deposit/{id}/finaliseValidate, publish, materialise
DELETE /deposit/{id}Abort an open deposit
GET /ro-cratesList RO-Crates
GET /ro-crate/{id}Retrieve an RO-Crate
GET /ro-crate/{id}/metadataThe deposited metadata document, verbatim
DELETE /ro-crate/{id}Delete an RO-Crate
DELETE /entity/{id}Delete a contributor-less entity

The guides walk through the flows:

Client Rules

  • Feature-detect before use: check capabilities.deposit.supported and read the declared modes rather than probing.
  • One code path for finalise: branch on the deposit's returned state, not on an expectation of sync or async behaviour.
  • Treat entityIds as live: the materialised-entity list reflects the current materialisation and can change as other deposits land.
  • Degrade gracefully: roCrateIds and roCrateId are optional; absent means the implementation (or that resource) doesn't carry them, not that the resource is invalid.