Organize and publish datasets

From managed records to a citable version

Roles

  • Visitors and viewers read or download content allowed by policy.
  • Collaborators work with authorized records but do not manage membership, policy, or publication.
  • Dataset administrators manage access, participants, metadata, and versions.
  • Super administrators can maintain any dataset, but scientific stewardship and publication decisions belong to the responsible team.

Dataset versus dataset version

A dataset is a managed collection that may keep changing. It organizes records and participants and controls access. Page visibility does not by itself publish every record or grant a reuse license.

A dataset version is a fixed distribution snapshot with its own UUID, date, and files. Use the dataset page for ongoing work and the version UUID for links, citations, and reproducible analyses.

Before publication, review title, description, project, participants, access, license, use policy, download agreement, ordered authors and roles, references, record scope, filters, taxonomic-list sharing, related datasets, and metadata.

To publish, define version, date, scope, and filters; review authorship, license, policy, citation, and metadata; generate and follow the UserJob; open the UUID page; download the main and media files; and check the README, field definitions, counts, sample records, agreement, and access as a non-administrator.

Version files are persistent and must be included in backups. Correct the managed dataset and publish a new version instead of silently replacing files from an already cited version. Usage logs support reporting but do not change permissions or prove how data were used.