This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

User guides

Workflows for users and data stewards

These guides explain complete OpenDataBio workflows. See Concepts for object definitions and the API for integration parameters.

1 - Discover and map data

How to discover, filter, visualize, and export data

Who can use it

Visitors can search public content. Authenticated users can also see records released to their access level or made available through projects, datasets, or biocollections. The same filters may therefore return different results to users with different permissions.

Choose a tool

  • Use record lists to search one object type and follow its relationships.
  • Use Data Explorer to combine filters across related data records.
  • Use Map Explorer to inspect spatial distributions and open Location and Individual details.
  • Use the GET API or OpenDataBio-R for reproducible queries and analysis.

Start with the smallest useful set of filters. Check whether Taxon and Location filters include descendants, and use Project or Dataset filters when provenance or data policy matters. Inspect representative records before exporting.

Maps and exports

Map Explorer displays Locations, Individuals, plots, and transects. Large results use vector tiles and simplified geometries for navigation; use the official record or export geometry for analysis.

Small exports may be returned directly; large exports run as UserJobs. Follow the job progress and log, review warnings, then keep the README and field metadata with the downloaded data. A query export is a result at one moment; cite a dataset version when fixed, published content is required.

Exports may include OpenDataBio and Darwin Core fields. Botanical installations may also offer the BRAHMS/INPA profile for Individual- or Voucher-based exchange. Select the base record explicitly and review its metadata: the profile does not turn incomplete records into complete curatorial data.

Continue with Getting data with R.

2 - Data import workflow

How to prepare, import, reconcile, and validate data in stages

A reliable import is iterative: prepare one dependency, search existing shared records, send a small batch, recover generated IDs, reconcile them with the source table, and validate the result before importing dependent objects.

Prepare and order dependencies

Define the destination Project and Dataset, confirm permissions, preserve an immutable source copy, and add a unique local key such as source_row_id. Normalize encoding, dates, missing values, decimals, and column names. Search shared People, References, Taxons, Locations, and Traits before creating them, and check the POST API. Keep both the source key and ODB ID or UUID.

StagePrepare or findUsed later by
1People and Bibliographic Referencescollection, identification, measurements, Taxons, Datasets
2Taxonsidentifications, measurements, vernacular names
3LocationsIndividuals, measurements, spatial validation
4Traits, units, categoriesmeasurements and forms
5Project and DatasetIndividuals, Vouchers, measurements, media
6Individuals and occurrencesVouchers, identifications, measurements, media
7Vouchers and identification historymeasurements, media, references
8Measurements, Media, vernacular namesfinal dataset

Validate coordinates first

POST locations-validation accepts decimal latitude and longitude and reports registered Locations containing each point. Use it before creating Individuals or automatic point Locations to find swapped axes, wrong signs, unexpected administrative areas or protected areas, duplicates, and unsuitable precision. It reports relationships with existing Locations; it does not decide whether a coordinate is scientifically correct.

The UserJob cycle

  1. Submit representative pilot rows and save the UserJob ID.
  2. Follow status, progress, and logs; do not resubmit while processing.
  3. Review every row’s status, ID, recognition fields, errors, and warnings.
  4. Join results back to source_row_id; never rely on position after sorting.
  5. GET created or reused records and compare essential fields and relationships.
  6. Correct and resubmit only pending rows, preserving every job ID.
  7. Use validated IDs in the next dependency stage.

A successful job may still contain warnings or reused records. A useful working table keeps source_row_id, odb_status, odb_id, odb_uuid, odb_error, and odb_warning. An import is complete only when every row is documented, warnings reviewed, IDs reconciled, server records checked, counts and permissions verified, and source files and UserJob results preserved.

Continue with the R import tutorials.

3 - Organize and publish datasets

From managed records to a citable version

Roles

  • Visitors and viewers read or download content allowed by policy.
  • Collaborators work with authorized records but do not manage membership, policy, or publication.
  • Dataset administrators manage access, participants, metadata, and versions.
  • Super administrators can maintain any dataset, but scientific stewardship and publication decisions belong to the responsible team.

Dataset versus dataset version

A dataset is a managed collection that may keep changing. It organizes records and participants and controls access. Page visibility does not by itself publish every record or grant a reuse license.

A dataset version is a fixed distribution snapshot with its own UUID, date, and files. Use the dataset page for ongoing work and the version UUID for links, citations, and reproducible analyses.

Before publication, review title, description, project, participants, access, license, use policy, download agreement, ordered authors and roles, references, record scope, filters, taxonomic-list sharing, related datasets, and metadata.

To publish, define version, date, scope, and filters; review authorship, license, policy, citation, and metadata; generate and follow the UserJob; open the UUID page; download the main and media files; and check the README, field definitions, counts, sample records, agreement, and access as a non-administrator.

Version files are persistent and must be included in backups. Correct the managed dataset and publish a new version instead of silently replacing files from an already cited version. Usage logs support reporting but do not change permissions or prove how data were used.

4 - Import phylogenies into the backbone

Experimental import of trees proposed for taxonomic incorporation

Import compares source-tree terminals and relationships with existing Taxons. Labels must be reconciled with local taxonomic concepts. Importing a tree does not incorporate it automatically; it creates a proposal for review.

Preserve the source file, register its reference or DOI, check terminal labels and homonyms, create or validate missing Taxons through the normal taxonomic workflow, and document the target clade and intended interpretation. After import, review unmatched terminals, ambiguous matches, taxonomic conflicts, and relationships incompatible with the current backbone.

Backbone incorporation affects an installation-wide shared library and requires super-administrator review and approval. While experimental, this documentation does not promise general publication, sharing, analysis, versioning, or export of phylogenies; visible interface features may only support import review.

5 - Curate shared libraries

How to review Taxons, Locations, People, references, and vernacular names

Taxons, People, Bibliographic References, Locations, Traits, and vernacular names are installation-wide libraries. Search before creating, and remember that a global correction may affect many Datasets.

External Taxon validation

Full users may validate a Project, Dataset, or taxonomic root. Work on a small scope, inspect the UserJob, separate safe results from conflicts and missing Taxons, and choose a source appropriate to the group. Review Index Fungorum for fungi, Tropicos and IPNI for plants, and use GBIF as a broad source without assuming it resolves every conflict. Accept parent changes only when the local hierarchy should change. Apply permitted changes and submit suggestions for the others. Conflicts between sources require curatorial judgment.

Duplicate-Taxon merging is restricted to super administrators. Compare authorship, publication, validity, accepted name, parent, external keys, and descendant use. Homonyms and distinct concepts are not duplicates.

Shared Locations

Countries, administrative units, protected areas, Indigenous lands, environmental layers, plots, and transects are shared. Search name, hierarchy, type, and geometry first. Do not duplicate a country or municipality merely for another spelling or language. Coordinate new countries and large administrative imports with super administrators because they affect parent detection and coordinate validation.

For protected or environmental layers, record source, date, geometry version, and correct type; use WGS84; validate polygons; and check overlaps and existing versions. For plots and transects, use distinctive names and verify parent, geometry, orientation, dimensions, subplots, search width, datum, and units.

Some imports automatically create point Locations from Individual coordinates. Check whether that mechanism fits before creating points in bulk. Full users may edit Locations only while they have no linked Individuals, Vouchers, Measurements, or Media. Once used, only a super administrator may alter them; deletion also requires no descendants or related data.

Duplicate People and References

Search name variants, abbreviation, institution, email, and ORCID before creating a Person. Super administrators may merge duplicates and redirect collection, authorship, identification, measurement, expertise, and unpublished name relationships. Confirm the same real person, select the most complete primary record, inspect User links and conflicting roles, then review all relationships. Never merge homonyms.

Search DOI and BibTeX key before creating a Bibliographic Reference. If either already exists, compare and correct the existing record when it represents the same publication; do not change a key just to force a duplicate.

Vernacular names

Record language and use citations for source, context, and regional variation. The same spelling in different languages or regions does not necessarily express the same use. Full users may create names; ordinary users may edit only their own, and deletion is blocked by citations owned by others. Search name and language, inspect linked objects, and prefer adding a relationship or citation to creating a duplicate.

6 - Vouchers, labels, and requests

Biological collection operations and label preparation

Batch identification and Vouchers

Before identifying Individuals in a batch, review the selection, Taxon, identifiers, date, modifier, reference, and notes. Follow the UserJob and inspect Identification History, which preserves previous determinations.

A Voucher represents physical material from an Individual deposited in a Biocollection. Check its source Individual, collection and catalog number, collectors and date, nomenclatural type, references, and media. Vouchers inherit the Individual’s identification and Location. Attach a Measurement or Medium to the Voucher only when it describes that physical sample.

Generate labels

Select a small test set, object and sheet format, dimensions, margins, and content. Preview and print one sheet at 100%, verify codes, names, collection numbers, and identifiers, and only then generate the full batch. A saved preset stores print configuration, not a frozen copy of the selected records. Public presets may be reused or duplicated, but only their owner may change them.

Biocollection requests

Requests are available when at least one Biocollection is managed by the installation. They support two distinct cases: depositing material already represented by Individuals, and requesting material already deposited in a collection. This is separate from requesting access to a Dataset.

Deposit material and register Vouchers

Use this workflow after registering collection data as Individuals:

  1. review each Individual’s Location, collectors, date, and identification;
  2. select the Individuals and destination Biocollection;
  3. request Voucher registration and follow the UserJob that creates the request;
  4. a collection administrator or collaborator reviews the items and may record corrections or apply the requested identification update;
  5. when accepted, the collection selects the new Vouchers’ Dataset and first catalog number; OpenDataBio creates one sequentially numbered Voucher per Individual;
  6. review which items were registered or denied.

Only administrators and collaborators of the Individual’s Dataset may include it. The interface also requires these Individuals to belong to a Dataset open to the public or to registered users so collection staff can evaluate them.

Registration changes editing responsibility in a managed collection:

  • only collection members may edit its Voucher;
  • only users belonging to all managed collections linked to an Individual may edit that Individual, including its identification and Location;
  • the installation super administrator retains administrative access;
  • Measurements and Media do not pass to collection control and continue to follow their own Dataset permissions.

The depositor retains authorship and access defined by the Datasets, but cannot continue editing the collection-curated record unless also on collection staff.

Request a material loan

Use this workflow for Vouchers already held by a managed Biocollection:

  1. find and select the Vouchers;
  2. provide institution, contact, purpose, and intended conditions;
  3. follow each Voucher’s status separately;
  4. collection staff check the items and record the loan or denial;
  5. staff record return, or donation when that is the agreed outcome.

Implemented states distinguish requested, checked, lent, denied, returned, and donated items. They document processing in OpenDataBio; packing, shipping, deadlines, and institutional agreements remain the collection’s responsibility.

The requester creates and tracks the request. Collection administrators and collaborators may annotate and process it; for multiple collections, the user must belong to all of them. Editing administrative request data requires being an administrator of every involved collection, except for a super administrator. Status and history belong to each item, so one request can have mixed outcomes.

7 - Tutorials

Reproducible workflows with OpenDataBio-R

The tutorials apply the concepts and user guides through reproducible workflows with the OpenDataBio-R package.

Before importing, also read the Data import workflow, which explains dependencies, validation, and reconciliation of UserJob results.