Data Access & Usage - Information for Data users
This page contains information for persons who want to search, download or reuse data published on Edmond.
General Notes
Accessibility
All datasets published on Edmond are openly accessible, without the need of having a user-account. An exception are draft version of datasets, which are not published yet.
Licences
Downloading data from Edmond requires acceptance of Edmond's terms of use, in particular, the acceptance of the licence which has been given to the dataset.
See also here for further legal notes.
The licence given to a dataset can be found in the "Terms" tab of the dataset:

See also here for recommended licences when publishing a dataset.
Data access and download
When being on the page of a dataset, several options are available for downloading files:

- Option 1: Download all files
For that, click on button "Access Dataset" and then on "Download ZIP".

- Option 2: Download only one file
For that, select the file of interest, then click on iconand choose the link in the middle (in this example: "Plain Text").

- Option 3: Download several files
For that, select the files of interest, then click on the "Download" button.

In case of options 1 and 3, the files are collected in one ZIP file for being downloaded. However, there is an upper size limit of 4.7 GB for these options. In case of option 2, the selected file is provided in its original state, without being zipped.
Table and Tree view
Per default, all files are in the root directory within a dataset. However, in some datasets, e.g. in https://doi.org/10.17617/3.7m, file paths are used for organizing the files in a folder structure. For such datasets, you can choose between the "Table View" and the "Tree View":

The Table View (see screenshot above) is presented per default, showing some file details, and for some image formats also thumbnails. Click the "Tree" button to switch to the Tree View (see screenshot below), which shows the folder structure. This view can be helpful for getting an overview and for finding key files like the Readme file, which is typically in the root directory of the dataset.

Preview URL for dataset drafts
A special case are drafts of datasets, which are not published yet. Read access to such a dataset draft can be granted by a so-called "Preview URL". If you have received such a Preview URL, you can use it to see the draft, also to download data as described above. However, be aware, that such a dataset draft may be changed or deleted, and that the Preview URL will not be available any more once the dataset is published. Thus, the Preview URL is not suitable for persistently citing a dataset.

See here for the description how to create a Preview URL.
Versioning
In order to provide the possibility to extend, update, improve or correct a dataset, Edmond supports versioning.

A published dataset may have one or more version. Once a dataset gets published the first time, it gets version 1.0.
In case of changes, a new version of the dataset is created.
Typical reasons for changes are updates in the metadata, like updates of links to other sources or extension or correction of the dataset description.
Usually, for such changes, the digit after the dot in the version number will be increased by one, i.e. version 1.0 is followed by 1.1, 1.2 and so on.
When files are added or removed (modifying a file would be interpreted as removing one and adding a new one), the digit before the dot is increased, with digit after the dot being 0 - e.g. version 1.2 will be followed by 2.0.
The DOI link of a dataset resolves to the newest published version of the dataset, but also the previous versions can be accessed. For that, go to the "Versions" tab of the dataset and select the desired version in column "Dataset Version".

Because all versions of the dataset share the same DOI, we recommend to add the version number when citing the dataset, in case you want to relate to a specific version, for example:
Walter, David, 2022, "Code for dealing with data format CARIBIC_NAmes_v02", https://doi.org/10.17617/3.WDVSU7, Edmond, V1
Persistence and Curation
Dataset drafts
Before a new dataset gets published, an unpublished dataset draft is created. Per default, such a draft is not publicly available, but needs login and rights for being accessed (cf here. However, read access can be provided via Preview URL. Such a draft is not considered to be persistent: It may be modified or even deleted. Although a DOI is already reserved for such an unpublished draft, the DOI link does not resolve yet.
Also between a published version and its sucessor version, a draft version is created. Again, such a draft version is not considered to be persistent. At that stage, the DOI link is still pointing to the latest published version.
If a draft is not changed or published for a longer than a year, it may get deleted by the Edmond team after contacting the contact person of the draft.
Curation levels
Each dataset version has a curation level, which is stored in metadata field "Curation Level". Currently, following levels are defined for Edmond:
- Pre26
Dataset versions having been published before 2026-07-20, i.e. before the establishment of the new Edmond curation workflow. In those cases, field "Curation Level" is empty. - PreliminaryReview
Dataset which has received a basic review by a DataReviewer, but where a version with updated metadata is expected.- Typical use case: For the last steps in publishing a paper, the supporting data may be requested to be already published on Edmond, while the DOI of the paper is not known yet. Then the first dataset version may get curation level "PreliminaryReview". Then, once the paper is published and its DOI is known, an updated dataset version 1.1 is created with curation level "BasicReview" or "EnhancedReview", where the paper is referenced properly.
- BasicReview
Dataset which has received a basic review by a DataReviewer.- Such a basic review includes a check for the existence of basic metadata like authors, dataset title, description etc. See page Data Guide for the expected metadata.
- However, the content of the files themselves is usually not or only partly checked by Reviewer. In particular, no guarantee can be given by the Edmond team related to the quality of the data.
- EnhancedReview
Dataset which has received an enhanced review by a DataReviewer.- Such an enhanced review includes the checks like for the BasicReview.
- The dataset shall contain an explicit description of each file of the dataset, giving information about the file format and content.
- For tabular data, each column has to be described.
- Usually, this information is given in one or several separate Readme files. In some cases, these information may be part of the dataset description or in the header of the indiviudal data files
- The files are checked for interoperability and their suitablility for long-term archiving
- Compromises may be made in case of large datasets, where specific file formats may be chosen due to storage consumption and performance aspects.
- Exceptions can be made for additional files.
- Example:
A dataset may contain the processed data of a lab experiment in form of csv tables with proper description of the individual columns and further relevant metadata, thus justifying for level "EnhancedReview". As bonus on top, proprietary raw data files of the used instruments (e.g. setting files) are added to the dataset. In that case, these addition file do not reduce (but rather enhance) the reusability of the dataset. Thus, "EnhancedReview" is justified for such a dataset although including proprietary (and maybe not well documented) files.
- Example:
- Additionally to DataSubmitter, a DataCarer is defined, who takes care about the dataset and who can act as contact for the case that DataSubmitter would not be reachable any more. See here for details.
Persistence of published datasets
Once a dataset is published, it is considered to be persistent in the following sense:
- New versions of a dataset may be published under the existing DOI, but the old versions still remain accessible.
- Such new versions are usually created by the DataSubmitter, but might also be created by the DataCarer, DataReviewer or the Edmond team for curation purposes or in the context of technical updates of Edmond. See here for the explanation of these curation roles.
- In exceptional cases, a published dataset or a published version of a dataset might get deaccessioned, see below.
- A certain published version of a dataset remains unchanged:
- Bitstream-preservation of the files: The content of the files of a dataset are considered to be the essential good to be preserved by Edmond, thus the file content remains unchanged. Here, file content means the sequence of bytes, thus a checksum like MD5 may be used for checking the integrity of the file.
- The contents of the metadata fields like Title, Authors, Description etc are intended to remain unchanged, but smaller modifications to the updates or general curation tasks of Edmond might occur. For example, currently the description field allows for some html-based formatting, a feature which is not guaranteed to remain unchanged in future.
- Therefore, we recommend to include content, which must remain untouched, within the uploaded files.
- Duration - how many years
- Following the rules of good scientific practice, the datasets will be preserved for at least 10 years.
- However, existing datasets are handled according to our preservation policy, thus the 10 years are the minimum preservation period. As Edmond started in 2014, some of the datasets are already older than 10 years, and also datasets are planned to be preserved for more than 10 further years.
- Deletion scenario
- In the unexpected case that there should be a need for deletion after 10 years due to resource considerations, datasets of level "PreliminaryReview" would be considered to be of lower priority to keep, while datasets of level "EnhancedReview" would be considered to be of higher priority to keep (thus at least 15 years would be seeked at).
- Furthermore, an email would be sent to DataSubmitter (and DataCarer, where applicable) for informing them about the planned deletion.
- Migration scenario
- In the unexpected case that Edmond should be closed within 10 years, the datasets would be migrated to another repository.
- This would include the redirection of the DOIs.
- Long-term preservation
- Edmond provides guidelines for long-term preservation. They shall help to find an archive solution for preserving data far beyond the aformentioned 10 or 15 years. Usually, only datasets of level "EnhancedReview" are considered to be suitable candidates for that process.
Deaccession
Once a dataset is published, it can normally not be withdrawn. However, for exceptional cases like unforseeable data protection issues, ownership conflicts or contravention, the Edmond team has the technical possibility to deaccession (i.e. withdraw) a dataset even after it was published. After that, only metadata and a comment are available for other users, but not the files anymore. The DOI is still resolving. However, such a deaccession must be a very exceptional case, which needs to be well-argumented and discussed between the persons claiming the deaccession, the Edmond team and potentially further affected parties. Furthermore, the persons claiming the deaccession must provide a statement which explains the reasons, and which gets published.
Metadata
As mentioned above, the content of the files in Edmond are the central good to be preserved. However, an Edmond dataset additionally contains various metadata, some of them being essential for the FAIRness of the dataset.
The following screenshot shows the overview page of an example dataset.
- In the upper part, title and version of the dataset are given
- The box below contains a proposed text for citing the dataset.
- Below that a selection of metadata is shown, namely the Description, Keywords, and the License.
- In the lower part, 4 tabs are shown, namely:

Tab "Files" - file-related metadata
For each file, following metadata exist:
- File Name
- File Path
- Per default, this is empty, meaning that all files are in the root directory within a dataset.
- However, file paths can be used to mimic a folder structure.
- Description
- A description of the file. An alternative or additional place for describing a file is the Readme file.
- Tags: A file can have an optional tag.
- The standard tags and their use-cases are:
- "Code": Files containing software code or scripts (e.g. Python, R etc)
- "Documentation": Intended for Data Processing Reports, Standard Operating Procedures, etc
- "Data": Raw data, processed data, e.g. in tabular form
- Also custom tags are possible.
- A file may have less or more than one tag.
- The standard tags and their use-cases are:

Besides those metadata, further file-related information are shown in the screenshot above, like the file size, and the MD5 hash of the file.
Tab "Metadata"

Following a list of metadata presented in that tab. The order of the fields and the naming slightly differs between the view of a user reading the dataset, and the DataSubmitter when modifying the metadata. For the latter case also see page Data Guide.
- Persistent Identifier
Edmond uses DOI as persistent identifier, e.g.doi:10.17617/3.2G. One DOI refers to one dataset. In case there are more than one versions of a dataset available, the DOI refers to all of them, but the DOI link (e.g.https://doi.org/10.17617/3.2G) resolves to the newest version of the dataset. - Publication Date
The publication date is the date when the first version of the dataset has been published. It is filled automatically by the Edmond system. - Title
Title of the dataset. - Author
One or several authors of the dataset, being responsible for creating the work contained in the dataset. Usually, these are persons, but also entities like teams or roles are in principle possible.- ORCID: In case the ORCID of the author is provided, the ORCID logo
is shown behind the author name, containing a link to the ORCID profile.
- Organization: Affiliation of the author
- ROR ID of the Organization: In case the ROR ID of the organization is provided, the ROR logo
is shown behind the organization name, containing a link to the ROR profile.
- ORCID: In case the ORCID of the author is provided, the ORCID logo
- Study Type
Gives a rough hint about the study type. Available choices are: "observational", "experimental", "simulation/modelling", "derived/compiled", "survey", "other". A dataset can have more than one of these study types. - Description
Text describing the dataset and its content. - Subject
Domain-specific subject - categories that are topically relevant to the dataset. Available choices are: "Humanities", "Social and Behavioural Sciences", "Biology", "Medicine", "Agriculture, Forestry and Veterinary Medicine", "Chemistry", "Physics", "Mathematics", "Geosciences", "Mechanical and Industrial Engineering", "Thermal Engineering/Process Engineering", "Materials Science and Engineering", "Computer Science, Systems and Electrical Engineering", "Construction Engineering and Architecture". A dataset can have more than one of these subjects. - Keyword
Key terms that describe important aspects of the Dataset. Choosing proper keywords help to increase the findability of the dataset: One the one hand, one can filter for keywords (facettes on the lift side at the Edmond main page). On the other hand, they are considered in the search function. - Topic Classification
Here, terms with a certain meaning according to some kind of controlled vocabulary can be provided. This can be used for providing additional metadata in a structured way. In particular, "Project" is often used as so-called "Controlled Vocabulary Name", in combination with the name of the project (as so-called "Term"). For example, in dataset doi:10.17617/3.P0PHTC, the Controlled Vocabulary Name is "Project" and the term is "ATTO", in the metadata tab shown asATTO (Project). - Language
Languages which are used in the files of the dataset. In most cases, this is English, also in case this metadata field is not filled out. - Funding Information
Information about the datasetʼs financial support. This is a pair of fields named "Agency" and "Identifier", where the first one contains the ageny providing the financial support (e.g. "Horizon Europe"), and the latter one some identifier (like award number or grant identifier) for the project or grant (e.g. "101133587"). In that example, they would be shown asHorizon Europe: 101133587. - Depositor
The entity, that deposited first version of the dataset in Edmond. It is automatically filled by the Edmond system and shows the Familiy Name and the Given Name having been stored in the the user account of the person who has created the first draft of the dataset. - Deposit Date
The date when the dataset was deposited into the repository. It is automatically filled by the Edmond system. - Software
Information about the software used to generate the Dataset. In case of processed data, this is usually the processing software, e.g. "R", "Python" etc. Also a version of the software may be provided. - Related Publication
Article or report that is related to the dataset. In case available, also a persistent identifier (e.g. a DOI) should be provided. - Related Dataset
Information about a related dataset, e.g. the raw data having been used to generate the data of this dataset. - Geolocation
Information about the Geolocation where the data were collected.
Further metadata
- Contact
The entity, e.g. a person or organization, that users of the dataset can contact with questions. There, an email address is provided. That email is not shown to the users, but it is used when a user clicks on button "Contact Owner". - Other Identifier
Further identifier for the dataset. This metadata field is predominantly used in older datasets.
Search for data
Search for projects
For some projects, the datasets provide a so-called "Topic Classification" as metadata (cf above in metadata description), with "Project" as Controlled Vocabulary Name.
In order to find datasets of project "ATTO", you can enter following in the search field: topicClassValue:"ATTO" AND topicClassVocab:"Project".

However, typically not all the datasets related to a project provide this metadata. Thus in such case more hits are expected when just entering the project name in the search field, because then also further metadata fields like title, description and keywords are searched through.
