Glossary (FAQ)
See also page Conventions & Terms for conventions and explanation of terms (words with certain meaning)
This page shall serve as a place to search for information related to Edmond.
In the following, miscellaneous information related to Edmond are collected, organized in a glossary style (some of them also in a question and answer style), with links to further information.
Name "Edmond"
Where does the name "Edmond" come from?
The first collection in Edmond was a set of images of the nucleus of the comet Halley. Therefore, the repository was named after Edmond Halley (1656-1742), an English astronomer who is best known for computing the orbit of the eponymous Halley's Comet and after whom the comet is named.
FAIR
The term "FAIR" stands for "Findable, Accessible, Interoperable, Reusable". One central goal of Edmond and other data repositories is to provide data in a FAIR way. More information about the FAIR principles can be found e.g. at www.go-fair.org.
Costs
What are the costs for using Edmond?
The usage of Edmond and also the Edmond support is free of charge to all members of the Max Planck Society. It is a service provided by the Max Planck Digital Library (MPDL), being a department of the Max Planck Information and Technology (MaxIT).
Data Management Plan
How can I refer to Edmond in my DMP?
You can mention Edmond as an open research data repository in your data management plan (DMP). Depending on the questions in your DMP, you can make use of the different answers given here. We also provide a summary of questions and answers in the Horizon Europe DMP template that affect Edmond. The document is only available in the MPG-IP-Range. If you need more information or feedback on your DMP please visit our Information Platform on Research Data Management. Also, please do not hesitate to contact our Edmond support team. They are happy to support you with everything you need.
Finding Edmond and its datasets
Can Edmond and the data published in Edmond been found globally?
Yes, Edmond is registered at the Registry of Research Data Repositories (re3data). Through re3data, Edmond is visible globally. All datasets published on Edmond get a persistent identifier (DOI), which is obtained by the DOI provider DataCite. Hence, all metadata for all research data published on Edmond appear in the DataCite Commons. Hereby, all published datasets in Edmond are listed in a global search engine. Edmond is also harvested by other aggregators like B2FIND.
Alternative repositories
There are many data repositories available - some are discipline-specific, others are institutional like Edmond. The Registry of Research Data Repositories (re3data) collects all known data repositories. This list can also be sorted based on the different disciplines.
Preservation Policy
Edmond provides a preservation policy here.
Archiving for long-term preservation
Edmond provides guidelines for long-term preservation here. They deal with the case that a DataSubmitter wants to preserve their data even beyond the preservation period of Edmond. See also here.
CoreTrustSeal
Edmond ist certified with the CoreTrustSeal.
Metadata schema
The metadata schema in Edmond is aligned to the DataCite Metadata Schema 4.4. The metadata schema is generic. Special requirements for metadata might be mapped within the provided metadata fields and metadata fields for controlled vocabularies.
ORCID
When you enter the author name in field "Name", you get a list of suggestions. Once you choose the approriate one, the ORCID field will be filled automatically. Otherwise you can fill in your ORCID manually. Please consider, that it must start with "https://orcid.org/".
ROR - Research Organization Registry
The ROR metadata field will be automatically filled by selecting your affiliation. If your affiliated institution is listed in ROR (which should be the case for all Max Planck Institutes), the ROR ID will be displayed in this metadata field.
Publishing requirements
All Max Planck employees and their collaboration partners can publish data on Edmond. A dataset must have at least one MPG affiliated author.
End of contract
What happens with the published research data, when the contract of the MPG author ends?
The published research data will still be available according to our preservation policy.
Storage limit
There is no explicit storage limitation in Edmond. However, when planning to store more than about 1 TB of data, please contact our Edmond support team beforehand. Furthermore, it is advisable for larger amounts of data to use the DVUploader or UpVerse, instead of uploading via the browser. Furthermore, there is an upper limit of 500 files per dataset. Please contact our Edmond support team if you have any questions.
Login, SSO
The standard way to log in to Edmond is using MPG Single Sing-on (SSO), thus all MPG accounts are ready to log in to Edmond without the need of an additional registration. For collaborators without MPG-account, local Edmond accounts are possible. See here for further information.
Password reset
You can reset your password for Edmond in the Single Sign-on (SSO) system. As an MPG employee and member, you can reset your password within your local IT infrastructure. Please contact your local IT staff for questions regarding your MPG account.
DOI
Digital Object Identifiers (DOI) are automatically provided for each dataset. The DOI is immediately reserved by creating a dataset. When a dataset is published, the DOI will then resolve. The DOI is always resolving towards the newest version of a dataset.
Preview URL for Reviewers
How can I share my dataset with others before publishing the dataset?
You can create a so-called "Preview URL" (sometimes also called "Private URL") for your unpublished dataset. Everyone having that link can access the dataset without login. This feature can be used in a review process, where reviewers want to see your unpublished data in advance. See also here for using a Private URL. How to create a Private URL is described here.
OAIS
Which data handling processes are defined for Edmond?
Edmond employs principles of the Open Archival Information System (OAIS) reference model. Details are given in the preservation policy.
Accessing published datasets
Who can search and access a published dataset?
When you published your dataset, anyone can search and access it. Edmond is an Open Research Data Repository, which includes openness by design.
Citing a dataset
You can find a citation suggestion for the dataset at the top of the dataset page. You can also download the citation in different file formats.
Mandatory metadata
What information is needed for a dataset?
Currently, per default following metadata are mandatory:
- Title of the dataset
- At least one author
- Name
- Organization
- Contact email address
- Description
However, for the sake of the FAIRness of the dataset, as much metadata as applicable should be provided. See page Guidelines for further information.
File Path - Structuring the files in a dataset
A folder structure is possible. You can generate the folder structure with different files in it, e.g. like in dataset https://doi.org/10.17617/3.7m. How to view the folder structure is shown here.
In order to create the folder structure, you can define a path in field "File Path", like desrcibed here. For larger files or multiple files, the tools UpVerse and DVUploader are comfortable alternatives to the web interface for creating such a file structure.
Deletion of a dataset
- An unpublished dataset draft can be deleted. All information and files will be deleted and cannot be restored.
- A published dataset cannot be deleted, however, a new version can be published, cf below.
- In exceptional cases one can ask for deaccessioning (i.e. withdrawing) a published dataset, cf here for details.
Changes and Versioning
As long as the dataset is in draft status, changes are possible, without storing previous states of the draft.
Once the dataset gets published, it gets a version (1.0 for the first version). This version is unchangable, i.e. neither the files nor the metadata can be changed any more. However, a new version can be published, where metadata may be updated or new files may be added. The new version will not receive a new DOI. Instead, the DOI for the dataset remains the same, regardless of the version. See here and here for details.
Licenses
Edmond supports flexible licensing. Following licenses are recommended for data:
- Creative Commons
Creative Commons licenses were originally created for the legally secure licensing of creative content, not primarily for data. But since version 4, this slightly different nature of data is kept in mind so that CC-BY 4.0 and CC-BY-SA 4.0 are also usable for research data. The advantage of CC licenses is their wide distribution and awareness.
By licensing your research outputs under CC-BY, your research is openly available, but it is required that others have to give you credit, in the form of a citation, should they use or refer to your research object. This license lets others distribute, remix, tweak, and build upon your work, even commercially, as long as they credit you for the original creation. This is the most accommodating of licenses offered. Recommended for maximum dissemination and use of licensed materials. - Open Data Commons
Open Data Commons has developed standard licenses directly for data, therefore they are likely to fit a bit better. But they are just not as widespread as CC licenses. - Public Domain
Not uncommon is the use of public domain waivers as CC0 or Open Data Commons Public Domain Dedication and License (PDDL) for research data. Here, the greatest possible freedom is granted in handling the data. Depending on the type of data, this can be useful, but it also leaves possibilities to misuse the data.
Although CC0 does not legally require users of the data to cite the source, it does not take away the moral responsibility to give attribution, as is common in scientific research. - Datenlizenz Deutschland
Through a collaboration between federal, state and municipal associations, Datenlizenz Deutschland has developed a recommendation for uniform conditions for the use of administrative data in Germany, which now exists as "data license Germany" in version 2.0. This version is officially accepted as an open license by the Council of Experts of the Open Definition.
Following licenses are recommended for software:
- MIT
The MIT License is a permissive software license, which puts few restrictions on reuse and has high license compatibility. - GPL
The GNU General Public Licenses (GNU GPL or simply GPL) are a series of widely used free software licenses. The GPL is a copyleft license, cf Wikipedia for further info.
If no license is manually assigned, the data within Edmond will automatically be released to the public domain under CC0.
See also:
- here is shown where to find the licence of a published dataset on Edmond.
- here for legal notes.
- A tabular overview of the different licenses can be found on the German page Forschungslizenzen.
Mybinder and Wholetale
Can I explore data on Edmond with other services?
Yes, it is possible to explore the published data on Edmond with other services. Edmond mainly supports the Mybinder service and Whole Tale. If you want to start data in Mybinder, please select "Dataverse" as the starting repository and enter the corresponding DOI of the dataset. The Mybinder service can then be launched. For analysis in Whole Tale, you must first register the desired data in your workspace. Please see the corresponding Whole Tale instructions.
However, these features are considered as "nice to have", without any guarantee related to their availability and persistence.
File types
What types of data files can I upload?
In principle, on Edmond, all files formats can be uploaded, including your software code. However, these should be formats that are reasonable for the use and reuse of your dataset, and which are suitable for long-term archiving. Therefore, only few file formats are accepted for curation level EnhancedReview. For recommendations regarding file formats see here, see also our Information Platform on Research Data Management.
Size limits
Currently there is no explicit file size or dataset size limit in Edmond. However, when planning to store more than about 1 TB of data, please contact our Edmond support team beforehand. Furthermore, we have set an upper limit of 500 files per dataset. See here for details, or contact our Edmond support in case of further questions.
Zip file
For bigger sets of file items we recommend to upload ZIP files instead of many single files. With an integrated ZIP viewer, the ZIP archives can directly be viewed on Edmond. The folder structure including files of a ZIP archive will automatically be displayed. At the same time, individual files can also be downloaded from the ZIP archive, so that it is no longer necessary to download the entire ZIP file. This gives users a quicker insight into the published ZIP file.
Direct download link
Edmond offers direct download links to individual file items. For datasets with CC0 license, the direct link for each file item is displayed in tab "Metadata" of the individual file.
For example, https://edmond.mpg.de/api/access/datafile/336721 is the direct download link for file "Marcgravia_Flowers_Interpolated.csv" in dataset doi:10.17617/3.VR8C1M.
A direct download link for the whole dataset can be generated with prefix https://edmond.mpg.de/api/access/dataset/:persistentId?persistentId=doi:10.17617/, followed by the last numbers and digits of the datasets DOI - for example: https://edmond.mpg.de/api/access/dataset/:persistentId?persistentId=doi:10.17617/3.VR8C1M for dataset doi:10.17617/3.VR8C1M.
Deleting files
How can I delete a file?
Deleting an item from a published dataset will lead to a new version of the general dataset. See also the answer on versioning in Edmond. Please note that the file is still available in the previous version.
Moving files
How can I move a file?
This is possible via edit metadata, where also the "File Path" is available. Through this, moving a file to different folders is possible.
Jupyter
Are there special features for Jupyter notebooks?
Yes, we offer additional features for datasets with Jupyter notebooks.
Datasets with data and and a .jypnb file can be executed in the Mybinder service.
In addition, a Mybinder button can also be integrated into the dataset description.
The following html syntax is necessary for this, whereby XXX must be replaced by the dataset DOI:
<a href="https://mybinder.org/v2/dataverse/10.17617/XXX/" target="_blank"><img src="https://mybinder.org/badge_logo.svg" class="img-responsive"></a>
This allows the dataset to be started directly in Mybinder without the user having to set the Mybinder configurations.
When this is entered, the Mybinder button appears in the desired position in the description, like this
Dataverse
Which technology is used for Edmond?
Edmond uses the Dataverse software, an open-source community-driven development.
Safety
Is my data safe with Edmond?
Edmond uses diverse strategies to back up the data (back up of the server and of the database itself). These backups are created in frequent intervals (the shortest is every 12 hours) and are stored in several locations on servers of the Max Planck Society as well as on tapes.
Upload tools
Does Edmond provide an upload client?
Yes, several tools are available. We recommend to use these tools instead of the web browser in case of larger files (more than one few GB):
- Command-Line and programmatic clients:
- Java DVUploader: A command-line bulk uploader for Dataverse installations based on Java.
- Python DVUploader: A Python library which can be used programmatically or via Command Line interface. Examples and instructions can be found on the GitHub page.
- Desktop Client
- UpVerse: If you prefer a graphical interface, you can download UpVerse, a desktop application for Windows, MacOS and Linux. UpVerse allows to easily upload folder structures to Edmond. For more information see the UpVerse Wiki.
How can I upload a larger file to a dataset using the Java DVUploader command-line client? ToDo: Markus fragen ob noch stimmt.
- Prerequisite: Java 8 (or greater)
- Download the DVUploader
- Upload your file to an existing dataset by executing the following command:
java -jar DVUploader-v1.1.0.jar -directupload -server=https://edmond.mpg.de -key=<api key> -did=<dataset doi> <path to file>
Therein, following parts need to be adapted:<api key>is your API key (to get your API key, login to Edmond, click on your name, select "Account Information" and then "API Token")<dataset doi>is the DOI of the dataset in the formatdoi:10.XXXXX/XXXXXX(e.g.doi:10.17617/3.ABCDE)
- Further information and configuration options can be found on the GitHub page.
API
Does Edmond provide an API?
Yes, an API is available. The starting point is https://edmond.mpg.de/api/. The documentation of the Dataverse API is available in the Dataverse Guide.
Community work and further reading
- Poster presented at Dataverse Community Meeting 2026 in Barcelona: "Establishment of a new curation process in Edmond - the Max Planck Society open research data repository"