Projects

Open Research Data Projects

Projects funded in the framework of the ORD Program

The joint ORD program of ETH Zurich, EPFL and the four research institutes of the ETH Domain has financially supported more than 60 research projects in the period 2020–2023. Funding supports researchers engaging in, or developing, ORD practices with and for their community and assists these researchers in becoming Open Research Data leaders in their field.

This page provides an overview of these projects. It highlights how researchers in the ETH Domain are currently applying ORD in exemplary ways. Some of the projects have already been completed, others are still in progress. The projects have been divided into three categories.

“Establish” projects help link existing ORD practices to a research agenda to establish them on a broader basis. They contribute to a shared and comprehensive understanding of ORD practices that can then become de facto standards.

“Explore” projects are the most extensive ventures in the program and are designed to explore and test early-stage ORD practices. The goal is to map processes of what an ORD practice might look like and develop prototypes. Through these projects, new teams form across disciplines and institutions.

“Contribute” projects help scientists integrate their research data into existing, often international, infrastructures. By standardizing the processes and making them generally accessible, the data are validated, and their potential is considerably expanded.

Filter

Category
Category
Institutions
MMS (Masonry MicroStructures database) - A 3D masonry microstructures database for advancing numerical research on irregular stone masonry structures

Category

Contribute

Institutions

EPFL

Data type

Microstructure database

Field

Materials Science

Researchers

Shah, Mati Ullah

Abstract

Stone masonry is an eco-friendly construction material, but its use has declined due to its vulnerability to earthquakes, mainly because of the poor arrangement of its microstructure. The microstructure includes the shape, size, and arrangement of stone units, which vary based on geographic, temporal, and material factors. Current building codes cannot fully account for this variability, and experimental studies are costly and impractical due to the diversity of masonry typologies. Numerical studies offer a solution, but creating realistic microstructures for modeling irregular stone masonry is complex and time-consuming. As a result, simplified microstructures are often used in simulations, which fail to capture the complexities of irregular masonry walls. To address this challenge, we have developed a 3D masonry microstructures database ready to use in numerical simulations. To enhance accessibility and usability, this project aims to create a web-based platform hosting this curated database of 3D microstructures and their geometric indices. The proposed web-based platform will also feature a tool for evaluating masonry quality using the Masonry Quality Index (MQI) from 2D images, promoting the preservation of historic structures and sustainable construction practices. Additionally, the platform will enable researchers to contribute and document new 3D microstructures, fostering collaboration and advancing numerical research on stone masonry.

Application Programming Interface for the River to Ocean Geodatabase for Education and Research

Category

Contribute

Institutions

ETH Zurich

Data type

Environnement

Field

Earth sciences

Researchers

Paradis, Sarah

Abstract

In order to advance our understanding of the carbon cycle, it is essential to evaluate the spatiotemporal variations of carbon between river and marine environments and gain insights into the pathways of carbon transfer from land to ocean. To do this, we need to work jointly with riverine and marine data, accounting for their temporal and spatial distribution. However, each of these systems have different data and metadata reporting strategies that need to be accounted for, which complicates their joint application. Efforts have been made to compile data from each of these systems into independent databases, but no attempt has yet been done to create a joint database of data of both of these systems while accounting for their different metadata. Hence, this project aims to bring together riverine and marine data into one database to easily query the data between both systems through the River to Ocean Geodatabase for Education and Research (ROGER). This database will be displayed in an interactive web-interface that queries riverine and/or marine data depending on the user’s requirements through a REST API. Harnessing the advanced geographical functions of PostgreSQL, the REST API will include functions that allow users to geospatially integrate riverine and marine data. This new database will provide a crucial step forward in the understanding of the carbon cycle along the land-ocean continuum, while ensuring that the data complies with best Open Research Data practices.

Development of standardized Respiratory Open Access Research

Category

Contribute

Institutions

EPFL

Data type

Medical data

Field

Life sciences

Researchers

Dan, Jonathan

Abstract

Chronic cough is a common condition globally. While efforts are being made to develop wearables to detect and quantify cough events automatically, such monitoring devices have not yet been incorporated into routine clinical practice due to a lack of consistency in their validation, resulting in slow progress and a lack of trust in reported results. We have identified three main reasons for this heterogeneity: 1) the clinical definition of different cough events and especially the delimitation of their beginning/end lacks standardization, 2) the data used is typically private and imbalanced with inadequate labelling as a result of the previous point, and 3) methodologies to assess the accuracy of event detection are different between research groups and often inappropriate. This proposal builds on ORD datasets, community guidelines, and standards to propose a unified framework for validating cough event detection algorithms. The main objective is the development of standards that will unify the workflow for validating respiratory event detection algorithms to ensure data adheres the principles of Findable, Accessible, Interpretable, and Reusable data. This will be distributed through a website, serving as a central hub and reference for standardizing clinical definitions and methodologies, leading to a future benchmarking platform for respiratory event detection algorithms.

Filter

Category
Category
Institutions
The Imaging Plaza: A Curated Online Catalog of FAIR Imaging Software

Category

Explore

Institutions

EPFL

Data type

Imaging codes

Field

Biomedical Imaging

Researchers

Michael Unser, Laurène Donati, Oksana Riba Grognuz

Abstract

This project envisions The Imaging Plaza, an online catalog of FAIR imaging software for ETH domain scientists. The need for enhanced visibility and accessibility of Swiss imaging research output inspires this concept. The Imaging Plaza's goal is to simplify non-experts' access to imaging code produced by peers, enabling confident navigation through available options, providing guidance and incentives. Its aim is not to host code but to facilitate discovery. The joint effort involves the EPFL Center for Imaging and the SDSC, with deployment test sites at EPFL, PSI, and ETHZ. Expert curation ensures software is reusable and well-documented. SDSC engineers create shareable runtime environments. Users can navigate the catalog via a tailored search system and a common language (ontology) for imaging. A "FAIR levels" framework indicates code accessibility for untrained users.

Explore AiiDA for Regional Inverse Modelling of Greenhouse Gases

Category

Explore

Institutions

Empa

Data type

AiiDA

Field

Materials Science

Researchers

Stephan Henne

Abstract

Inverse modeling of greenhouse gas emissions, combining atmospheric observations and transport simulations, yields vital real-world emission estimates for policy relevance. However, the process currently lacks traceability, repeatability, and user-friendliness. This project proposes using the 'Automated Interactive Infrastructure and Database for Computational Science' (AiiDA) system to create prototype workflows and automation plugins for atmospheric modeling. AiiDA, well-established in computational material science at ETH, will be adapted for the inverse modeling community. Collaborating with Swiss and European partners, we aim to design generic workflows for broad adoption. This implementation enhances the traceability, repeatability, findability, and sharing of emission estimates, crucial for climate policy assessment. The project's impact extends to international and European collaborations beyond our group.

Development of Open Research Data Analysis Services supplementing astronomical Open Research Data

Category

Explore

Institutions

EPFL

Data type

Cloud-based ORD data analysis services (ORDAS)

Field

Astronomy

Researchers

Jean-Paul Kneib, Emma Tolley, Andrii Neronov, Volodymyr Savchenko

Abstract

In the past decade, astronomers have pioneered multi-messenger astronomy, merging electromagnetic signals from radio to gamma-ray wavelengths with neutrino and gravitational wave signals. This approach enhances our understanding of astronomical phenomena. The challenge lies in managing the diverse data sets, typically available as ORD, used in multi-messenger data analysis. Reproducibility and traceability of results to raw observational data are essential for implementing the "Findable-Accessible-Interoperable-Reusable" (FAIR) principle. The project offers a solution: cloud-based ORD data analysis services (ORDAS) to ensure result reproducibility and ORD reusability. These services will be developed for two next-generation facilities, the Cherenkov Telescope Array (CTA) and Square Kilometre Array (SKA). The power of ORDAS will be showcased through a Multi-Messenger Online Data Analysis (MMODA) platform, engaging the multi-messenger astronomy community in ORD practices for transient multi-messenger astronomical source studies requiring on-the-fly analysis of multi-messenger ORD.

Open and Reproducible Materials Science Research (PREMISE)

Category

Establish

Institutions

ETH Zurich, Empa, PSI

Data type

Electronic lab notebooks (ELNs)/lab information management systems (LIMSs), Workflow management systems (WFMSs)

Field

Materials Science

Researchers

Giovanni Pizzi, Edan Bainglass, Bernd Rinn, Caterina Barillari, Mihai-Cosmin Danaila, Juan Fuentes, Adam Laskowski, Henry Lütcke, Carlo Pignedoli, Fábio Da Costa Lopes, Aliaksandr Yakutovich, Simone Baffelli, Bruno Schuler, Corsin Battaglia, Nukorn Plainpan, Peter Kraus

Abstract

This project seeks to promote and establish FAIR ORD practices in Materials Science, with a focus on streamlining the treatment of experimental and simulation data on the same footing. Metadata standards for interoperability between electronic lab notebooks (ELNs)/lab information management systems (LIMSs) and workflow management systems (WFMSs) will be developed in the field of Materials Science, and combined with ontological semantic annotations. Best practices for integrating ORD into the research process will be collected, designed, and disseminated. Pilot use cases (focusing in particular on microscopies, spectroscopies, and battery research) that are applicable broadly to Materials Science research will demonstrate the deliverables. Open platforms like openBIS and AiiDA+AiiDAlab will be enhanced to meet FAIR requirements. These advancements will enable seamless interoperability between ELN/LIMS and WFMS. The ultimate goal of the project is to contribute to autonomous laboratories, where AI-driven simulations and robotic experiments expedite materials discovery and characterization.

See the website: https://ord-premise.org/

Open EM Data Network

Category

Establish

Institutions

ETH Zurich, EPFL, Empa, PSI, with swissuniversities partners UNIGE, UNIBE, UNIBAS, UNIL

Data type

Electron micrographs, tomograms, tilt series, diffraction images, spectrograms, electron density maps, atomic models, and processing workflows.

Field

Electron Microscopy

Researchers

Alexandra Radenovic, Andrzej J. Rzepiela, Bilal Qureshi, Christophe Copéret, Christophe Briand, Daniel Böhringer, Elisabeth Müller, Gebhard Schertler, Gregor Cicchetti, Henning Stahlberg, Marco Cantoni, Miroslav Peterek, Nicolas Blanc, Rolf Erni, Spencer Bliven, Stephan Egli, and Volodymyr Korkhov

Abstract

The Open EM Data Network (OpenEM) in Switzerland will implement ORD practices for Electron Microscopy (EM). Cryo-EM has revolutionized protein structure determination, while materials science has seen a surge in possibilities, including 4D STEM data. This has led to increased data volumes and computational needs. OpenEM will extend PSI’s SciCat Data Catalog to provide open and FAIR access to EM data Swiss-wide. It aims to standardize data and metadata collection, streamline acquisition, facilitate data sharing, automate deposition in international ORD repositories, offer user training, and establish a sustainable structure beyond the project’s closure. A parallel initiative through the swissuniversities CHORD program funds the participation of partners outside the ETH domain. OpenEM complements the “EM frontiers” initiative to advance electron microscopy technology in Switzerland, included in the Swiss Roadmap for Research Infrastructures 2023.

See the website: https://swissopenem.github.io/

Initiative for primary bio-NMR open research data

Category

Explore

Institutions

ETH Zurich

Data type

Protein Data Bank (PDB), Biological Magnetic Resonance Bank (BMRB), NMR spectra

Field

Nuclear Magnetic Resonance (NMR) spectroscopy in biology

Researchers

Roland Riek, Peter Güntert

Abstract

Nuclear Magnetic Resonance (NMR) spectroscopy, vital in structural biology, lacks an open database for primary data – multidimensional NMR spectra. These spectra are fundamental for in-depth protein analysis but remain largely inaccessible. The NMRprime initiative seeks to establish a FAIR-compliant database, integrating with the Protein Data Bank (PDB) and Biological Magnetic Resonance Bank (BMRB). NMRprime will facilitate spectrum uploads, automated annotation via machine learning, search capabilities, and open data access. The goal, in collaboration with PDB, BMRB, and journals, is to mandate NMR spectra deposition for bio-NMR projects, paralleling protein structure deposition in the PDB as is well-established for other methods in structural biology such as X-ray crystallography. NMRprime's expertise in bio-NMR and automated spectral analysis makes it the ideal candidate for this mission.

Scroll to Top

Filter

Category
Category
Institutions