Projects

Open Research Data Projects

Projects funded in the framework of the ORD Program

The joint ORD program of ETH Zurich, EPFL and the four research institutes of the ETH Domain has financially supported more than 60 research projects in the period 2020–2023. Funding supports researchers engaging in, or developing, ORD practices with and for their community and assists these researchers in becoming Open Research Data leaders in their field.

This page provides an overview of these projects. It highlights how researchers in the ETH Domain are currently applying ORD in exemplary ways. Some of the projects have already been completed, others are still in progress. The projects have been divided into three categories.

“Establish” projects help link existing ORD practices to a research agenda to establish them on a broader basis. They contribute to a shared and comprehensive understanding of ORD practices that can then become de facto standards.

“Explore” projects are the most extensive ventures in the program and are designed to explore and test early-stage ORD practices. The goal is to map processes of what an ORD practice might look like and develop prototypes. Through these projects, new teams form across disciplines and institutions.

“Contribute” projects help scientists integrate their research data into existing, often international, infrastructures. By standardizing the processes and making them generally accessible, the data are validated, and their potential is considerably expanded.

Filter

Category
Category
Institutions
MMS (Masonry MicroStructures database) - A 3D masonry microstructures database for advancing numerical research on irregular stone masonry structures

Category

Contribute

Institutions

EPFL

Data type

Microstructure database

Field

Materials Science

Researchers

Shah, Mati Ullah

Abstract

Stone masonry is an eco-friendly construction material, but its use has declined due to its vulnerability to earthquakes, mainly because of the poor arrangement of its microstructure. The microstructure includes the shape, size, and arrangement of stone units, which vary based on geographic, temporal, and material factors. Current building codes cannot fully account for this variability, and experimental studies are costly and impractical due to the diversity of masonry typologies. Numerical studies offer a solution, but creating realistic microstructures for modeling irregular stone masonry is complex and time-consuming. As a result, simplified microstructures are often used in simulations, which fail to capture the complexities of irregular masonry walls. To address this challenge, we have developed a 3D masonry microstructures database ready to use in numerical simulations. To enhance accessibility and usability, this project aims to create a web-based platform hosting this curated database of 3D microstructures and their geometric indices. The proposed web-based platform will also feature a tool for evaluating masonry quality using the Masonry Quality Index (MQI) from 2D images, promoting the preservation of historic structures and sustainable construction practices. Additionally, the platform will enable researchers to contribute and document new 3D microstructures, fostering collaboration and advancing numerical research on stone masonry.

Application Programming Interface for the River to Ocean Geodatabase for Education and Research

Category

Contribute

Institutions

ETH Zurich

Data type

Environnement

Field

Earth sciences

Researchers

Paradis, Sarah

Abstract

In order to advance our understanding of the carbon cycle, it is essential to evaluate the spatiotemporal variations of carbon between river and marine environments and gain insights into the pathways of carbon transfer from land to ocean. To do this, we need to work jointly with riverine and marine data, accounting for their temporal and spatial distribution. However, each of these systems have different data and metadata reporting strategies that need to be accounted for, which complicates their joint application. Efforts have been made to compile data from each of these systems into independent databases, but no attempt has yet been done to create a joint database of data of both of these systems while accounting for their different metadata. Hence, this project aims to bring together riverine and marine data into one database to easily query the data between both systems through the River to Ocean Geodatabase for Education and Research (ROGER). This database will be displayed in an interactive web-interface that queries riverine and/or marine data depending on the user’s requirements through a REST API. Harnessing the advanced geographical functions of PostgreSQL, the REST API will include functions that allow users to geospatially integrate riverine and marine data. This new database will provide a crucial step forward in the understanding of the carbon cycle along the land-ocean continuum, while ensuring that the data complies with best Open Research Data practices.

Development of standardized Respiratory Open Access Research

Category

Contribute

Institutions

EPFL

Data type

Medical data

Field

Life sciences

Researchers

Dan, Jonathan

Abstract

Chronic cough is a common condition globally. While efforts are being made to develop wearables to detect and quantify cough events automatically, such monitoring devices have not yet been incorporated into routine clinical practice due to a lack of consistency in their validation, resulting in slow progress and a lack of trust in reported results. We have identified three main reasons for this heterogeneity: 1) the clinical definition of different cough events and especially the delimitation of their beginning/end lacks standardization, 2) the data used is typically private and imbalanced with inadequate labelling as a result of the previous point, and 3) methodologies to assess the accuracy of event detection are different between research groups and often inappropriate. This proposal builds on ORD datasets, community guidelines, and standards to propose a unified framework for validating cough event detection algorithms. The main objective is the development of standards that will unify the workflow for validating respiratory event detection algorithms to ensure data adheres the principles of Findable, Accessible, Interpretable, and Reusable data. This will be distributed through a website, serving as a central hub and reference for standardizing clinical definitions and methodologies, leading to a future benchmarking platform for respiratory event detection algorithms.

Filter

Category
Category
Institutions
MiShMASh: Microbiome sequence and metadata availability standards

Category

Contribute

Institutions

ETH Zurich

Data type

Microbiome data

Field

Microbiome research

Researchers

Lina Kim

Abstract

Microbiome research relies on vast datasets, demanding unrestricted data access and consistent metadata. This project addresses two core issues: ineffective sequence data statements and inconsistent metadata standards. It proposes a dual solution - a tier-based FAIR ORD standard and compliance assessment software. The project contributes open resources for diverse users, including researchers, journals, and funders. The validation software allows users to evaluate adherence to data and metadata standards, enhancing data reporting for better accessibility, interoperability, and future reusability.

Enabling compliance with ORD standards for cutting-edge time-resolved experiments at high data-rates

Category

Contribute

Institutions

PSI

Data type

X-rays

Field

X-ray science

Researchers

Filip Leonarski

Abstract

The upcoming Swiss Light Source 2.0 machine upgrade and the advent of free electron lasers (SwissFEL) enable novel advancements in X-ray science. One emerging technique is time-resolved serial crystallography, providing insight into biomolecular processes at micro- and millisecond timescales but generating extensive data. A single experiment can produce a continuous stream of X-ray images at 2,000 images per second (17 GB/s), leading to terabytes of data. Managing such large datasets in public repositories is challenging. This project aims to enhance data accessibility by creating reduced datasets. Existing protein diffraction labeling algorithms will filter images to include only high-quality diffraction images, approximately 0.1-10% of the total. These reduced datasets will be available in the PSI public data repository, alongside the complete dataset, improving data interoperability and findability by adding labeling results to the metadata.

Pycsou FAIR: A Community Marketplace for Discovering and Sharing Image Reconstruction Plugins

Category

Contribute

Institutions

EPFL

Data type

Image reconstruction plugins

Field

Computational imaging

Researchers

Matthieu Simeoni

Abstract

Pycsou is an open-source Python computational imaging software framework. It natively supports hardware acceleration and distributed computing. Its microservice architecture ensures optimized and scalable computational imaging tools that are easy to share across modalities. The framework's domain-agnosticity keeps it lightweight, accessible, and portable. However, some imaging communities face challenges due to this generic nature. This project introduces Pycsou FAIR, a web platform, meta-programming framework, and interoperability protocol. It aims to enhance the discovery, development, and sharing of FAIR-compliant image reconstruction plugins at scale. Imaging scientists can readily integrate modern computational imaging methods into their processing pipelines.

Establishing Structures for an Efficient Management of Materials and Processing Data

Category

Contribute

Institutions

Empa

Data type

Coatings and thin films

Field

Coating Technologies (CT)

Researchers

Sebastian Siol

Abstract

Empa's Coating Technologies (CT) group excels in developing coatings and thin films for diverse applications. Our research focuses on advancing thin films through combinatorial techniques, yielding extensive datasets. However, many published datasets lack precise metadata. This project aims to automate metadata creation by developing software that interfaces with OpenBIS to record crucial parameters from deposition and measurement tools. Raw data will be stored on LAN servers, with standardized file structures. Software tools will enable data analysis (e.g., CombIgor) to access metadata and measurements. High-quality datasets in an open format can be used internally and published in open repositories like NOMAD or Zenodo. This solution benefits labs working on thin film deposition, with growing demand from partner organizations.

Standardizing Encoding Elements for Lensless Cameras

Category

Contribute

Institutions

EPFL

Data type

Open-source hardware and software toolkit

Field

Lensless imaging

Researchers

Eric Bezzam

Abstract

In this project, the objective is to expand an open-source hardware and software toolkit known as LenslessPiCam, which is designed for lensless imaging. Lensless imaging relies on an optical element that serves as a substitute for a conventional lens, typically a thin mask. Although various methods for creating these masks exist, there is currently no standardized approach that ensures both design compatibility and reproducibility. The aim of this project is to incorporate mask-design tools into LenslessPiCam, enhancing its relevance to the broader research community. Additionally, our methodology is influenced by the FAIR principles and utilizes readily available resources. This approach will enable LenslessPiCam to remain an affordable and high-performing toolkit for lensless imaging, catering to educators, hobbyists, and researchers.

Mitigating spaceborne radio frequency interference through satellite database

Category

Contribute

Institutions

ETH Zurich

Data type

Satellite signals

Field

Space Geodesy

Researchers

Matthias Schartner

Abstract

The tremendous growth of satellite mega-constellations like Starlink and OneWeb, emitting radio signals, poses a significant threat to radio astronomy. These satellite signals can lead to radio frequency interference (RFI) or oversaturation of the broad-frequency receivers in sensitive radio telescopes, causing data loss or observation failures. In response, the International Very Long Baseline Interferometry Service for Geodesy and Astronomy (IVS) community and the International Astronomical Union (IAU) have initiated working groups dedicated to measuring and mitigating satellite RFI. A promising strategy involves avoiding observations near potentially disruptive satellites. In this initiative, the plan is to contribute by establishing an open database containing satellite orbits and the associated frequency spectra emitted by satellites. This database will amalgamate existing orbit data with available frequency information and measurements obtained from observatories. It will be made openly accessible through a web interface and an application programming interface (API), ensuring easy integration into modern software workflows.

Traceable thermodynamic datasets for chemical modelling

Category

Contribute

Institutions

PSI

Data type

Thermodynamic datasets

Field

Chemical thermodynamic modeling

Researchers

George-Dan Miron

Abstract

Currently, thermodynamic datasets lack adherence to ORD FAIR principles. ThermoHub database offers access to meticulously curated and expert-documented thermodynamic datasets in an open-standard JSON format. This project's goal is to optimize ThermoHub and showcase its ORD capabilities by creating a comprehensive, user-ready database from various widely-used thermodynamic datasets for chemical modeling. The project also strives to design and provide a documented semi-automated workflow for future expansion and maintenance. This endeavor will standardize and harmonize the chemical thermodynamic modeling workflow, enhancing data quality, reliability, and traceability. Offering FAIR-compliant datasets will simplify modeling, eliminating the need for researchers to manually collect thermodynamic values from extensive literature or develop complex scripts for data import. ThermoHub will enhance collaboration across Swiss, European, and international projects by providing traceable thermodynamic data for diverse modeling applications. Additionally, it will support ongoing work at PSI/LES, EPMA, and ETHZ on thermodynamic database and modeling code development.

Building Open-Source Tools for reproducible interaction with biological ORD databases

Category

Contribute

Institutions

ETH Zurich

Data type

NCBI, EBI, MGnify

Field

Life sciences

Researchers

Nicholas Bokulich

Abstract

The life sciences increasingly rely on centralized online databases for sharing biological information (e.g., NCBI, EBI, MGnify). These databases enable the exchange of primary and secondary biological datasets for downstream re-use. However, technical challenges for depositing and accessing open research data (ORD) and issues with traceability and reproducibility hinder scientific progress and adoption of ORD/FAIR practices in the life sciences. This project aims to develop software tools to facilitate remote, programmatic, and fully FAIR interactions with prominent ORD resources in the biological sciences. These tools will remove existing barriers, promote ORD sharing and re-use, and foster community engagement in ORD practices. While this project focuses on microbiome research, its multidisciplinary nature ensures its relevance across various research domains, aligning with the Contribute program's objectives of contributing software tools for established ORD databases.

Open JMP - unlocking the potential of global indicator data

Category

Contribute

Institutions

ETH Zurich

Data type

Manual data structuring

Field

Process Engineering

Researchers

Elizabeth Tilley

Abstract

Decades of manual data structuring have resulted in the most comprehensive and internationally-comparable information on Water, Sanitation, and Hygiene (WASH) coverage. The WHO/UNICEF Joint Monitoring Programme for Water Supply, Sanitation, and Hygiene (JMP) maintains the database. The data are shared openly but in spreadsheet-based proprietary software, not following FAIR data principles. Data stored in spreadsheets underutilizes the potential those data could have for purposes other than the national, regional, and global progress monitoring in WASH. This project aims to unlock this potential by developing open-source data and software packages adhering to FAIR data principles for sharing within the WASH community and beyond. The process involves hosting free learning events using open-source computational tools to empower community members with FAIR data competencies.

Web

Application Programming Interface for the Modern Ocean Sedimentary Inventory and Archive of Carbon database

Category

Contribute

Institutions

ETH Zurich

Data type

Modern Ocean Sediment Archive and Inventory of Carbon (MOSAIC)

Field

Earth sciences

Researchers

Sarah Paradis

Abstract

The growing use of data repositories in marine geosciences has led to dispersed and unstandardized datasets, making compilation and harmonization time-consuming. To address this, the Modern Ocean Sediment Archive and Inventory of Carbon (MOSAIC) database was created, focusing on factors impacting organic carbon distribution in marine sediments. It has expanded significantly in coverage and complexity, stored as a PostgreSQL database. To enhance accessibility, this project will develop an Application Programming Interface (API) with a user-friendly web interface for querying MOSAIC. Python and R packages will also be created for researchers to integrate MOSAIC into their analysis workflows, ensuring reproducibility. This project promotes Open Research Data practices in marine geosciences.

Scroll to Top

Filter

Category
Category
Institutions