Modern biodiversity science generates enormous quantities of data, from DNA sequences and specimen images to species occurrence records and environmental measurements. But how do we manage that data responsibly, make it FAIR (findable, accessible, interoperable, and reusable), and integrate it across research infrastructures?
The Data Competence Center (DCC) is Naturalis's central hub for this expertise. We support researchers and projects in managing, analyzing, and sharing biodiversity data. By combining expertise in data stewardship, bioinformatics, and data architecture, we help Naturalis move from fragmented data swamps toward connected, transparent, and reusable data streams.
Servicesand expertise
The DCC provides tools, skills, and governance frameworks across three main areas, supporting researchers from the earliest stages of data collection through long-term archiving and publication:
Bioinformatics
Open Science
Data / Systems Architecture
Explore the sections below to learn more about our work and expertise within each domain.
Bioinformatics at DCC
We develop, maintain, and run bioinformatics pipelines for the standardized, reproducible processing and quality control of large-scale biological datasets. Our expertise includes DNA barcoding, metabarcoding, genome skimming, and DNA-based species identification.
We also collaborate with researchers to design custom analyses and support students and early-career researchers in developing computational skills. We manage access to high-performance computing (HPC) and workflow systems – including a shared Galaxy server – and maintain internal tools and services for Naturalis employees.
Open Scienceat DCC
We help researchers, projects, and infrastructures follow good open science practices. We develop metadata models and policies, advise on data management plans, and apply metadata standards. This ensures data is archived and published in line with the requirements of funding agencies and international repositories like GBIF, ENA, and BOLD.
We also supervise students and interns, and organize workshops for Naturalis staff on key open science topics – including research data management, computational reproducibility, and version control. Overall, these initiatives contribute to capacity building and culture change within Naturalis and research consortia.
Data / Systems Architectureat DCC
We model and integrate dataflows from diverse information systems, develop integrated data products, and maintain reference databases connecting collection, taxonomy, laboratory, and molecular data.
We contribute to international biodiversity data standards through active participation in bodies like DiSSCo, GBIF, TDWG, and RDA. We lead Naturalis's work on FAIR data policies for infrastructure projects and guide global processes around digital sequence information (DSI) and access and benefit sharing (ABS).
Whoworks here?
Dick Groenenberg, Bioinformatician
Dominika Kresa, PO, responsible for product vision, backlog, and stakeholder alignment
Jeroen Creuwels, Data manager
Julián López Gordillo, Data specialist
Luka Lenaroto, Bioinformatics employee
Pierre-Etienne Cholley, Bioinformatics employee
Rutger Vos, Department head, associate member
Sharif Islam, Data and system architect
Victor Heijke, Data and system architect
Wouter Addink, Coordinator international e-infrastructures and data
Yunhai Yi, Data steward
Partnersand networks
The DCC collaborates with a wide range of national and international partners to develop a shared biodiversity informatics infrastructure and exchange data, tools, and expertise. Our network currently includes partners such as:
GBIF — Global Biodiversity Information Facility
BOLD / iBOL — Barcode of Life Data System, International Barcode of Life Consortium
ENA / INSDC — European Nucleotide Archive / International Nucleotide Sequence Database Collaboration
DiSSCo — Distributed System of Scientific Collections
LifeWatch — European e-Science Infrastructure for Biodiversity
SURF — Dutch national high-performance computing and data infrastructure
NLeSC — Netherlands eScience Center
CoL — Catalogue of Life
TDWG — Biodiversity Information Standards
RDA — Research Data Alliance
OSNL — National Open Science Programme
TDCC NES — Thematic Digital Competence Center for National and Engineering Sciences
TDCC LSH — Thematic Digital Competence Center for Life Sciences and Health