{"type": "FeatureCollection", "features": [{"id": "10.1038/s41598-020-58025-3", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:16:49Z", "type": "Journal Article", "created": "2020-01-28", "title": "Building de novo reference genome assemblies of complex eukaryotic microorganisms from single nuclei", "description": "Abstract<p>The advent of novel sequencing techniques has unraveled a tremendous diversity on Earth. Genomic data allow us to understand ecology and function of organisms that we would not otherwise know existed. However, major methodological challenges remain, in particular for multicellular organisms with large genomes. Arbuscular mycorrhizal (AM) fungi are important plant symbionts with cryptic and complex multicellular life cycles, thus representing a suitable model system for method development. Here, we report a novel method for large scale, unbiased nuclear sorting, sequencing, and de novo assembling of AM fungal genomes. After comparative analyses of three assembly workflows we discuss how sequence data from single nuclei can best be used for different downstream analyses such as phylogenomics and comparative genomics of single nuclei. Based on analysis of completeness, we conclude that comprehensive de novo genome assemblies can be produced from six to seven nuclei. The method is highly applicable for a broad range of taxa, and will greatly improve our ability to study multicellular eukaryotes with complex life cycles.</p>", "keywords": ["0301 basic medicine", "Evolutionary Biology", "0303 health sciences", "Genome", "Fungi", "Computational Biology", "Eukaryota", "Genomics", "Article", "Workflow", "Evolutionsbiologi", "03 medical and health sciences", "13. Climate action", "Algorithms"]}, "links": [{"href": "https://www.nature.com/articles/s41598-020-58025-3.pdf"}, {"href": "https://doi.org/10.1038/s41598-020-58025-3"}, {"rel": "related", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/Scientific%20Reports", "name": "related record", "description": "related record", "type": "application/json"}, {"rel": "self", "type": "application/geo+json", "title": "10.1038/s41598-020-58025-3", "name": "item", "description": "10.1038/s41598-020-58025-3", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/10.1038/s41598-020-58025-3"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2020-01-28T00:00:00Z"}}, {"id": "10.1101/2024.10.22.619569", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:17:23Z", "type": "Journal Article", "created": "2025-07-17", "title": "Metagenomics-Toolkit: the flexible and efficient cloud-based metagenomics workflow featuring machine learning-enabled resource allocation", "description": "Abstract                   <p>The metagenome analysis of complex environments with thousands of datasets, such as those in the Sequence Read Archive, requires substantial computational resources for it to be completed within a reasonable time frame. Efficient use of infrastructure is essential, and analyses must be fully reproducible with publicly available workflows to ensure transparency. Here, we introduce the Metagenomics-Toolkit, a scalable, data-agnostic workflow that automates the analysis of short and long metagenomic reads from Illumina and Oxford Nanopore Technology devices, respectively. The Metagenomics-Toolkit provides standard features such as quality control, assembly, binning, and annotation, along with unique capabilities including plasmid identification, recovery of unassembled microbial community members, and discovery of microbial interdependencies through dereplication, co-occurrence, and genome-scale metabolic modeling. Additionally, the Metagenomics-Toolkit includes a machine learning-optimized assembly step that adjusts peak RAM usage to match actual requirements, reducing the need for high-memory hardware. It can be executed on user workstations and includes optimizations for efficient cloud-based cluster execution. We compare the Metagenomics-Toolkit with five widely used metagenomics workflows and demonstrate its capabilities on 757 sewage metagenome datasets to investigate a possible sewage core microbiome. The Metagenomics-Toolkit is open source and available at https://github.com/metagenomics/metagenomics-tk.</p", "keywords": ["Pipelines and Workflows for Biological Data Analysis"]}, "links": [{"href": "https://doi.org/10.1101/2024.10.22.619569"}, {"rel": "related", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/NAR%20Genomics%20and%20Bioinformatics", "name": "related record", "description": "related record", "type": "application/json"}, {"rel": "self", "type": "application/geo+json", "title": "10.1101/2024.10.22.619569", "name": "item", "description": "10.1101/2024.10.22.619569", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/10.1101/2024.10.22.619569"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2024-10-25T00:00:00Z"}}, {"id": "10.1105/tpc.20.00318", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:17:24Z", "type": "Journal Article", "created": "2020-10-10", "title": "ARADEEPOPSIS, an Automated Workflow for Top-View Plant Phenomics using Semantic Segmentation of Leaf States", "description": "Linking plant phenotype to genotype is a common goal to both plant breeders and geneticists. However, collecting phenotypic data for large numbers of plants remain a bottleneck. Plant phenotyping is mostly image based and therefore requires rapid and robust extraction of phenotypic measurements from image data. However, because segmentation tools usually rely on color information, they are sensitive to background or plant color deviations. We have developed a versatile, fully open-source pipeline to extract phenotypic measurements from plant images in an unsupervised manner. ARADEEPOPSIS (https://github.com/Gregor-Mendel-Institute/aradeepopsis) uses semantic segmentation of top-view images to classify leaf tissue into three categories: healthy, anthocyanin rich, and senescent. This makes it particularly powerful at quantitative phenotyping of different developmental stages, mutants with aberrant leaf color and/or phenotype, and plants growing in stressful conditions. On a panel of 210 natural Arabidopsis (Arabidopsis thaliana) accessions, we were able to not only accurately segment images of phenotypically diverse genotypes but also to identify known loci related to anthocyanin production and early necrosis in genome-wide association analyses. Our pipeline accurately processed images of diverse origin, quality, and background composition, and of a distantly related Brassicaceae. ARADEEPOPSIS is deployable on most operating systems and high-performance computing environments and can be used independently of bioinformatics expertise and resources.", "keywords": ["0301 basic medicine", "0303 health sciences", "Genotype", "Large-Scale Biology Articles", "Arabidopsis", "Computational Biology", "Semantics", "Workflow", "Plant Leaves", "03 medical and health sciences", "Phenotype", "Image Processing", " Computer-Assisted", "Phenomics", "Software", "Genome-Wide Association Study"]}, "links": [{"href": "https://doi.org/10.1105/tpc.20.00318"}, {"rel": "related", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/The%20Plant%20Cell", "name": "related record", "description": "related record", "type": "application/json"}, {"rel": "self", "type": "application/geo+json", "title": "10.1105/tpc.20.00318", "name": "item", "description": "10.1105/tpc.20.00318", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/10.1105/tpc.20.00318"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2020-10-09T00:00:00Z"}}, {"id": "10.1128/msystems.00859-24", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:17:59Z", "type": "Journal Article", "created": "2024-09-10", "title": "A novel barcoded nanopore sequencing workflow of high-quality, full-length bacterial 16S amplicons for taxonomic annotation of bacterial isolates and complex microbial communities", "description": "ABSTRACT                                     <p>               Due to recent improvements, Nanopore sequencing has become a promising method for experiments relying on amplicon sequencing. We describe a flexible workflow to generate and annotate high-quality, full-length 16S rDNA amplicons. We evaluated it for two applications, namely, (i) identification of bacterial isolates and (ii) species-level profiling of microbial communities. We assessed the identification of single bacterial isolates by sequencing, using a set of barcoded full-length 16S rRNA gene primer pairs (pair A), on 47 isolates encompassing multiple genera and compared those results with matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS)-based identification. Species-level community profiling was tested with two sets of barcoded full-length 16S primer pairs (A and B) and compared to the results obtained with shotgun Illumina sequencing using 27 stool samples. We developed a Nextflow pipeline to retain high-quality reads and taxonomically annotate them. We found high agreement between our workflow and MALDI-TOF data for isolate identification (positive predictive value = 0.90, Cram\uffc3\uffa9r\uffe2\uff80\uff99s               V               = 0.857, and Theil\uffe2\uff80\uff99s               U               = 0.316). For species-level community profiling, we found strong correlations (               r                                s                              &gt; 0.6) of alpha diversity indices between the two primer sets and Illumina sequencing. At the community level, we found significant but small differences when comparing sequencing techniques. Finally, we found a moderate to strong correlation when comparing the relative abundances of individual species (average               r                                s                              = 0.6 and 0.533 for primers A and B). Despite identified shortcomings, the proposed workflow enabled accurate identification of single bacterial isolates and prominent features in microbial communities, making it a worthwhile alternative to MALDI-TOF MS and Illumina sequencing.             </p>                            IMPORTANCE               <p>A quick, robust, simple, and cost-effective method to identify bacterial isolates and communities in each sample is indispensable in the fields of microbiology and infection biology. Recent technological advances in Oxford Nanopore Technologies sequencing make this technique an attractive option considering the adaptability, portability, and cost-effectiveness of the platform, even with small sequencing batches. Here, we validated a flexible workflow to identify bacterial isolates and characterize bacterial communities using the Oxford Nanopore Technologies sequencing platform combined with the most recent v14 chemistry kits. For bacterial isolates, we compared our nanopore-based approach to matrix-assisted laser desorption ionization-time of flight mass spectrometry-based identification. For species-level profiling of complex bacterial communities, we compared our nanopore-based approach to Illumina shotgun sequencing. For reproducibility purposes, we wrapped the code used to process the sequencing data into a ready-to-use and self-contained Nextflow pipeline.</p>", "keywords": ["DNA", " Bacterial", "1303 Biochemistry", "gut microbiome", "610 Medicine & health", "Microbiology", "Workflow", "1311 Genetics", "RNA", " Ribosomal", " 16S", "1312 Molecular Biology", "1706 Computer Science Applications", "DNA Barcoding", " Taxonomic", "Humans", "DNA sequencing", "Bacteria", "10179 Institute of Medical Microbiology", "Microbiota", "2404 Microbiology", "1314 Physiology", "bioinformatics", "QR1-502", "Nanopore Sequencing", "1105 Ecology", " Evolution", " Behavior and Systematics", "Spectrometry", " Mass", " Matrix-Assisted Laser Desorption-Ionization", "570 Life sciences; biology", "2611 Modeling and Simulation", "Research Article"]}, "links": [{"href": "https://doi.org/10.1128/msystems.00859-24"}, {"rel": "related", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/mSystems", "name": "related record", "description": "related record", "type": "application/json"}, {"rel": "self", "type": "application/geo+json", "title": "10.1128/msystems.00859-24", "name": "item", "description": "10.1128/msystems.00859-24", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/10.1128/msystems.00859-24"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2024-04-11T00:00:00Z"}}, {"id": "10.1186/s12302-024-00873-1", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:18:08Z", "type": "Journal Article", "created": "2024-03-11", "title": "SWAT\u2009+\u2009input data preparation in a scripted workflow: SWATprepR", "description": "Abstract<p>Input data collection, quality assurance and preparation are central but time_consuming steps in environmental modeling. Errors due to manual processing of model input data can result in an incorrect representation of an environmental system and may consequently lead to implausible model simulations. Correct input data preparation and thorough quality check at an early stage of the model setup procedure are essential to build confidence in model simulation results. Typically, in environmental model applications, many steps in the input data preparation phase have to be repeated with the inflow of new, additional or corrected data. In this study, we selected the widely used SWAT\uffe2\uff80\uff89+\uffe2\uff80\uff89ecohydrological model as an illustrative example to investigate challenges related to input data preparation. To assist in these tasks, we developed an R package named SWATprepR, which provides functions for typical and repeating SWAT\uffe2\uff80\uff89+\uffe2\uff80\uff89model input data preparation tasks. The package supports the preparation of weather input files, atmospheric deposition, soil parameters, crop rotations, and observed (control or calibration) data, to name a few, presently with focus on European applications. The SWATprepR functions are integrated in R script workflows and can help SWAT\uffe2\uff80\uff89+\uffe2\uff80\uff89modelers to avoid repetitive tasks, secure reproducibility and transparently document the data processing steps. Application of the package is illustrated with a test case of a SWAT\uffe2\uff80\uff89+\uffe2\uff80\uff89model for a small catchment in central Poland.</p", "keywords": ["Environmental sciences", "SWAT\u2009+\u2009model", "Environmental law", "R package", "0208 environmental biotechnology", "0207 environmental engineering", "GE1-350", "02 engineering and technology", "Input data processing", "K3581-3598", "Reproducibility", "Workflow"]}, "links": [{"href": "https://link.springer.com/content/pdf/10.1186/s12302-024-00873-1.pdf"}, {"href": "https://doi.org/10.1186/s12302-024-00873-1"}, {"rel": "related", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/Environmental%20Sciences%20Europe", "name": "related record", "description": "related record", "type": "application/json"}, {"rel": "self", "type": "application/geo+json", "title": "10.1186/s12302-024-00873-1", "name": "item", "description": "10.1186/s12302-024-00873-1", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/10.1186/s12302-024-00873-1"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2024-03-11T00:00:00Z"}}, {"id": "10.18419/opus-2935", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:18:36Z", "type": "Report", "title": "Datenmanagementpatterns in multi-skalaren Simulationsworkflows", "description": "In den vergangenen Jahren haben sich im unternehmerischen Umfeld Workflows zur Beschreibung und Ausf\u00fchrung von (Gesch\u00e4fts-)Prozessen durchgesetzt. Seit kurzem wird diese Technologie auch in der Wissenschaft eingesetzt. Z.B. werden Simulationsabl\u00e4ufe als Workflows modelliert. Charakteristisch f\u00fcr solche Simulationen bzw. Simulationsabl\u00e4ufe sind komplexe mathematische Berechnungen sowie verschiedene Aufgaben im Bereich der Datenverwaltung und Datenbereitstellung. Oftmals m\u00fcssen gro\u00dfe Datenmengen, die in propriet\u00e4ren Formaten vorliegen, aus verschiedenen Quellen verarbeitet werden. Damit diese Daten durch einen Simulationsworkflow und den von ihm eingebundenen Programmen und Diensten verarbeitet werden k\u00f6nnen, m\u00fcssen sie in passende Eingabeformate transformiert werden. Gerade bei umfangreichen Simulationen, die eine Vielzahl an Datenquellen ben\u00f6tigen, f\u00fchrt dies aufgrund der enormen Komplexit\u00e4t zu Problemen. Um diese Probleme zu l\u00f6sen, wurde das SIMPL-Rahmenwerk (SimTech - Information Management, Processes and Languages) entwickelt. Das SIMPL-Rahmenwerk ist in ein Scientifc Workflow Management System eingebettet und schafft eine Abstraktionsebene f\u00fcr die Defnition des Datenmanagements. SIMPL bietet einheitliche Zugriffsmethoden, um, aus einem Simulationsworkflow heraus, auf beliebige Datenquellen zuzugreifen. Ein weiterer Bestandteil des SIMPL-Rahmenwerks sind Datenmanagementpatterns. Dabei handelt es sich um vorgefertigte Datenmanagement-Operationen, die nur noch parametrisiert werden m\u00fcssen. Auf diese Weise wird eine neue Abstraktionsebene geschaffen. In einer vorherigen Arbeit wurden bereits erste Datenmanagementpatterns erarbeitet. So k\u00f6nnen z.B. Daten zwischen zwei Datenressourcen ausgetauscht werden. Des Weiteren wurde ein Konzept erarbeitet, um Datenmanagementpatterns auf ausf\u00fchrbare Workflow-Fragmente abzubilden. Dieses Konzept nutzt Transformationsregeln sowie gespeicherte Metadaten \u00fcber beteiligte Ressourcen als Basis. Im Rahmen dieser Diplomarbeit wird das bereits entwickelte Konzept erweitert und wenn n\u00f6tig angepasst, um auf multi-skalare Simulationen angewendet werden zu k\u00f6nnen. Dar\u00fcber hinaus wird die prototypische Umsetzung des SIMPL-Rahmenwerks um Datenmanagementpatterns erweitert.", "keywords": ["000", "Heterogeneous Databases (CR H.2.5)", "Datenmanagementpatterns", "Software Engineering Software Architectures (CR D.2.11)", "wissenschaftliche Workflows", "Office Automation (CR H.4.1)", "Datenmanagement", "Simulationsworkflows", "Simulation Support Systems (CR I.6.7)", "Datenbereitstellung", "004"], "contacts": [{"organization": "Pietranek, Henrik Andreas", "roles": ["creator"]}]}, "links": [{"href": "https://doi.org/10.18419/opus-2935"}, {"rel": "self", "type": "application/geo+json", "title": "10.18419/opus-2935", "name": "item", "description": "10.18419/opus-2935", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/10.18419/opus-2935"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2012-01-01T00:00:00Z"}}, {"id": "10.5194/bg-19-3505-2022", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:19:57Z", "type": "Journal Article", "created": "2022-07-28", "title": "Reviews and syntheses: The promise of big diverse soil data, moving current practices towards future potential", "description": "<?xml version='1.0' encoding='UTF-8'?><article><p>Abstract. In the age of big data, soil data are more available and richer than ever, but \u2013 outside of a few large soil survey resources \u2013 they remain largely unusable for informing soil management and understanding Earth system processes beyond the original study. Data science has promised a fully reusable research pipeline where data from past studies are used to contextualize new findings and reanalyzed for new insight. Yet synthesis projects encounter challenges at all steps of the data reuse pipeline, including unavailable data, labor-intensive transcription of datasets, incomplete metadata, and a lack of communication between collaborators. Here, using insights from a diversity of soil, data, and climate scientists, we summarize current practices in soil data synthesis across all stages of database creation: availability, input, harmonization, curation, and publication. We then suggest new soil-focused semantic tools to improve existing data pipelines, such as ontologies, vocabulary lists, and community practices. Our goal is to provide the soil data community with an overview of current practices in soil data and where we need to go to fully leverage big data to solve soil problems in the next century.                     </p></article>", "keywords": ["FOS: Computer and information sciences", "0301 basic medicine", "Data Sharing", "Information Systems and Management", "literature review", "1904 Earth-Surface Processes", "Social Sciences", "data set", "01 natural sciences", "Decision Sciences", "Data science", "Life", "QH501-531", "910 Geography & travel", "soil analysis", "database", "QH540-549.5", "2. Zero hunger", "QE1-996.5", "000", "Ecology", "communication", "Physics", "Earth", "Geology", "[SDU.ENVI] Sciences of the Universe [physics]/Continental interfaces", " environment", "World Wide Web", "10122 Institute of Geography", "soil survey", "Physical Sciences", "Data Reuse", "environment", "Information Systems", "Evolution", "future prospect", "Data management", "Data Sharing and Stewardship in Science", "Database", "Big data", "03 medical and health sciences", "Behavior and Systematics", "Data mining", "0105 earth and related environmental sciences", "[SDU.OCEAN]Sciences of the Universe [physics]/Ocean", "Management and Reproducibility of Scientific Workflows", "Metadata", "Data curation", "Atmosphere", "[SDU.OCEAN] Sciences of the Universe [physics]/Ocean", " Atmosphere", "Acoustics", "15. Life on land", "Computer science", "1105 Ecology", " Evolution", " Behavior and Systematics", "Surface Processes", "Harmonization", "FOS: Biological sciences", "Computer Science", "Environmental Science", "[SDU.ENVI]Sciences of the Universe [physics]/Continental interfaces", "soil management", "Research Data", "Environmental DNA in Biodiversity Monitoring"]}, "links": [{"href": "https://doi.org/10.5194/bg-19-3505-2022"}, {"rel": "related", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/Biogeosciences", "name": "related record", "description": "related record", "type": "application/json"}, {"rel": "self", "type": "application/geo+json", "title": "10.5194/bg-19-3505-2022", "name": "item", "description": "10.5194/bg-19-3505-2022", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/10.5194/bg-19-3505-2022"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2022-07-28T00:00:00Z"}}, {"id": "10.5281/zenodo.13986640", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:20:29Z", "type": "Dataset", "title": "Map of soil organic carbon loss of mineral soils in Estonia", "description": "Open AccessThe internal EJP SOIL project SERENA contributed to the evaluation of soil multifunctionality aiming at providing assessment tools for land planning and soil policies at different scales. By co-working with relevant stakeholders, the project provided co-developed indicators and associated cookbooks to assess and map them, to report both on soil degradation, soil-based ecosystem services and their bundles, under actual conditions and for climate and land-use changes, at the regional, national, and European scales.  The map was generated to evaluate soil organic carbon (SOC) loss in Estonian agricultural soils. It is directly related to SERENA project WP3, T3.2, D3.3 with the aim of applying cookbooks to assess soil threats or ecosystem services. This map is the outcome of applying a cookbook developed by ISRIC (Genova, G., Poggio, L., Kempen, B., & Colman, B. DSM Workflow Seedling. ISRIC - World Soil Information. https://doi.org/10.17027/ISRIC-FSX2-2691).  The generated map of SOC loss expressed as absolute sequestration rate (t C ha-1 a-1) between 2015 and 2021 is in GEOTIFF format at the resolution of 100m. The input data for the cookbook was from the PANDA database, which contains regular soil monitoring and voluntary soil sampling data by farmers in Estonia. To achieve the aim for accounting SOC loss in agricultural soils temporal pairs were selected resulting in 1037 paired points where the interval between second sampling was more than 5 years. SOC stocks were calculated for the depth of 20 cm using the equation by Adams (1973) to calculate soil bulk density. The calculated SOC stock for time0 and time2 (> 5 years resampled locations) were used as input points for digital soil mapping, that is the ISRIC cookbook.", "keywords": ["Estonia", "Task3.2", "ESTONIA", "EJPSOIL", "WP3", "SERENA project", "D3.3", "SOC loss", "Grant n 862695", "DSM Workflow Seedling"], "contacts": [{"organization": "Putku, Elsa", "roles": ["creator"]}]}, "links": [{"href": "https://doi.org/10.5281/zenodo.13986640"}, {"rel": "self", "type": "application/geo+json", "title": "10.5281/zenodo.13986640", "name": "item", "description": "10.5281/zenodo.13986640", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/10.5281/zenodo.13986640"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2024-10-24T00:00:00Z"}}, {"id": "10.5281/zenodo.14989604", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:20:45Z", "type": "Software", "title": "Metagenomics-Toolkit: The Flexible and Efficient Cloud-Based Metagenomics Workflow featuring Machine Learning-Enabled Resource Allocation", "description": "The Metagenomics-Toolkit is a scalable, data agnostic workflow that automates the analysis of short and long metagenomic reads obtained from Illumina or Oxford Nanopore Technology devices, respectively. The Toolkit offers not only standard features expected in a metagenome workflow, such as quality control, assembly, binning, and annotation, but also distinctive features, such as plasmid identification based on various tools, the recovery of unassembled microbial community members, and the discovery of microbial interdependencies through a combination of dereplication, co-occurrence, and genome-scale metabolic modeling. Furthermore, the Metagenomics-Toolkit includes a machine learning-optimized assembly step that tailors the peak RAM value requested by a metagenome assembler to match actual requirements, thereby minimizing the dependency on dedicated high-memory hardware.  Quickstart and documentation can be found at https://github.com/metagenomics/metagenomics-tk", "keywords": ["Metagenomics/methods", "Cloud Computing", "Workflow"], "contacts": [{"organization": "Belmann, Peter, Osterholz, Benedikt, Kleinb\u00f6lting, Nils, P\u00fchler, Alfred, Schl\u00fcter, Andreas, Sczyrba, Alexander,", "roles": ["creator"]}]}, "links": [{"href": "https://doi.org/10.5281/zenodo.14989604"}, {"rel": "self", "type": "application/geo+json", "title": "10.5281/zenodo.14989604", "name": "item", "description": "10.5281/zenodo.14989604", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/10.5281/zenodo.14989604"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2025-03-07T00:00:00Z"}}, {"id": "20.500.11850/562259", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:22:36Z", "type": "Journal Article", "created": "2022-07-28", "title": "Reviews and syntheses: The promise of big diverse soil data, moving current practices towards future potential", "description": "<?xml version='1.0' encoding='UTF-8'?><article><p>Abstract. In the age of big data, soil data are more available and richer than ever, but \u2013 outside of a few large soil survey resources \u2013 they remain largely unusable for informing soil management and understanding Earth system processes beyond the original study. Data science has promised a fully reusable research pipeline where data from past studies are used to contextualize new findings and reanalyzed for new insight. Yet synthesis projects encounter challenges at all steps of the data reuse pipeline, including unavailable data, labor-intensive transcription of datasets, incomplete metadata, and a lack of communication between collaborators. Here, using insights from a diversity of soil, data, and climate scientists, we summarize current practices in soil data synthesis across all stages of database creation: availability, input, harmonization, curation, and publication. We then suggest new soil-focused semantic tools to improve existing data pipelines, such as ontologies, vocabulary lists, and community practices. Our goal is to provide the soil data community with an overview of current practices in soil data and where we need to go to fully leverage big data to solve soil problems in the next century.</p></article>", "keywords": ["FOS: Computer and information sciences", "0301 basic medicine", "Data Sharing", "Information Systems and Management", "literature review", "1904 Earth-Surface Processes", "Social Sciences", "data set", "01 natural sciences", "Decision Sciences", "Data science", "Life", "QH501-531", "910 Geography & travel", "soil analysis", "database", "QH540-549.5", "2. Zero hunger", "QE1-996.5", "000", "Ecology", "communication", "Physics", "Earth", "Geology", "[SDU.ENVI] Sciences of the Universe [physics]/Continental interfaces", " environment", "World Wide Web", "10122 Institute of Geography", "soil survey", "Physical Sciences", "Data Reuse", "environment", "Information Systems", "Evolution", "future prospect", "Data management", "Data Sharing and Stewardship in Science", "Database", "Big data", "03 medical and health sciences", "Behavior and Systematics", "Data mining", "0105 earth and related environmental sciences", "[SDU.OCEAN]Sciences of the Universe [physics]/Ocean", "Management and Reproducibility of Scientific Workflows", "Metadata", "Data curation", "Atmosphere", "[SDU.OCEAN] Sciences of the Universe [physics]/Ocean", " Atmosphere", "Acoustics", "15. Life on land", "Computer science", "1105 Ecology", " Evolution", " Behavior and Systematics", "Surface Processes", "Harmonization", "FOS: Biological sciences", "Computer Science", "Environmental Science", "[SDU.ENVI]Sciences of the Universe [physics]/Continental interfaces", "soil management", "Research Data", "Environmental DNA in Biodiversity Monitoring"]}, "links": [{"href": "https://doi.org/20.500.11850/562259"}, {"rel": "related", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/Biogeosciences", "name": "related record", "description": "related record", "type": "application/json"}, {"rel": "self", "type": "application/geo+json", "title": "20.500.11850/562259", "name": "item", "description": "20.500.11850/562259", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/20.500.11850/562259"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2022-07-28T00:00:00Z"}}, {"id": "PMC11494973", "type": "Feature", "geometry": null, "properties": {"updated": "2026-09-21T16:24:39Z", "type": "Journal Article", "created": "2024-09-10", "title": "A novel barcoded nanopore sequencing workflow of high-quality, full-length bacterial 16S amplicons for taxonomic annotation of bacterial isolates and complex microbial communities", "description": "ABSTRACT                                                             <p>                       Due to recent improvements, Nanopore sequencing has become a promising method for experiments relying on amplicon sequencing. We describe a flexible workflow to generate and annotate high-quality, full-length 16S rDNA amplicons. We evaluated it for two applications, namely, (i) identification of bacterial isolates and (ii) species-level profiling of microbial communities. We assessed the identification of single bacterial isolates by sequencing, using a set of barcoded full-length 16S rRNA gene primer pairs (pair A), on 47 isolates encompassing multiple genera and compared those results with matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS)-based identification. Species-level community profiling was tested with two sets of barcoded full-length 16S primer pairs (A and B) and compared to the results obtained with shotgun Illumina sequencing using 27 stool samples. We developed a Nextflow pipeline to retain high-quality reads and taxonomically annotate them. We found high agreement between our workflow and MALDI-TOF data for isolate identification (positive predictive value = 0.90, Cram\uffc3\uffa9r\uffe2\uff80\uff99s                       V                       = 0.857, and Theil\uffe2\uff80\uff99s                       U                       = 0.316). For species-level community profiling, we found strong correlations (                       r                                                s                                              &gt; 0.6) of alpha diversity indices between the two primer sets and Illumina sequencing. At the community level, we found significant but small differences when comparing sequencing techniques. Finally, we found a moderate to strong correlation when comparing the relative abundances of individual species (average                       r                                                s                                              = 0.6 and 0.533 for primers A and B). Despite identified shortcomings, the proposed workflow enabled accurate identification of single bacterial isolates and prominent features in microbial communities, making it a worthwhile alternative to MALDI-TOF MS and Illumina sequencing.                     </p>                                            IMPORTANCE                       <p>A quick, robust, simple, and cost-effective method to identify bacterial isolates and communities in each sample is indispensable in the fields of microbiology and infection biology. Recent technological advances in Oxford Nanopore Technologies sequencing make this technique an attractive option considering the adaptability, portability, and cost-effectiveness of the platform, even with small sequencing batches. Here, we validated a flexible workflow to identify bacterial isolates and characterize bacterial communities using the Oxford Nanopore Technologies sequencing platform combined with the most recent v14 chemistry kits. For bacterial isolates, we compared our nanopore-based approach to matrix-assisted laser desorption ionization-time of flight mass spectrometry-based identification. For species-level profiling of complex bacterial communities, we compared our nanopore-based approach to Illumina shotgun sequencing. For reproducibility purposes, we wrapped the code used to process the sequencing data into a ready-to-use and self-contained Nextflow pipeline.</p>", "keywords": ["DNA", " Bacterial", "1303 Biochemistry", "gut microbiome", "610 Medicine & health", "Microbiology", "Workflow", "1311 Genetics", "RNA", " Ribosomal", " 16S", "1312 Molecular Biology", "1706 Computer Science Applications", "DNA Barcoding", " Taxonomic", "Humans", "DNA sequencing", "Bacteria", "10179 Institute of Medical Microbiology", "Microbiota", "2404 Microbiology", "1314 Physiology", "bioinformatics", "QR1-502", "Nanopore Sequencing", "1105 Ecology", " Evolution", " Behavior and Systematics", "Spectrometry", " Mass", " Matrix-Assisted Laser Desorption-Ionization", "570 Life sciences; biology", "2611 Modeling and Simulation", "Research Article"]}, "links": [{"href": "https://doi.org/PMC11494973"}, {"rel": "related", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/mSystems", "name": "related record", "description": "related record", "type": "application/json"}, {"rel": "self", "type": "application/geo+json", "title": "PMC11494973", "name": "item", "description": "PMC11494973", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items/PMC11494973"}, {"rel": "collection", "type": "application/json", "title": "Collection", "name": "collection", "description": "Collection", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main"}], "time": {"date": "2024-04-11T00:00:00Z"}}], "links": [{"rel": "self", "type": "application/geo+json", "title": "This document as GeoJSON", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items?keywords=Workflow&f=json", "hreflang": "en-US"}, {"rel": "alternate", "type": "text/html", "title": "This document as HTML", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items?keywords=Workflow&f=html", "hreflang": "en-US"}, {"rel": "collection", "type": "application/json", "title": "Collection URL", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main", "hreflang": "en-US"}, {"type": "application/geo+json", "rel": "first", "title": "items (first)", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items?keywords=Workflow&", "hreflang": "en-US"}, {"rel": "last", "type": "application/geo+json", "title": "items (last)", "href": "https://repository.soilwise-he.eu/cat/collections/metadata:main/items?keywords=Workflow&offset=11", "hreflang": "en-US"}], "numberMatched": 11, "numberReturned": 11, "distributedFeatures": [], "timeStamp": "2026-09-22T13:57:44.742436Z"}