Breeding for Stress Tolerance and Quality
In current agriculture, abiotic and biotic stress are the main reasons that yield potential and quality aspects are difficult to realize for many crops. Resistance breeding focuses on the use of genetic resources for improving plant defence against stress factors. Breeding for biotic stress resistance addresses with defence mechanisms and strategies that protect host plants
Breeding for Stress Tolerance and Quality
In current agriculture, abiotic and biotic stress are the main reasons that yield potential and quality aspects are difficult to realize for many crops. Resistance breeding focuses on the use of genetic resources for improving plant defence against stress factors. Breeding for biotic stress resistance addresses with defence mechanisms and strategies that protect host plants against pests and pathogens, inheritance of resistance genes, and durable effectiveness of resistance genes. Abiotic stress is caused by environmental factors (a.o. temperature, water, nutrients, minerals). Breeding for abiotic stress resistance and tolerance to such factors addresses concepts such as adaptability and stability of crop plants, mechanisms of stress tolerance and phenotyping for selection, genotype by environment interaction and selection in multi-environment trials. Differences between plant genotypes and assessment of the importance of these effects can be quantified in experiments with appropriate experimental designs and using statistical analysis of the collected data.
Quality breeding is mainly directed at improving plant compounds like carbohydrates, proteins, vegetable fats and oils, fibres and secondary metabolites, that are all synthesized in metabolic pathways. Breeding objectives include improved product quality (e.g. taste, shelf life), enhanced production of flavours, fragrances and health-supporting components, absence of allergens and other undesirable compounds and improvement of processing characteristics of plant raw materials.
Program/Minor
BSc Minor Plant Breeding; Master Biobased Sciences; Master Plant Biotechnology; Master Plant SciencesLearning Outcomes
Explain the major characteristics of various resistance, tolerance and quality traits; Define appropriate selection strategies for specific target traits; Apply relevant analytic and statistical screening techniques for trait evaluation; Use this knowledge to develop breeding strategies for improved resistance, tolerance and quality
Prerequisites
Assumed: MAT20306 Advanced Statistics; PBR21803 Pre-breeding; PBR22303 Plant Breeding Mandatory: ZSS06100 Laboratory Safety; ZSS06200 Fieldwork Safety
Details
Institution: WUR
Period: Period 5
Level: Advanced
EC: 6 ec
BSc Minor Climate-Resilient Crops: Interdisciplinary Approaches
Plant breeding has been enormously successful in increasing the yield, variety, and quality of crops we consume on a daily basis. However, it is a major challenge to meet the growing global demand for affordable agricultural products while adapting to climate change (leading to heat waves, droughts, floods, diseases, pests and poor soil) and increasing
BSc Minor Climate-Resilient Crops: Interdisciplinary Approaches
Plant breeding has been enormously successful in increasing the yield, variety, and quality of crops we consume on a daily basis. However, it is a major challenge to meet the growing global demand for affordable agricultural products while adapting to climate change (leading to heat waves, droughts, floods, diseases, pests and poor soil) and increasing sustainability-driven constraints on agriculture. This challenge is further exacerbated by growing populations, dietary changes, and declining farmlands. A key element in addressing these challenges is the development of climate-resilient crops that, thanks to new genomic makeups and cultivation methods, thrive even under more variable, more unpredictable, and more often extreme abiotic and biotic stresses. Resilience is, however, a highly complex trait with multiple genes and processes interacting simultaneously and/or over time, involving many trade-offs. Even the most advanced current plant breeding techniques lack the ability to efficiently select for such traits. Moreover, actors in food systems around the world may have diverging views on what resilience traits are most important for their specific context, while the development and uptake of climate-resilient crops may also face various non-technological challenges related to organizational routines, government regulations, and market structures.
Thus, the CropXR institute was funded to disentangle these complex traits, develop a generation of more resilient crops, and train scientists in state-of-the-art methodologies which provide insights into the most important processes underlying plant performance under stress and subsequent societal deployment of these innovations. Such state-of-the-art approaches are highly multidisciplinary, combining recent developments in plant science, social science, and data science and modelling, and thus requires trained experts with affinity for combining all these aspects. In this minor, students will learn fundamental concepts of all three areas, the fields’ latest developments, and major challenges related to their integration. With this knowledge, we expect them to be able to tackle stakeholder needs and problems by providing sustainable and innovative solutions to develop climate-change resilient and future proof crops.
Program/Minor
missingLearning Outcomes
Discuss the most relevant interdisciplinary aspects of climate resilient crop development; Explain and discuss the relevance of developments and challenges at the intersection of plant sciences, data sciences and social sciences; Advise stakeholders on a plant resilience-related problem, integrating aspects of plant sciences, data sciences and social sciences To give structure to the minor we have created three clusters of courses. Students are encouraged to choose courses from the clusters they are least familiar with as this will provide a basic understanding on how different fields approach crop resilience-related problems and will allow them to communicate with students from other backgrounds more effectively. NB: Clusters are only a suggestion to give students structure. If they want, they can ignore or combine clusters to tailor the minor towards their specific interests. Cluster Quantitative Resilient Crop Development: This cluster focusses on teaching students about quantitative, modelling and/or computational techniques that are used for resilient crop development. Cluster Biology of Resilient Crops: This cluster focusses on the (molecular) Plant biology that underlies resilient crops. Cluster Societal Impact of Resilient Crops: This cluster teaches students about different societal factors that play a role in how innovations can contribute to society, the different trade-offs that are involved, and about how to innovate responsibly.
Prerequisites
Basic knowledge on biology (high-school level) is assumed. Anyone with an interest in learning more about the challenge to develop climate change resilient crops and the various disciplines that support this. Basic knowledge on biology (high-school level) is assumed. Students can come from a social science, plant science, or data science background, e.g. any student from these programs: BSc (WUR): Biology, Biotechnology, Plant Sciences, Molecular Life Sciences, Communication and Life Sciences, Data Science for Global Challenges, International Development Studies etc. BSc (other universities): Life Sciences, Data Science, Computer Science, Global Sustainability Science, Artificial Intelligence, Molecular and Biophysical Life Sciences, Molecular Science and Technology, Computational Social Science, Future Planet Studies, Natuurwetenschap en Innovatiemanagement etc. HBO: Biology, Agriculture, Horticulture, Food Systems Innovation, Life Sciences etc.
Details
Institution: WUR
Period: Period 1 - 3
Level: missing
EC: 24 ec
Climate Smart Agriculture
Agriculture contributes significantly to global warming through large scale greenhouse gas emissions. At the same time many agriculture systems are vulnerable to climate change and without adaptation global food production could significantly reduce affecting food security. In response to these challenges the concept of climate smart agriculture has been developed. Climate-smart agriculture (CSA) aims to
Climate Smart Agriculture
Agriculture contributes significantly to global warming through large scale greenhouse gas emissions. At the same time many agriculture systems are vulnerable to climate change and without adaptation global food production could significantly reduce affecting food security. In response to these challenges the concept of climate smart agriculture has been developed. Climate-smart agriculture (CSA) aims to address the interlinked challenges of food security and climate change through an integrated approach. For example an better integration of land and water management is necessary to adapt to future climate change while at the same time reducing emissions. Climate smart agriculture combines three different objectives (1) sustainable increase of agricultural production (2) Adaptation of agricultural and food security systems to climate change; and (3) reducing greenhouse gas emissions from agriculture. To improve both climate change mitigation and adaptation of the agricultural sector these issues need to addressed at different spatial scales (from farm to landscape & local to global) and at time scales (from season to decades).
During the course the students will learn about the main principles of climate smart agriculture. The student will learn how agricultural systems including both plant and animal systems contribute to climate change through the emissions of CO2, N2O and CH4. The course will address the main processes causing these emissions and which mitigation measures can be used to reduce greenhouse gas emissions. Similarly, the course will address the main impacts of climate change on on agricultural systems and food security, how these impacts can be assessed and which adaptation measures can be used the reduce the vulnerability of agricultural systems and improve food security. The course will address how climate change affects water and land use management including issues such as future water availability and plant growth and which land and water use changes are needed to adapt to future climate. The key of climate smart agriculture is integrating adaptation and mitigation into existing farming systems and land and water use practices. Especially in developing countries ideally this would combined with higher agricultural production and improved food security. The course will address cases from both developed (intesive, high inputs, large scale, market-oriented) and the developing (extensive, minimal inputs, small scale, subsistence) world focusing on for example dryland systems in Africa, rice based farming systems in Asia and mixed farming systems in Europe.
During the course the students will develop a climate smart agricultural farming systems for their case. The student will assess the greenhouse gas emission and the climate change impacts using simple farming systems models. Based on these analyses the student will design different adaptation and mitigation measures which need to be integrated into a climate smart strategy for their case.
Program/Minor
Master Climate Studies; Master International Land and Water Management; Master Plant SciencesLearning Outcomes
Apply the main principles of climate smart agriculture; Analyse the impacts of climate variability and climate change on agricultural systems; Describe the essential processes that are important in crop-climate interactions; Perform simple analyses on the CO2, CH4, & N2O emission potential of agricultural systems; Develop and critically assess adaptation and mitigation measures related to agricultural systems; Integrate adaptation and mitigation measures into a climate SMART agricultural system
Details
Institution: WUR
Period: Period 3
Level: Advanced
EC: 6 ec
Computational Biology
This course focuses on using computational modelling to explore biological systems and test specific hypotheses. Students learn to construct exact models and analyse their behaviour to gain insight into the original biological system. The course draws on a broad range of biological questions across evolutionary, developmental, ecological, and molecular biology. Topics include evolutionary dynamics such
Computational Biology
This course focuses on using computational modelling to explore biological systems and test specific hypotheses. Students learn to construct exact models and analyse their behaviour to gain insight into the original biological system. The course draws on a broad range of biological questions across evolutionary, developmental, ecological, and molecular biology.
Topics include evolutionary dynamics such as genome evolution, robustness, and host-pathogen co-evolution; developmental dynamics such as pattern formation, cell differentiation, morphogenesis, and the evolution of development (EVO-DEVO); network dynamics including gene regulation, metabolic networks, and RNA interference; and behaviour, including self-structuring through local interactions and the relationship between learning and evolution. Both plant and animal models are used.
A central theme throughout is spatial pattern formation and emergent properties, which are introduced as a general theoretical module. Model formalisms taught include non-linear differential and difference equations (ODE, maps), partial differential equations (PDE), cellular automata, event-based models, individually oriented models, evolutionary computation, and hybrid models. Analysis methods include bifurcation analysis, sensitivity analysis, and pattern analysis techniques.
Program/Minor
Biology (BIOL) (B); Mathematics and Applications (WSKT); Minor Complex Systems (MINOR-CS-WIS); Molecular and Biophysical Life Sciences (MBLS)Learning Outcomes
Understand how computational models can be used to investigate biological systems; Formulate and analyse exact models based on specific biological hypotheses; Interpret model results to gain biological insights; Recognize and apply different modelling formalisms including ODE, PDE, cellular automata, event-based and hybrid models; Use analysis techniques such as bifurcation analysis, sensitivity analysis, and pattern analysis; Describe biological processes in evolution, development, networks, and behaviour using computational theory; Critically evaluate modelling approaches and assumptions; Present findings clearly and work effectively in a collaborative research setting.
Prerequisites
Mandatory prerequisite knowledge of Biological models and statistics (former Quantitative Biology) and Genomics. It is strongly recommended to do course Biological Modeling (year 2) beforehand. Students with similar knowledge can be accepted too. For students outside of the UU Biology programme, please contact the coordinator to discuss your pre-knowledge level.
Details
Institution: UU
Period: Period 3
Level: Advanced
EC: 7.5 ec
Control Engineering
Besides a correct design or layout, good control systems are essential to guarantee that production systems operate and produce according the desired specifications. This course gives an introduction to classical control engineering approaches and discusses the standard methods and tools that are usually applied. The methods discussed in the course have a very wide application
Control Engineering
Besides a correct design or layout, good control systems are essential to guarantee that production systems operate and produce according the desired specifications. This course gives an introduction to classical control engineering approaches and discusses the standard methods and tools that are usually applied. The methods discussed in the course have a very wide application area. Examples are greenhouse climate, bioreactors, food production, robotics, environmental systems etc. This makes that the course fits in the curricula of several studies.
The course starts with a refresher on dynamic models of systems represented by differential equations. These differential equations will be solved by transformation to the Laplace domain. The system representation in the Laplace domain by transfer functions offers several new possibilities to interpret and to analyze the characteristics of systems and to design controllers.
Classical control is discussed and analyzed for the PID controller family. Controller tuning, stability and performance are central items to qualify the controllers, and methods to find these qualifications are introduced (response times, pole placement, root-locus and frequency response). At the end of the course the use of control systems is extended from single-input single-output systems to the more complex multiple-input multiple-output systems.
During the course theory will be explained by examples from practice, exercising problems, working on a design case and a computer practical on controller tuning as it would be done in a practical situation. The course provides an important initial step towards more advanced control that can be applied in Mechanical Engineering (combining hardware and software).
Program/Minor
Bachelor Biosystems EngineeringLearning Outcomes
Solve differential equations by using Laplace transformations; Translate differential equations into transfer functions; Derive stability and response characteristics from transfer functions; Design, analyse, and tune PID controllers by using step response and root-locus methods; Design, analyse, and tune PID controllers by using frequency response method, Bode, and Nyquist; Improve the performance of controllers; Analyse a process and design a controller configuration; Use special types of controllers as feedforward, and cascade controllers; Propose the controller structure for multiple-input multiple-output systems; Apply these concepts during practical exercises and a design case
Prerequisites
Assumed knowledge on: Modelling Dynamic Systems; Mathematics 2; Mathematics 3
Details
Institution: WUR
Period: Period 4
Level: Intermediate
EC: 6 ec
Controlled Environment Agriculture
Life-long learning:6 courses on image analysis, greenhouse horticulture (summer school)
Controlled Environment Agriculture
Life-long learning:6 courses on image analysis, greenhouse horticulture (summer school)
Details
Institution: WUR education for professionals
EC: LLO ec
Data Analysis & Visualization
Much data is quantitative, and there is a wide range of methods available for the analysis of such data. After a brief introduction to data types and normalisation, a number of visualisation methods will be discussed. Next, methods will be introduced to find groups (clustering), dependencies (regression), significant differences between conditions (hypothesis testing) and to
Data Analysis & Visualization
Much data is quantitative, and there is a wide range of methods available for the analysis of such data. After a brief introduction to data types and normalisation, a number of visualisation methods will be discussed. Next, methods will be introduced to find groups (clustering), dependencies (regression), significant differences between conditions (hypothesis testing) and to predict classes (classification). In addition, ways of assessing the relevance of findings and of interpreting results will be discussed. Students will learn to apply all these methods in practice in R.
Program/Minor
BSc Minor BioinformaticsLearning Outcomes
Thanks for helping us update the CropXR course inventory! Please enter all you updates in this sheet, since any edits made directly in the other sheets will probably go unnoticed. You can check what courses we have currently assembled in the tab "Courses", we also included a list of courses that were part of the inventory 2 years ago but that we couldn't find in your institutions catalogue for this year under "Removed". Please check if all of the CropXR-related courses have been included in the overview through the tables below. Make sure to at least provide a course name and link so that we can fill out all the info. If a course is not available online, do make sure to fill out all of the information in the table below.
Prerequisites
BIF21806 Practical Computing for Biologists or INF22306 Programming in Python
Details
Institution: WUR
Period: Period 2
Level: Inermediate
EC: 6 ec
Data Analysis for Biosystems Engineering
The following topics will be addressed in the course: linear regression and multiple linear regression, including model formulation, meaning of model parameters, checking model assumptions and prediction; data transformation; experimental design, including completely randomized design, block design and factorial design, and calculating the required sample size to obtain a certain precision; analysis of variance and
Data Analysis for Biosystems Engineering
The following topics will be addressed in the course: linear regression and multiple linear regression, including model formulation, meaning of model parameters, checking model assumptions and prediction; data transformation; experimental design, including completely randomized design, block design and factorial design, and calculating the required sample size to obtain a certain precision; analysis of variance and pair-wise testing; non-parametric tests, including Wilcoxon (Mann-Whitney), Spearman rank correlation, and Kruskal-Wallis; proportion analysis for one population, test for difference between two proportions, and the binomial distribution; contingency tables and the chi-squared tests for goodness of fit, for independence and for homogeneity; multiple linear regression model comparison; experimental design involving factorial design in blocks; selection of variables (quantitative and/or qualitative) to find the optimal linear regression model, including checking assumptions; repeated measurements; and calibration, validation and cross-validation.
These methods are relevant for further data analysis in the biosystems engineering domain. The theory of the course will be supported by practicals in which relevant data sets from the biosystems engineering domain will be analyzed.
This course is tailormade for Bachelor Biosystems Engineering and part of the course is in Dutch. Students of other programs, please follow MAT-20306 Advanced Statistics which covers nearly the same statistics but with more general examples.
Program/Minor
Bachelor Biosystems EngineeringLearning Outcomes
Formulate a statistical hypothesis basen on a research question; Recognize a valid experimental design or sampling procedure for data collection; Select an appropriate statistical model that allows valid estimation, quantification of uncertainty, and hypothesis testing given the research question and properties of the data; Analyze the data with R Studio for a given data set and research question; Interpret the outcome of the statistical analysis and draw conclusions with respect to the stated problem
Prerequisites
Statistics 1 & Statistics 2
Details
Institution: WUR
Period: Period 1
Level: Intermediate
EC: 6 ec
Data Analysis for Plant and Animal Breeding
Data analysis is central to both plant and animal breeding, and the size and complexity of phenotypic and genomic data sets continue to increase. Thus, the ability to analyze and interpret such large data sets is an essential skill for breeders, both in science and industry. In this course you will become familiar with state-of-the-art
Data Analysis for Plant and Animal Breeding
Data analysis is central to both plant and animal breeding, and the size and complexity of phenotypic and genomic data sets continue to increase. Thus, the ability to analyze and interpret such large data sets is an essential skill for breeders, both in science and industry. In this course you will become familiar with state-of-the-art methods and skills for quantitative genetic analysis of breeding data, both for animals and plants.
This is a hands-on course, where you develop the skills to analyze real-life data and handle real-life problems in genetic analysis. Next to genetic analysis, this will include developing the skills to competently curate data sets in the R software environment. At the same time, you will develop an understanding of the statistical methods on an applied, practically relevant and intuitive level. This includes being able to choose an appropriate analysis based on the research question and the data at hand, understanding the statistical model and its assumptions, interpreting the results and becoming aware of common pitfalls. You will achieve this by working on illustrative real-life data sets that link to modern animal and plant breeding. The course covers the most important categories of statistical models and the associated methods for genetic and genomic data analysis and for model validation. During the course, you will gradually build up the required R-skills.
We make use of plenary lectures and computer tutorials focused on application using real-life data, and you will also work on two case studies. In each of the two case studies, you will analyze an actual data set and write a short report on the analysis. In the tutorials, you will learn how to use the R-software for data handling, editing, filtering and quantitative genetic analyses. You will also become familiar with more advanced methods for genetic analysis, with complex pedigreed and large genomic data, using dedicated software.
The course consists of six one-week modules. In the first week, you will become familiar with data handling, visualization and editing, and model building and model validation using linear models. In the next weeks you will become familiar with more advanced statistical models and tools, with a major focus on Linear Mixed Models, and also including Generalized Linear Models and Maximum Likelihood, and the use of these tools for quantitative genetic analysis of breeding data. In the final weeks, you will become familiar with more advanced analysis of genetic, genomic and phenotypic (big) data in animals and plants. This includes the estimation of genetic parameters such as heritability, QTL mapping, genomic prediction, and genome-wide association studies.
Program/Minor
European Master in Animal Breeding and Genetics; Master Animal Sciences; Master Bioinformatics; Master Plant SciencesLearning Outcomes
Apply data handling skills necessary to competently curate data sets in the R software environment; Choose a model category, build a model for quantitative genetic analysis of a given data set and research question, and execute the analysis; Interpret and explain the results of your data analysis; Perform model validation by evaluating model assumptions and/or cross validation (for genomic prediction), using illustrative plots; Explain the differences between a linear model (LM), linear mixed model (LMM) and a generalized linear model (GLM) in terms of model assumptions and purpose of the analysis; Explain the principles of maximum likelihood and restricted maximum likelihood; Explain the difference between fixed and random effects; Explain how the heritability of a trait can be estimated in pedigreed or genotyped populations in animals or plants; Design an experiment for estimating heritabilities, for QTL mapping and for genome-wide association studies (GWAS); Explain how genomic prediction can be performed, and propose statistical models for that purpose; Explain how genome-wide association studies or QTL-detection can be used to detect genomic regions of interest in outbred populations or in line crosses
Prerequisites
It is assumed that students who take this course have a basic understanding of statistics and some understanding of genetics. It is recommended to take the courses MAT15303 + MAT15403 and MAT20306 Advanced Statistics before taking part in the present course. Some experience with the R-software is helpful, but not mandatory.
Details
Institution: WUR
Period: Period 5
Level: Master
EC: 6 ec
Data Driven Discovery in the Life Sciences: Hypothesis Generation from Omics Data
Across the life sciences, scientists utilize omics data to study biological phenomena in humans, plants, animals and microbes. This results in large and heterogeneous data sets that can be analyzed using a variety of algorithms and statistical methods. Making sense of the data, extracting biological knowledge out of the results of these analyses and formulating
Data Driven Discovery in the Life Sciences: Hypothesis Generation from Omics Data
Across the life sciences, scientists utilize omics data to study biological phenomena in humans, plants, animals and microbes. This results in large and heterogeneous data sets that can be analyzed using a variety of algorithms and statistical methods. Making sense of the data, extracting biological knowledge out of the results of these analyses and formulating new hypotheses and research questions based on them is not trivial. However, when basic data science skills are combined with domain knowledge, either from literature or from databases accessed in a high-throughput manner, omics data constitute a goldmine for data-driven discovery of novel insights and hypotheses that can be tested in follow-up experiments.
This course will train students in linking domain knowledge to data using data science techniques and skills, in order to design omics experiments, evaluate the quality of the resulting data, interpret them in the light of literature and domain databases, and mine them to make discoveries and compose new research questions and hypotheses. Domain-specific case studies will allow students to directly apply their skills on data relevant to their specialization.
Program/Minor
Master Bioinformatics; Master Biology; Master Data Science for Food and Health; Master Nutrition and Health; Master Plant Biotechnology; Master Plant SciencesLearning Outcomes
Describe the advantages and limitations of different types of omics data; Access, in a high throughput manner, databases commonly used in the life sciences for the interpretation of omics data; Design effective omics experiments, with appropriate replicates, controls, controlling for batch effects, etc.; Evaluate the quality and limitations of omics data (quality of the raw data, technical/biological variation, etc.) based on the outcomes of statistical analyses; Interpret (processed) omics analysis results using domain knowledge, data mining (commonly used biological databases ) and literature mining; Extract knowledge from the data and synthesize this into a (possible) biological story and compose new research questions and hypotheses based on this (data-driven)
Prerequisites
Statistics and data analysis as applied to omics data, as treated in Data Analysis and Visualization or Molecular Systems Biology and in Statistics for Data Science; Basic coding skills in R and/or Python are considered very useful. The course is considered most useful for life science students who have recently acquired basic data science skills (e.g., through the data science track) and want to learn to apply these, or for bioinformatics students with a limited background in the life sciences who want to improve their data interpretation skills.
Details
Institution: WUR
Period: Period 6
Level: Advanced
EC: 6 ec
Data Management
This course covers database design and the use of databases in applications, with a focus on applications in the life sciences. Topics include the relational model, database design principles, the structured query language (SQL), including temporal and spatial queries. Data lifecycle topics and contemporary issues for data scientists and practitioners are also introduced, i.e. big
Data Management
This course covers database design and the use of databases in applications, with a focus on applications in the life sciences. Topics include the relational model, database design principles, the structured query language (SQL), including temporal and spatial queries. Data lifecycle topics and contemporary issues for data scientists and practitioners are also introduced, i.e. big data, FAIR principles, data governance, licensing, and privacy.
The course includes extensive practical work in the design, construction, and use of databases in the students’ field of study. Practical work involves MySQL and PowerBI. Students tend to value these hands-on database design and building skills as the most valuable part of the course.
Program/Minor
BSc Minor Data Science; BSc Minor Geo-information for Environment & Society; Bachelor Biosystems Engineering; Master Food Technology; Master Geo-Information Science; Master Plant Biotechnology; Master Plant Sciences (2025)Learning Outcomes
Demonstrate a managerial perspective on an organization's memory; Explain key concepts of data modelling and databases (i.e. entities, relationships, primary and foreign keys); Interpret data model diagrams using different notations (E-R diagrams); Compile Database queries with SQL: including joins, subqueries, arithmetic, logical and spatial operations; Analyse a realistic information system problem and propose a data design solution; Design and implement a Database for a problem in their field of study; Debate some of the contemporary challenges in data management, such as FAIR principles, security, big data, privacy, licensing
Prerequisites
At ease with computers and basic mathematics.
Details
Institution: WUR
Period: Period 1, Period 5
Level: Intermediate
EC: 6 ec
Data Mining
The goal of the course is to teach students how to think like a data miner. Intuitively, this means you have the mindset and skills to find practical solutions to common problems you encounter when extracting knowledge, patterns, and models from large data sets. To make such solutions effective, you must understand both the underlying
Data Mining
The goal of the course is to teach students how to think like a data miner. Intuitively, this means you have the mindset and skills to find practical solutions to common problems you encounter when extracting knowledge, patterns, and models from large data sets. To make such solutions effective, you must understand both the underlying intuition and its mathematical foundation. The field of data mining contains far too many of such practical solutions to teach in one course. We therefore focus on core techniques that demonstrate some of the magic behind modern data mining solutions: matrix decomposition, sketching and hashing, and embeddings and distances.
We cover the basic mathematical skills required to use these techniques effectively and you have to demonstrate mastery by developing solutions in 3 large lab assignments from scratch: Anomaly detection in system logs, recommender systems for profile matching and clustering in social networks
Data from these different domains often needs special (pre)processing methods to be able to apply machine learning/data mining methods. We discuss the main (pre)processing methods and their considerations in the course. Importantly, different distance measures and their effect on the mining outcome is a recurring theme. Also, special consideration is given to being able to deal with huge datasets through Smart approximations. Ethical considerations of data mining are discussed. In all of these topics, the course will cover key algorithms for similar-item retrieval, dimension reduction, large scale clustering, collaborative filtering, locality-sensitive hashing, outlier detection, profiling, and graph mining.
Program/Minor
BSc Computer Science and EngineeringLearning Outcomes
Explain and manually apply key data mining algorithms, including techniques for dimensionality reduction (SVD), distance measurement (DTW), clustering (kMeans++, DBSCAN), anomaly detection (Isolation Forest, PCA), sketching and hashing (CountMin Sketch, MinHashing), recommendation (Collaborative Filtering, NMF), text representation (Word2Vec), and graph embeddings; Make informed and justified design choices in applying data mining methods, including preprocessing, distance computation, clustering, dimensionality reduction, anomaly detection, sketching, recommendation systems, text mining, and graph mining; Apply data mining techniques to real-world challenges such as anomaly detection, recommendation systems, and large-scale graph mining; Critically evaluate ethical considerations when using data mining techniques in practical applications.
Prerequisites
Courses considered required, i.e. students are strongly recommended to have passed all: Introduction to Programming; Algorithms and Data Structures; Calculus; Linear Algebra; Probability Theory and Statistics; Machine Learning; Big Data Processing. Specific topics that are assumed as prior knowledge include: Discrete mathematics: set intersections, unions, and differences; Linear algebra: matrix multiplication, linear systems, eigendecompositions; Probability and statistics: multivariate Gaussian distribution and correlation and covariance (matrices); Programming: Java or Python programming skills; Data structures: arrays, linked lists, hash tables, and trees. These requirements are not enforced and for information purposes only. Gaps in prior knowledge will impede success in this course. Students are expected to be responsible for their own study success. We are transparent in our expectations to enable students to make the right choice for their skill set.
Details
Institution: TU Delft
Period: Period 2
Level: Intermediate
EC: 5 ec