Statistics and Data Science for Health and Environment

Health and environmental data are often heterogeneous, multivariate, longitudinal, spatially and temporally correlated, and incomplete. They provide a rich field of application for statistical learning and data science. L2S develops statistical learning methods to extract meaningful knowledge from complex data and to build interpretable models.

This research relies on long-term collaborations with various partners, including the Brain Institute (ICM), Assistance Publique-Hôpitaux de Paris (APHP), CEA Neurospin, Gustave Roussy, Centre National de Recherche en Génomique Humaine (CNRGH), The Institut Pasteur, and the European ATHLETE project, to name a few. Our partners provide access to diverse multimodal cohorts and health data warehouses.

Research themes

L2S develops a general statistical framework for the joint analysis of multiple sources of information. The framework, called Regularized Generalized Canonical Correlation Analysis (RGCCA), enables the integration of heterogeneous data types, including tensor-valued, matrix-valued, functional and longitudinal data. It also incorporates variable selection to identify the most relevant variables, thereby facilitating the interpretation of the results and improving the interpretability of the resulting models. Our work also investigates structural equation models to move beyond correlation and study causal relationships. In addition, we study time-series models for count data, multivariate observations, and outliers. These methods are used to analyze relationships between air pollution and health, model respiratory diseases, and characterize the effects of environmental exposure over time.

Data integration for neurodegenerative diseases


Neurodegenerative diseases (for instance: Alzheimer’s, Parkinson’s, Multiple Sclerosis, etc.) have multifactorial background. In that context, the observation of the patient in its various facets (e.g. multimodal imaging, omics data, clinical data) is crucial. Therefore, it is mandatory to develop advanced statistical methods for the joint analysis of these large and heterogeneous data sources.
Partnership. ICM, Neurospin, CNRGH, Institut Pasteur

Selected publications.

Girka et al. (2024) Tensor generalized canonical correlation analysis, Information Fusion, Volume 102, 102045
Gloaguen et al. (2022). Multiway generalized canonical correlation analysis. Biostatistics, Volume 23, Issue 1, Pages 240–256
Fransson et al. (2024). Multiple sclerosis patient macrophages impaired metabolism leads to an altered response to activation stimuli. Neurology: Neuroimmunology & Neuroinflammation, 11(6), e200312.
Xicota, L. et al. (2019). Multi-omics signature of brain amyloid deposition in asymptomatic individuals at-risk for Alzheimer’s disease: The INSIGHT-preAD study. EBioMedicine, 47, 518-528.
Garali et al. (2018). A strategy for multimodal data integration: application to biomarkers identification in spinocerebellar ataxia. Briefings in Bioinformatics, 19 (6), 1356-1369

Chair in Artificial Intelligence and Health



The goal of the AP-HP Health Data Repository (EDS) is to integrate all the medical data collected from patients hospitalized in one of the 39 AP-HP hospitals. The main challenge related to the exploitation of such longitudinal, heterogeneous, multi-centric biomedical data is to improve the field of healthcare for more personalized medicine. In this context, it is crucial to use/develop statistical methods to address the various questions of the clinicians from this massive amount of data. In addition, we explore the use of pre-trained models to make AI more data-efficient in medical applications.
Partnership: APHP – L2S- INRIA – DataIA

Selected publications.

Sort et al. (2026) Latent Functional PARAFAC for Modeling Multidimensional Longitudinal Data. Psychometrika. 26:1-25.
Sort et al. (2024) Functional Generalized Canonical Correlation Analysis for studying multiple longitudinal variables, Biometrics, 80(4)

Relationship between exposome and health

The exposome is defined as the set of environmental exposures received over the course of lifetime. The study of the effects of the exposome on health is a major public health issue. This field is undergoing a data revolution, making any statistical analysis difficult. L2S offers advanced statistical tools for the analysis of exposome data.
Partnership: L2S – Athlete European project – UFES (Brésil)

Selected publications.

Tenenhaus et al. (2026). Structural equation modeling with factors and composites within the framework of the basic design. Advances in Data Analysis and Classification, 20(1), 227-254.
Amine et al. (2025). Early-life exposome and health-related immune signatures in childhood. Environment International, 202, 109668
Camara et al (2025). Robust estimate for count time series using GLARMA models: An application to environmental and epidemiological data. Applied Mathematical Modelling, 137, 115658.
Danilevicz et al. (2026). Adaptive LASSO quantile regression with fixed effects. Applied Mathematical Modelling, 153, 116600.
Reisen et al. (2024). A dimension reduction factor approach for multivariate time series with long-memory: a robust alternative method. Statistical Papers, 65(5), 2865-2886.
Danilevicz et al. (2024). A longitudinal study of the influence of air pollutants on children: a robust multivariate approach. Journal of Applied Statistics, 51(11), 2178-2196.

Open-source softwares

L2S provides the scientific community with several open-source software packages:

  • The RGCCA R package provides a comprehensive implementation of the RGCCA framework, including highly efficient algorithms, graphical displays, and tools to assess the robustness and significance of results. It is also distributed through CRAN.
  • The semFC R package implements highly efficient algorithm with sound statistical properties for Structural Equation Modeling with Factors and Composites.
  • Quantile regression with fixed effects provides a flexible framework for analyzing longitudinal data while accounting for individual-specific heterogeneity. The prqfe and alqfe R packages implement several estimation methods based on different loss functions and model structures, including LASSO-penalized fixed effects. Adaptive LASSO provides desirable oracle properties, while BIC can be used to select the optimal tuning parameter.

Selected publications

Girka et al. (2026). Multiblock data analysis with the RGCCA package. Journal of Statistical Software, 1-36

Contact


Arthur TENENHAUS

Professor – CentraleSupélec

Signaux et statistiques – Signal & Stat

.

Bât. Breguet .