Three days of lectures, workshops and conversations at IME-USP, bringing together mathematicians, statisticians and AI researchers to examine the mathematical foundations beneath the recent successes of artificial intelligence — successes which are almost as impressive as they are poorly understood.

Speakers

Click on a presentation title to view its abstract.

 

Prof. Rafael Izbicki

UFSCAR – São Carlos

Abstract :

Modern AI systems are increasingly used in settings where point predictions are insufficient: scientific inference, medical risk assessment, forecasting, and decision-making require calibrated uncertainty, predictive distributions, and valid prediction sets. In this talk, I will discuss how statistical and mathematical tools can help move AI beyond point prediction. I will focus on conditional density estimation, calibration diagnostics, and conformal prediction, with emphasis on recent work incorporating epistemic uncertainty into conformal scores. I will also discuss how recent tabular foundation models can be evaluated as conditional density estimators, and why post-hoc calibration and distribution-free guarantees remain central even as predictive models become more powerful. The talk will argue that one important contribution of mathematics to AI is to provide the language, diagnostics, and guarantees needed to decide when AI predictions can be trusted.

 

Prof. Paola Bermolen

Universidad de la República – Montevideo

Abstract :

A simple yet sufficiently expressive and interpretable model for random graphs is the Random Dot Product Graph (RDPG), which also comes with an inference method with statistical guarantees known as Adjacency Spectral Embedding (ASE). In this talk, we present this model alongside generalisations that allow us to overcome some of its best-known limitations, in all cases seeking to preserve its fundamental properties of interpretability and guarantees. We frame ASE as an optimisation problem, which enables the definition of inference methods that handle missing data or dynamic networks, where the number of nodes may, for example, increase or decrease over time. We also propose new inference methods for directed graphs, and analyse the corresponding optimisation landscape,  proving that it is benign. Furthermore, we introduce a nonparametric latent position model for random weighted graphs by relating the inner product of the nodal latent position vectors to the moment generating function of the edge weight distribution, proving statistical guarantees such as consistency and asymptotic normality. These spectral embeddings can also be used to generate new graphs that statistically resemble the structure and edge weight distribution of a given real graph. Finally, we discuss some further limitations and new directions,  including heterophilious graphs and hyperbolic geometry. 

Joint work with: Marcelo Fiori, Federico La Rocca, Bernardo Marenco and Gonzalo Mateos.

 

Prof. Mauricio Velasco

Universidad de la República – Montevideo

Abstract :

In this talk I will recall the basic definitions of statistical learning theory and then survey the proof, due to Bach, that statistical learning is universal for one-layer neural networks (of possibly infinite width). Crucially Bach’s results come with both explicit algorithms and with quantitative guarantees on the amount of data sufficient for learning, and fully solve the “mystery” of machine learning via shallow networks. In the last part of the talk I will present some recent work on Laplace-Barron spaces, which generalizes Bach’s work to produce a theory of learning on general convex cones. I will discuss some applications of this theory for learning Laplace transforms of random variables and for some problems in physics related to the computation of Feynman integrals.
The talk will be self-contained and will not assume any prior knowledge beyond multivariable calculus and probability. 

 

Prof. Joaquin Fontbona

Universidad de Chile – Santiago

Abstract :

Generative diffusion models are one of the most successful and flexible families of AI tools, with applications today going far beyond their original goal of sampling from an empirically learned image distribution. They are also one of the few modern AI models that originated from well-established mathematical ideas—namely, the theory of Markov diffusion processes and their time-reversals. Despite this early connection, and many subsequent developments drawing inspiration from probabilistic ideas in case-specific ways, the theory of generative diffusion models as probabilistic objects themselves is surprisingly far from complete.
We will first overview the basic ideas of generative diffusion models from a machine learning practitioner’s viewpoint. Then, we will propose a purely probabilistic perspective on these models, their aims, and their training methods, leveraging stochastic analysis tools and results—some of them classic, such as Girsanov’s theorem and Doob’s h-transform, and some less standard, like Föllmer processes and the enlargement of filtrations. We will also discuss how these concepts are (explicitly or not) built-in and utilized by goal-specific sampling methods, such as classifier-free guidance and Schrödinger Bridge matching.
On this basis, we will propose a general probabilistic framework that unifies several generative diffusion settings, their sampling goals, and their training methods. If time permits, we will also discuss some potential new (yet to be developed) applications of this framework, including in multimodal conditional sampling (e.g., text-to-image and vice versa).
 Based on ongoing works with David Felipe and Lucas Villanueva (U. of Chile)

 

Prof. Paulo Oreinstein

IMPA – Rio de Janeiro

Abstract :

Averaging is everywhere in modern AI (e.g., attention layers, ensembling, cross-validation, gradient aggregation) and every average implicitly answers a question: how should observations be weighted? The sample mean is efficient and minimax but fragile under heavy tails; robust estimators resist heavy tails but typically sacrifice efficiency. We show a soft-shrinkage calibration improves on both fronts: observations are smoothly downweighted by their distance from a pilot estimate, with the shrinkage scale calibrated so that the total rejected weight equals a prescribed budget. Any reasonable pilot is upgraded to sub-Gaussian concentration with the minimax constant $\sqrt{2}$ under finite variance alone, and the budget doubles as an exact breakdown point. The mechanism trades variance for bias, and the bias can be quantified exactly. This yields a second-order theory establishing an MSE expansion that identifies when shrinkage beats the sample mean, Berry–Esseen and moderate deviation approximations, and asymptotically exact confidence intervals. Beyond the theory, the results give concrete guidance for designing mean estimators in statistics and AI pipelines, and bridge modern sub-Gaussian estimation with classical robust statistics. This is joint work with Antônio Catão. 

 

Prof. Thiago Ramos 

UFSCAR- São Carlos

Abstract :

Many learning problems require more than predicting an output from an individual input: the goal is to learn an operator that acts on functions and captures the relationship between input and output distributions. We propose Functional Newton Methods for Operator Learning, a framework that learns such operators through second-order optimization in function space. The resulting representation supports applications including regression, uncertainty quantification, and dependence assessment.

 

Prof. Marcelo Finger

IME – USP

Abstract :

This talk presents an overview of a research line on representing neural networks in Łukasiewicz logic and applying this representation to formal property verification. The approach relies on the correspondence between neural network computations and rational McNaughton functions, enabling the translation of certain neural architectures into logical formulas. Formal verification is then performed via automated theorem proving in Łukasiewicz infinitely-valued logic. We present published results on the logical representation of ReLU–TId neural networks and on the encoding of reachability and robustness properties in Łukasiewicz logic.
 
This is joint work with Sandro Preto.

 

Prof. Hamed Yazdanpanah

IME – USP

Abstract :

Time series data underpin decision-making across energy, finance, healthcare, retail, and traffic domains, yet the dominant modeling paradigms have evolved through three distinct stages: local statistical models, global deep learning models, and, most recently, time series foundation models (TSFMs). TSFM are models pretrained on massive, heterogeneous corpora spanning domains, frequencies, and scales, intended to learn general-purpose temporal representations that transfer to unseen data with little or no task-specific training. This talk surveys the foundations, methodology, and open challenges of TSFMs. We formalize the foundation-model objective as learning a representation mapping rather than a direct predictor, and discuss the architectural landscape, key design axes, and the self-supervised objectives that drive representation and transfer learning. We then argue that strong average predictive accuracy does not imply statistical reliability, and discuss four open problems that scaling alone does not resolve: generalization under distribution shift, calibration of uncertainty under temporal dependence, robustness to nonstationarity, and evaluation validity in the presence of benchmark contamination. We close by outlining four directions for a statistical theory of TSFMs: transferability, uncertainty, adaptation, and scaling, needed to move the field from empirical leaderboard performance toward principled, reliable deployment.

 

Prof. Nina S. T. Hirata

IME – USP

Abstract :

Modern deep learning has achieved remarkable performance mainly through training very large architectures on huge volumes of data. However, this parameter-heavy paradigm is increasingly unsustainable and remains prone to overfitting whenever data is scarce. At its core, deep learning can be understood as representation learning: transforming high-dimensional inputs into structured, meaningful feature spaces. When data is limited, finding an effective representation in an unconstrained search space becomes exceptionally difficult. In this talk, we examine how inductive bias serves as a principled mathematical mechanism to guide this learning process. We will present some examples and discuss how embedding these priors into model architectures improves sample efficiency, reduces parameter counts, and helps prevent overfitting, offering a more sustainable path for model design.

Prof. Tom Hanika

University of Hildesheim

Abstract :

Intrinsic dimension aims to capture the effective complexity of a data set independently of the dimension of its ambient space. Most approaches assume that the data is sampled from a low- dimensional manifold embedded in Euclidean space. We follow a different route and build on a link discovered by Pestov, who suggested that the curse of dimensionality might be a manifestation of the concentration of measure phenomenon. Together with Gromov’s metric measure geometry, this yields a notion of intrinsic dimension whose central instrument is the  observable diameter. It requires neither a manifold assumption nor Euclidean structure, and is defined for vector data, graphs and binary data alike.
We discuss the mathematical properties of this dimension and how it can be estimated for large data sets. As applications, we analyze the benchmark data used in graph neural network (GNN) research, where the analysis has consequences for the reproducibility of published results, and derive an unsupervised criterion for feature selection in GNNs and beyond.
This is joint work with Friedrich Martin Schneider.

Program

Time

09:00 to 10:15 

Monday, 10 August

J. Fontbona

Tuesday, 11 August

 M. Velasco

Wednesday, 12 August

 H. Yazdanpanah

10:15 to 10:45

 Coffee break

10:45 to 12:00 

P. Bermolen

R. Izbicki

N. Hirata

12:00 to 14:00

Lunch break

14:00 to 15:15

 T. Ramos

P. Orenstein

 INCT meeting

15:15 to 15:45

Coffee break

15:45 to 17:00 

 M. Finger

T. Hanika

Working sessions

Scientific Comittee

  • Florencia Leonardi / USP – São Paulo
  • Hedibert Freitas Lopes / INSPER – São Paulo
  • Roberto Imbuzeiro Oliveira / IMPA – Rio de Janeiro

Organizing Comittee

  • Aline Duarte / USP – São Paulo
  • Florencia Leonardi / USP – São Paulo
  • Morgan André / USP – São Paulo

Support staff

  • Lourdes Vaz da Silva / USP – São Paulo
  • Renata Stella Khouri / USP – São Paulo
  • Liena Valero Bello / USP – São Paulo
  • Arthur Henrique Dias Rodrigues / USP – São Paulo
  • Matheus Teixeira / USP – São Paulo

Organized by:

INCT_novo_colorido

Supported by:

Registration for the workshop is free of charge, and some financial resources are available for participants outside the city of São Paulo. To ask for financial support, please register before July 5th. 

For registration, click here:

Results about acceptance of participation and confirmation of financial support can be expected to arrive about July 15th.
 For other inquiries, please contact us at:

inct.events@ime.usp.br