Three days of lectures, workshops and conversations at IME-USP, bringing together mathematicians, statisticians and AI researchers to examine the mathematical foundations beneath the recent successes of artificial intelligence — successes which are almost as impressive as they are poorly understood.
Speakers
Click on a presentation title to view its abstract.
Prof. Rafael Izbicki
UFSCAR – São Carlos
Abstract :
Prof. Paola Bermolen
Universidad de la República – Montevideo
Abstract :
A simple yet sufficiently expressive and interpretable model for random graphs is the Random Dot Product Graph (RDPG), which also comes with an inference method with statistical guarantees known as Adjacency Spectral Embedding (ASE). In this talk, we present this model alongside generalisations that allow us to overcome some of its best-known limitations, in all cases seeking to preserve its fundamental properties of interpretability and guarantees. We frame ASE as an optimisation problem, which enables the definition of inference methods that handle missing data or dynamic networks, where the number of nodes may, for example, increase or decrease over time. We also propose new inference methods for directed graphs, and analyse the corresponding optimisation landscape, proving that it is benign. Furthermore, we introduce a nonparametric latent position model for random weighted graphs by relating the inner product of the nodal latent position vectors to the moment generating function of the edge weight distribution, proving statistical guarantees such as consistency and asymptotic normality. These spectral embeddings can also be used to generate new graphs that statistically resemble the structure and edge weight distribution of a given real graph. Finally, we discuss some further limitations and new directions, including heterophilious graphs and hyperbolic geometry.
Prof. Mauricio Velasco
Universidad de la República – Montevideo
Abstract :
In this talk I will recall the basic definitions of statistical learning theory and then survey the proof, due to Bach, that statistical learning is universal for one-layer neural networks (of possibly infinite width). Crucially Bach’s results come with both explicit algorithms and with quantitative guarantees on the amount of data sufficient for learning, and fully solve the “mystery” of machine learning via shallow networks. In the last part of the talk I will present some recent work on Laplace-Barron spaces, which generalizes Bach’s work to produce a theory of learning on general convex cones. I will discuss some applications of this theory for learning Laplace transforms of random variables and for some problems in physics related to the computation of Feynman integrals.
The talk will be self-contained and will not assume any prior knowledge beyond multivariable calculus and probability.
Prof. Joaquin Fontbona
Universidad de Chile – Santiago
Abstract :
We will first overview the basic ideas of generative diffusion models from a machine learning practitioner’s viewpoint. Then, we will propose a purely probabilistic perspective on these models, their aims, and their training methods, leveraging stochastic analysis tools and results—some of them classic, such as Girsanov’s theorem and Doob’s h-transform, and some less standard, like Föllmer processes and the enlargement of filtrations. We will also discuss how these concepts are (explicitly or not) built-in and utilized by goal-specific sampling methods, such as classifier-free guidance and Schrödinger Bridge matching.
On this basis, we will propose a general probabilistic framework that unifies several generative diffusion settings, their sampling goals, and their training methods. If time permits, we will also discuss some potential new (yet to be developed) applications of this framework, including in multimodal conditional sampling (e.g., text-to-image and vice versa).
Based on ongoing works with David Felipe and Lucas Villanueva (U. of Chile)
Prof. Paulo Oreinstein
IMPA – Rio de Janeiro
Abstract :
Averaging is everywhere in modern AI (e.g., attention layers, ensembling, cross-validation, gradient aggregation) and every average implicitly answers a question: how should observations be weighted? The sample mean is efficient and minimax but fragile under heavy tails; robust estimators resist heavy tails but typically sacrifice efficiency. We show a soft-shrinkage calibration improves on both fronts: observations are smoothly downweighted by their distance from a pilot estimate, with the shrinkage scale calibrated so that the total rejected weight equals a prescribed budget. Any reasonable pilot is upgraded to sub-Gaussian concentration with the minimax constant $\sqrt{2}$ under finite variance alone, and the budget doubles as an exact breakdown point. The mechanism trades variance for bias, and the bias can be quantified exactly. This yields a second-order theory establishing an MSE expansion that identifies when shrinkage beats the sample mean, Berry–Esseen and moderate deviation approximations, and asymptotically exact confidence intervals. Beyond the theory, the results give concrete guidance for designing mean estimators in statistics and AI pipelines, and bridge modern sub-Gaussian estimation with classical robust statistics. This is joint work with Antônio Catão.
Prof. Thiago Ramos
UFSCAR- São Carlos
Abstract :
Many learning problems require more than predicting an output from an individual input: the goal is to learn an operator that acts on functions and captures the relationship between input and output distributions. We propose Functional Newton Methods for Operator Learning, a framework that learns such operators through second-order optimization in function space. The resulting representation supports applications including regression, uncertainty quantification, and dependence assessment.
Prof. Marcelo Finger
IME – USP
Abstract :
Prof. Hamed Yazdanpanah
IME – USP
Abstract :
Time series data underpin decision-making across energy, finance, healthcare, retail, and traffic domains, yet the dominant modeling paradigms have evolved through three distinct stages: local statistical models, global deep learning models, and, most recently, time series foundation models (TSFMs). TSFM are models pretrained on massive, heterogeneous corpora spanning domains, frequencies, and scales, intended to learn general-purpose temporal representations that transfer to unseen data with little or no task-specific training. This talk surveys the foundations, methodology, and open challenges of TSFMs. We formalize the foundation-model objective as learning a representation mapping rather than a direct predictor, and discuss the architectural landscape, key design axes, and the self-supervised objectives that drive representation and transfer learning. We then argue that strong average predictive accuracy does not imply statistical reliability, and discuss four open problems that scaling alone does not resolve: generalization under distribution shift, calibration of uncertainty under temporal dependence, robustness to nonstationarity, and evaluation validity in the presence of benchmark contamination. We close by outlining four directions for a statistical theory of TSFMs: transferability, uncertainty, adaptation, and scaling, needed to move the field from empirical leaderboard performance toward principled, reliable deployment.
Prof. Nina S. T. Hirata
IME – USP
Abstract :
Modern deep learning has achieved remarkable performance mainly through training very large architectures on huge volumes of data. However, this parameter-heavy paradigm is increasingly unsustainable and remains prone to overfitting whenever data is scarce. At its core, deep learning can be understood as representation learning: transforming high-dimensional inputs into structured, meaningful feature spaces. When data is limited, finding an effective representation in an unconstrained search space becomes exceptionally difficult. In this talk, we examine how inductive bias serves as a principled mathematical mechanism to guide this learning process. We will present some examples and discuss how embedding these priors into model architectures improves sample efficiency, reduces parameter counts, and helps prevent overfitting, offering a more sustainable path for model design.
Prof. Tom Hanika
University of Hildesheim
Abstract :
This is joint work with Friedrich Martin Schneider.
Program
Time
09:00 to 10:15
Monday, 10 August
J. Fontbona
Tuesday, 11 August
M. Velasco
Wednesday, 12 August
H. Yazdanpanah
10:15 to 10:45
Coffee break
10:45 to 12:00
P. Bermolen
R. Izbicki
N. Hirata
12:00 to 14:00
Lunch break
14:00 to 15:15
T. Ramos
P. Orenstein
INCT meeting
15:15 to 15:45
Coffee break
15:45 to 17:00
M. Finger
T. Hanika
Working sessions
Scientific Comittee
- Florencia Leonardi / USP – São Paulo
- Hedibert Freitas Lopes / INSPER – São Paulo
- Roberto Imbuzeiro Oliveira / IMPA – Rio de Janeiro
Organizing Comittee
- Aline Duarte / USP – São Paulo
- Florencia Leonardi / USP – São Paulo
- Morgan André / USP – São Paulo
Support staff
- Lourdes Vaz da Silva / USP – São Paulo
- Renata Stella Khouri / USP – São Paulo
- Liena Valero Bello / USP – São Paulo
- Arthur Henrique Dias Rodrigues / USP – São Paulo
- Matheus Teixeira / USP – São Paulo
Organized by:
Supported by:
inct.events@ime.usp.br