Loading Events

« All Events

The Geometry of Machine Learning 2026

September 8, 2026 @ 9:00 am - September 11, 2026 @ 5:00 pm

The Geometry of Machine Learning 2026

Dates: September 8–11, 2026

Location: Harvard CMSA, Room G10, 20 Garden Street, Cambridge MA 02138 & via Zoom Webinar

Register to attend in person

Register for Zoom Webinar

Large language models are presently, and will increasingly, be complemented by other dimensions of intelligence: formal verification and energy-based optimizers, becoming parts of larger ecosystems. Can AIs reason geometrically and can we use geometry to reveal how data is currently processed in NNs? Can AIs reveal the geometry of mathematics, as well as studying geometry as a subject within math. This conference is intended to continue the discussion of these topics.

Confirmed Speakers:

Organizers: Michael R. Douglas (CMSA) and Mike Freedman (CMSA)

 

Schedule

Tuesday, Sep. 8, 2026

8:15–8:45 am
Breakfast

8:45–9:30 am
Mike Mulligan, UCR, Logical Intelligence

9:45–10:30 am
Boris Hanin, Princeton

10:30–11:00 am
Break

11:00–11:45 am
Slava Krushkal, University of Virginia

12:00–12:45 pm
Robert Koirala and Bennett Chow, UCSD
AI for Ricci flow: Discovery and Formalization

 

Wednesday, Sep. 9, 2026

8:15–8:45 am
Breakfast

8:45–9:30 am
Surya Ganguli, Stanford

9:45–10:30 am
Mathew Vanherreweghe, Logical Intelligence

10:30–11:00 am
Break

11:00–11:45 am
Michael Brenner, Harvard and Google

12:00–12:45 pm
Mattiew Wyart, JHU
Learn from your own latents, not from tokens
Abstract: Language models need more than a hundred thousand times the data a child does. One explanation is that predicting raw tokens is simply the wrong level: methods like data2vec and JEPA instead train a network to predict its own internal representations, with strong empirical results but no theory of why. Using a hierarchical grammar that models that language and images have a hidden hierarchical structure, we quantify the gain exactly. Token-level learning needs a number of examples growing exponentially with the depth of the hierarchy; latent prediction needs a number independent of it. We also show data2vec performs this hierarchical prediction implicitly, which suggests that explicitly stacking levels – as in H-JEPA – yields little.

 

Thursday, Sep. 10, 2026

8:15–8:45 am
Breakfast

8:45–9:30 am
Dmitry Krotov, Dynamical Mind
Dense Associative Memory: Physical systems for novel AI architectures
Abstract: Dense Associative Memories are recurrent neural networks with fixed-point attractor states that are described by an energy function. In contrast to conventional Hopfield Networks, which were popular in the 1980s, Dense Associative Memories have a very large information storage capacity, making them appealing tools for many problems in AI. In this talk, I will provide an intuitive understanding and mathematical framework for this class of models and give examples of problems in AI that can be tackled using these new ideas. Specifically, I will explore the relationship between Dense Associative Memories and transformers. I will present a neural network called the Energy Transformer, which unifies energy-based modeling, associative memories, and transformers in a single architecture. I will demonstrate how Energy Transformers can be used for challenging tasks in image processing, solve partial differential equations, and serve as computational modules for energy-based language modeling. I will also discuss an exciting possibility of mapping these models onto analog hardware accelerators, which could enable much more energy-efficient inference compared to GPUs.

9:45–10:30 am
Roi Holtzman, Oxford

10:30–11:00 am
Break

11:00–11:45 am
Sean Welleck, CMU (via Zoom)

12:00–12:45 pm
Gabriel Poesia, University of Michigan (via Zoom)

Friday, Sep. 11, 2026

8:15–8:45 am
Breakfast

8:45–9:30 am
Nada Amin, Harvard

9:45–10:30 am
Randall Balestriero, Brown University

10:30–11:00 am
Break

11:00–11:45 am
Luca Pesce, Harvard CMSA
A spiked perspective on feature learning with gradient-based methods
Abstract: Depth is widely believed to give neural networks a clear computational advantage over shallow models, and making this belief precise is a central problem in learning theory. We study a controlled high-dimensional setting where this can be done. The targets are hierarchical: the relevant structure is distributed across latent subspaces of decreasing dimension, so that a shallow model must resolve all of it simultaneously, while a deep network does not. Each layer forms an intermediate representation, allowing learning to proceed in stages, with every stage reducing the effective dimension of the remaining problem. We analyze this staged mechanism through the gradient descent dynamics of a deep network yielding a sharp separation in sample complexity between shallow and deep architectures. The result is a concrete account of why depth allows such functions to be learned from substantially fewer samples than shallow methods require.

12:00–12:45 pm

Talk tba

 

Support provided by Logical Intelligence.

 

 

Details

  • Start: September 8, 2026 @ 9:00 am
  • End: September 11, 2026 @ 5:00 pm
  • Event Categories: ,