Loading Events

« All Events

Big Data Conference 2026

September 3, 2026 @ 9:00 am - September 4, 2026 @ 5:00 pm

Big Data Conference 2026

Dates: Sep. 3–4, 2026

Location: Harvard University CMSA, 20 Garden Street, Cambridge MA & via Zoom

The Big Data Conference features speakers from the Harvard community as well as scholars from across the globe, with talks focusing on computer science, statistics, math and physics, and economics.

Register to attend in person

Register for Zoom Webinar

 

Confirmed Speakers

 

Organizers

 

 

Thursday, Sep. 3, 2026

8:45–9:10 am
Breakfast

9:10–9:15 am
Introductions

9:15–10:15 am
Rex Ying, Yale

10:15–10:30 am
Break

10:30–11:30 am
Adit Radhakrishnan, MIT
Toward universal steering and monitoring of AI models
Abstract: Artificial intelligence (AI) models contain much of human knowledge. Understanding the representation of this knowledge will lead to improvements in model capabilities and safeguards. Building on advances in feature learning, we developed an approach for extracting linear representations of semantic notions or concepts in AI models. We showed how these representations enabled model steering, through which we exposed vulnerabilities and improved model capabilities. We demonstrated that concept representations were transferable across languages and enabled multiconcept steering. Across hundreds of concepts, we found that larger models were more steerable and that steering improved model capabilities beyond prompting. We showed that concept representations were more effective for monitoring misaligned content than for using judge models. Our results illustrate the power of internal representations for advancing AI safety and model capabilities.

11:30 am–12:45 pm
Lunch

12:45–1:45 pm
tba

1:45–2:00 pm
Break

2:00–3:00 pm
Bailey Flanigan, MIT
Algorithmic Tools for Trading Off Sortition Ideals
Abstract: Citizens’ assemblies and other deliberative minipublics — representative groups of everyday people convened to deliberate on a policy issue and then make recommendations — are now used by governments around the world. Choosing who sits on these panels is the problem of sortition: randomly selecting a small group of citizens that represents the broader population. Sortition has been the subject of substantial computer science research in recent years, and the resulting algorithms are now widely used in practice. This talk will describe the key challenges that arise in the practice of sortition, and the algorithmic tools that have been developed to navigate them optimally.

3:00–3:15 pm
Break

3:15–4:15 pm
Chris Wiggins, Columbia

 

Friday, Sep. 4, 2026

8:45–9:15 am
Breakfast

9:15–10:15 am
Ariel Procaccia, Harvard
No Generation Without Representation
Abstract: AI systems and democratic processes are confronting similar challenges around representation. I examine two related questions that cut across both domains. First, how can AI enable democratic processes that handle vast spaces of opinions or statements while ensuring proportional representation of a population’s views? Second, when AI systems themselves provide normative guidance, whose viewpoints do they reflect, and can we make this precise? Drawing on social choice theory, I present formal frameworks and algorithms for both problems, showing that meaningful representation guarantees are feasible and practical.

10:15–10:30 am
Break

10:30–11:30 am
Wei Zhou, Harvard
Biobank-scale genetic discovery: from association testing to global meta-analysis
Abstract: Biobanks linking genomic data with electronic health records provide unprecedented opportunities for genetic discovery for complex human diseases, but they also pose analytical challenges that extend well beyond sample size. Within a biobank, association studies must account for population structure and relatedness, highly unbalanced case–control ratios, rare genetic variants, longitudinal and censored outcomes, and the computational demands of analyzing hundreds of thousands of individuals and millions of genetic variants. Across biobanks, additional challenges arise from differences in ancestry, phenotype definitions, recruitment strategies and genetic effects.In this talk, I will discuss statistical and computational methods developed to address these challenges at successive stages of biobank analysis. These include scalable generalized linear mixed models for binary traits, survival mixed models for censored time-to-event outcomes, and gene- and region-based tests that aggregate rare variants. I will describe statistical approximations and computational strategies that make these analyses feasible at biobank scale while maintaining calibration in the presence of relatedness and highly unbalanced phenotypes. I will then introduce the Global Biobank Meta-analysis Initiative (GBMI) and describe how genetic evidence can be combined across biobanks without sharing individual-level data. Examples from GBMI will illustrate how combining evidence across biobanks can increase statistical power through larger sample sizes and broaden genetic discovery through greater ancestral diversity.

11:30 am–12:45 pm
Lunch

12:45–1:45 pm
Sergey Ovchinnikov, MIT

1:45–2:00 pm
Break

2:00–3:00 pm
SueYeon Chung, Harvard

3:00–3:15 pm
Break

3:15–4:15 pm
Andrew Sutherland, MIT


 

Details

Venue