From Data, to Models, and Back: Making Machine Learning Predictably Reliable

Ilyas, Andrew

dc.contributor.advisor	Daskalakis, Constantinos
dc.contributor.advisor	Mądry, Aleksander
dc.contributor.author	Ilyas, Andrew
dc.date.accessioned	2025-03-27T17:00:55Z
dc.date.available	2025-03-27T17:00:55Z
dc.date.issued	2025-02
dc.date.submitted	2025-03-04T17:21:19.744Z
dc.identifier.uri	https://hdl.handle.net/1721.1/158964
dc.description.abstract	Machine learning systems exhibit impressive performance, but we currently lack scalable ways to anticipate their successes, failure modes, and biases. This position limits our ability to deploy these systems in the appropriate contexts, and to build systems which we can confidently deploy in high-risk settings. Motivated by this state of affairs, this thesis aims to develop design principles for predictably reliable machine learning. Our ultimate goal is to enable developers to know when their models will work, anticipate when they will fail, and understand “why” in both cases. In pursuit of this goal, this thesis combines large-scale experiments with theoretical analysis to form a precise understanding of the ML “pipeline,” from training data (and the way we collect it), to learning algorithms, to deployment. Fully realized, such an understanding would allow us to build ML systems the same way we build buildings or airplanes—safely, scalably, and with a robust grasp of the underlying principles. In this thesis, we focus on four design choices within this pipeline: model deployment (Part I), dataset creation (Part II), data collection (Part III), and algorithm selection (Part IV). For each of these design choices, we use targeted experiments to uncover the corresponding principles that actually underlie the behavior of ML systems. We distill these principles into concise conceptual models which allow us to both reason about existing systems and design improved ones. Along the way, we will revisit, challenge, and refine various aspects of conventional wisdom surrounding ML model development.
dc.publisher	Massachusetts Institute of Technology
dc.rights	In Copyright - Educational Use Permitted
dc.rights	Copyright retained by author(s)
dc.rights.uri	https://rightsstatements.org/page/InC-EDU/1.0/
dc.title	From Data, to Models, and Back: Making Machine Learning Predictably Reliable
dc.type	Thesis
dc.description.degree	Ph.D.
dc.contributor.department	Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
mit.thesis.degree	Doctoral
thesis.degree.name	Doctor of Philosophy

Files in this item

Name:: ilyas-ailyas-phd-eecs-2025-the ...
Size:: 146.9Mb
Format:: PDF
Description:: Thesis PDF

View/Open

This item appears in the following Collection(s)

Doctoral Theses

Show simple item record