Department of Statistics and Data Science
Rebecca Nugent, Department Head
Peter Freeman, Director of Undergraduate Studies
Zach Branson, Assistant Director of the Undergraduate Program
Samantha Nielsen, Associate Director of Academic Programs
Amanda Mitchell, Lead Senior Academic Advisor
Glenn Clune, Academic Program Manager
Sylvie Aubin, Academic Program Manager
Peter Long, Academic Advisor
Email: statadvising@andrew.cmu.edu
Location: Baker Hall 129
www.stat.cmu.edu/
In Statistics and Data Science, we train students to be the people who determine the data to collect, visualize and analyze that data, make decisions based on that data, and communicate that data-driven process to collaborators and stakeholders. Because of this, our students regularly get sought-after jobs in many industries and attend top graduate programs.
Curriculum and Training in Statistics and Data Science
Our faculty have built (and continue to build) an innovative curriculum to ensure students have state-of-the-art training and can adapt in an ever-changing, data-driven world. Our curriculum focuses on four areas that are at the foundation of Statistics & Data Science:
-
Critical Thinking: Students formulate questions that can be answered with data, and pinpoint how to collect, visualize, and analyze data to answer those questions. Furthermore, students assess when visualization, analysis, and other tools (e.g., models, algorithms, computing, AI) are appropriate, reliable, and trustworthy.
-
Methods and Applications: Students implement state-of-the-art methods for processing, visualizing, and analyzing data with computing tools to do these tasks quickly-but-reliably. Students work with real-world scenarios, such that they learn how method choices are informed by the data and the context in which they are used.
-
Theory: Students gain the mathematical training necessary to characterize the effectiveness of data analysis methods in different settings (via probability, calculus, and statistical models), including new methods they develop or encounter in the world.
-
Communications: Students communicate, argue, and defend complex analyses to collaborators and stakeholders (in writing, presentations, and other modalities), such that data science work informs real-world decisions. This includes open-ended projects via experiential learning opportunities.
Our faculty have nurtured collaborations in many areas (e.g., astronomy, climate science, epidemiology, genetics, medicine, public policy, and sports analytics, to name a few), and the entire faculty teach courses at all levels, giving students different perspectives from world-renowned researchers and educators. Beyond courses, students gain expertise from research and experiential learning opportunities (e.g., our capstone courses, where students build data science products for non-statistical clients) and participate in our friendly, energetic community via student-led clubs and events.
On the remainder of this site, we outline the requirements for our degrees. This includes the core major in Statistics and Data Science, which can be augmented with several disciplines (economics, machine learning, mathematics, or neuroscience). Importantly, the core major involves a self-defined concentration of the student’s choosing (e.g., computer science, business, computational humanities, physics, psychology, public policy), such that students can explore how statistics and data science intersect with other disciplines. We also outline requirements for our minor in Statistics and Data Science.
Note for Current Students: Use this information as a general guide, and then schedule a meeting with a Statistics and Data Science Undergraduate Advisor (statadvising@stat.cmu.edu) to discuss these requirements, so that we can build a program that is tailored to your interests and strengths.
Department Policies
Transfer Credit
Undergraduate students who wish to transfer in credit for courses taken at another institution should familiarize themselves with both the university policies and their home college’s transfer credit policies. Statistics and Data Science students can find the Dietrich college policies here.
We encourage students considering taking courses outside of CMU to contact their academic advisor to ensure proper planning and guidance. Students must complete this form to request a course to transfer in as any undergraduate course in the Department of Statistics and Data Science (36-000 - 36-599).
Statistics and Data Science students who wish to transfer coursework from another institution for core Statistics and Data Science credit must adhere to the following guidelines:
-
At most four (4) core courses may be transferred in and used to fulfill degree requirements. This applies to primary and additional major students. For the minor, this is reduced to two (2) courses.
-
Advanced placement or equivalent and courses taken prior to matriculation to Carnegie Mellon University do not factor into the 4 course maximum.
-
Exceptions may be sought through consultation with an academic advisor and the undergraduate program director(s).
-
-
The following courses must be taken in residence at Carnegie Mellon University and are not eligible for transfer:
Substitutions and Waivers
Statistics and Data Science department majors and minors seeking substitutions or waivers must consult their academic advisor and complete the appropriate form, provided by the advisor, for any waiver to be considered. The department does not offer course waiver assessments and will not waive prerequisites for prior knowledge unless approved by an outside department.
Additionally, the department does not provide approval or permission for substitutions or waivers for majors outside of the Department of Statistics and Data Science. If a student wants to replace a Statistics & Data Science course that is a requirement in their primary major, and the substitute course has not been approved as an equivalent through the required form noted in the “Transfer Credit” section above, students must consult their home department primary advisor.
Dietrich Honors Thesis - Department Requirements
These guidelines apply to students who are completing an honors thesis that has been approved through the Statistics and Data Science department (i.e. our department signs off on the thesis paperwork and the work is completed with a department faculty member). If you are a Statistics and Data Science student pursuing a Dietrich senior honors thesis through another department (i.e. a different department is signing off on it) then these guidelines do not apply.
Eligibility is determined by Dietrich College. Students who are eligible will be notified prior to their senior year. Dietrich College Requirements can be found here. In order to be approved for a thesis with the Statistics and Data Science department the project needs to have a significant statistical component.
The Dietrich College senior honors thesis is a year-long project. As such, after the fall semester of a student’s senior year a progress report will be due to the Undergraduate Program Director for review by the last day of classes for the fall semester. The progress report should build substantially on the proposal, and lay out what work has been completed as well as an action plan for the spring semester. The progress report should be a minimum of five pages in length.
In alignment with a typical advanced data analysis (ADA) project in the field of Statistics the minimum required length of the final thesis must be fifteen (15) written pages, no more than 18 single-spaced pages, 12-point font. This does not include figures. Further requirements are outlined on the department website.
All theses are due to the Undergraduate Program Director and Department Head at the end of the twelfth (12th) week of classes in spring semester (typically the first week of April).
Research and Graduate School
Research as an Undergraduate
The Statistics and Data Science program encourages students to gain research experience. Opportunities within the department include Summer Undergraduate Research Apprenticeships (SURA), run in association with the university's Office of Undergraduate Research and Scholar Development, and the departmental capstone courses 36-490 Undergraduate Research, 36-493 Sports Analytics Capstone, or 36-497 Corporate Capstone Project (note that these courses require an application). Additionally, students can pursue independent study. For those students who maintain a quality point average of 3.25 overall or above, there is also the Dietrich College Senior Honors Program.
The faculty in the Statistics and Data Science department largely work within the domains of statistical theory and methodological development, areas that require advanced mathematical training. Thus we encourage students to search broadly for research opportunities: faculty, post-doctoral researchers, and graduate students in many departments throughout the university have data to analyze and would welcome the help of undergraduate Statistics and Data Science students.
Recommendations for Prospective Ph.D. Students
Students interested in pursuing a Ph.D. in Statistics or related programs after completing their undergraduate degree are strongly recommended to pursue the B.S. in Statistics and Data Science (Mathematical Sciences Track) or to take additional Mathematics courses. Although 21-240 Matrix Algebra with Applications is recommended for Statistics and Data Science majors, students interested in PhD programs should consider taking 21-241 Matrices and Linear Transformations or 21-242 Matrix Theory instead. Additional courses to consider are 21-228 Discrete Mathematics, 21-341 Linear Algebra, 21-355 Principles of Real Analysis I, and 21-356 Principles of Real Analysis II. We also recommend that students interested in pursuing a Ph.D. gain some research experience during their undergraduate degree, as discussed further in the Research section above. Internships that involve meaningful real data analysis are also looked upon favorably in PhD programs.
Course Descriptions
About Course Numbers:
Each Carnegie Mellon course number begins with a two-digit prefix that designates the department offering the course (i.e., 76-xxx courses are offered by the Department of English). Although each department maintains its own course numbering practices, typically, the first digit after the prefix indicates the class level: xx-1xx courses are freshmen-level, xx-2xx courses are sophomore level, etc. Depending on the department, xx-6xx courses may be either undergraduate senior-level or graduate-level, and xx-7xx courses and higher are graduate-level. Consult the Schedule of Classes each semester for course offerings and for any necessary pre-requisites or co-requisites.
- 36-198 Research Training: Writing in Statistics
- Intermittent
TBD
Prerequisite: 36-200
- 36-200 Reasoning with Data
- All Semesters: 9 units
This course is an introduction to learning how to make statistical decisions and now to reason with data. The approach will emphasize the thinking-through of empirical problems from beginning to end and using statistical tools to look for evidence for/against explicit arguments/hypotheses. Types of data will include continuous and categorical variables, images, text, networks, and repeated measures over time. Applications will largely drawn from interdisciplinary case studies spanning the humanities, social sciences, and related fields. Methodological topics will include basic exploratory data analysis, elementary probability, significance tests, and empirical research methods. There will be once-weekly computer lab for additional hands-on practice using an interactive software platform that allows student-driven inquiry.
- 36-202 Methods for Statistics & Data Science
- All Semesters: 9 units
This course builds on the principles and methods of statistical reasoning that were developed in a first-semester intro statistics course, and will cover regression analysis (simple and multiple), logistic regression, one-way analysis of variance, and some elementary machine learning topics. The course will revisit in more detail the methods for examining the relationship between two variables and will also expand the methods to cases where there is more than one explanatory variable. The course includes an introduction to the R studio statistical programming environment (through labs and homeworks). Skills in course methodology, concepts, and data scientific verbal expression will be synthesized in two written data analysis projects. Students produce and engage with an array of representations in R of descriptions, displays, models, and predictions, of real data. In each lab as well as on all homeworks and projects, data is explored, described, and modeled using R. Coding skills relevant to the course are developed from the ground up. In addition to relevant statistical and machine learning applications, students practice some elementary data cleaning, learn and experience the importance of reproducibility, and are introduced to the importance of commenting code. Issues of model selection, and the danger and issues of overfitting, are covered in a variety of contexts. Issues involving algorithmic bias and data ethics are also discussed.
Prerequisites: 36-247 or 36-207 or 70-207 or 36-200 or 36-220
- 36-204 Discovering the Data Universe
- Intermittent: 3 units
Every day we wake up in the data universe, we use the information around us to make decisions. We are constantly evaluating and interpreting data from our environment, in everything from spreadsheets to Instagram posts. At the same time, our own personal data are being observed and recorded and #8212;through websites we visit online, our smart devices, and even our interactions with other students and faculty at CMU. Navigating this data universe requires knowledge of what data is and how to use it responsibly. For example, can a plant be a data set? Discovering the truth behind a piece of data, including who made it, what it looks like, and what we can learn from it, is a critical skill. Understanding data can be the difference between being able to distinguish truth from lies; and the key to identifying your data footprint and succeeding in research and in your career. In this course, we will explore the data universe from multiple angles and across several types of data. We will define, find, and analyze data, and most importantly, identify narratives within data to tell stories about the world around us. We will examine data using the following questions: How can we tell multiple stories from the same dataset? What biases can exist in data? And, who creates or decides what data matters enough to collect, preserve, and share? NOTE: There will be one in person and one virtual pre-recorded lecture each week.
- 36-218 Probability Theory for Computer Scientists
- Fall and Spring: 9 units
Probability theory is the mathematical foundation for the study of both statistics and of random systems. This course is an intensive introduction to probability,from the foundations and mechanics to its application in statistical methods and modeling of random processes. Special topics and many examples are drawn from areas and problems that are of interest to computer scientists and that should prepare computer science students for the probabilistic and statistical ideas they encounter in downstream courses and research. A grade of C or better is required in order to use this course as a pre-requisite for 36-226, 36-326, and 36-410. If you hold a Statistics primary/additional major or minor you will be required to complete 36-226. For those who do not have a major or minor in Statistics, and receive at least a B in 36-218, you will be eligible to move directly onto 36-401.
Prerequisites: (21-112 and 21-111) or 21-120 or 21-256 or 21-259
Course Website: http://www.stat.cmu.edu/academics/courselist
- 36-219 Probability Theory and Random Processes
- Spring: 9 units
This course provides an introduction to probability theory. It is designed for students in electrical and computer engineering. Topics include elementary probability theory, conditional probability and independence, random variables, distribution functions, joint and conditional distributions, limit theorems, and an introduction to random processes. Some elementary ideas in spectral analysis and information theory will be given. A grade of C or better is required in order to use this course as a pre-requisite for 36-226 and 36-410.
Prerequisites: (21-111 and 21-112) or 21-120 or 21-256 or 21-259
- 36-220 Engineering Statistics and Quality Control
- Fall and Spring: 9 units
This is a course in introductory statistics for engineers with emphasis on modern product improvement techniques. Besides exploratory data analysis, basic probability, distribution theory and statistical inference, special topics include experimental design, regression, control charts and acceptance sampling.
Prerequisites: (21-120 or 21-112) and (02-120 or 15-110 or 15-112)
- 36-225 Introduction to Probability Theory
- All Semesters: 9 units
This course is the first half of a year-long course which provides an introduction to probability and mathematical statistics for students in the data sciences. Topics include elementary probability theory, conditional probability and independence, random variables, distribution functions, joint and conditional distributions, law of large numbers, and the central limit theorem. This course is open to those with sophomore standing or higher.
Prerequisites: (21-111 and 21-112) or 21-120 or 21-256 or 21-259
Course Website: http://coursecatalog.web.cmu.edu/schools-colleges/dietrichcollegeofhumanitiesandsocialsciences/depar
- 36-226 Introduction to Statistical Inference
- Spring and Summer: 9 units
This course is the second half of a year-long course in probability and mathematical statistics. Topics include estimation, confidence intervals, hypothesis testing, linear regression, and analysis of variance. A grade of C or better is required in order to advance to 36-401, 36-402 or any 36-46x course.
Prerequisites: 36-219 Min. grade C or 36-225 Min. grade C or 36-218 Min. grade C or 15-259 Min. grade C or 36-217 Min. grade C or 21-425 Min. grade C or 21-325 Min. grade C
- 36-235 Probability and Statistical Inference I
- Fall: 9 units
This class is the first half of a two-semester, calculus-based course sequence that introduces theoretical aspects of probability and statistical inference to students. The material in this course and in 36-236 (Probability and Statistical Inference II) is organized so as to provide repeated exposure to essential concepts: the courses cover specific probability distributions and their inferential applications one after another, starting with the normal distribution and continuing with the binomial and Poisson distributions, etc. Topics specifically covered in 36-235 include basic probability, random variables, univariate and multivariate distribution functions, point and interval estimation, hypothesis testing, and regression, with the discussion being supplemented with computer-based examples and exercises (e.g., visualization and simulation). Given its organization, the course is only appropriate for those taking the full two-semester sequence, and thus it is currently open only to statistics majors (primary, additional, dual) and minors. (Check with the statistics advisors for the exact declaration deadline.) Non-majors/minors requiring a probability course are directed to take 36-225 or one of its analogues. A grade of C or better in 36-235 is required in order to advance to 36-236 (or 36-226) and/or 36-410. This course is not open to students who have received credit for 36-217, 36-218, 36-219, or 36-700, or for 21-325 or 15-259. This course is open to those with sophomore standing or higher.
Prerequisites: (21-111 and 21-112) or 21-256 or 21-259 or 21-120
- 36-236 Probability and Statistical Inference II
- Spring: 9 units
This class is the second half of a two-semester, calculus-based course sequence that introduces theoretical aspects of probability and statistical inference to students. The material in this course and in 36-235 (Probability and Statistical Inference I) is organized so as to provide repeated exposure to essential concepts: the courses cover specific probability distributions and their inferential applications one after another, starting with the normal distribution and continuing with the binomial and Poisson distributions, etc. Topics specifically covered in 36-236 include the binomial and related distributions, the Poisson and related distributions, and the uniform distribution, and how they are used in point and interval estimation, hypothesis testing, and regression. Also covered in 36-236 are topics related to multivariate distributions: marginal and conditional distributions, covariance, and conditional distribution moments. All discussion is supplemented with computer-based examples and exercises (e.g., visualization and simulation). Given its organization, the course is only appropriate for those who first take 36-235, and thus it is currently open only to statistics majors (primary, additional, dual) and minors, as well as to CS majors using both 36-235 and 36-236 to complete their probability requirement. All others are directed to take 36-226. A grade of C or better in 36-236 is required in order to advance to 36-401.
Prerequisite: 36-235 Min. grade C
- 36-290 Introduction to Statistical Research Methodology
- Fall: 9 units
This is a first course in statistical practice, targeted to first-semester sophomores. It is designed as a high-level introduction to the ways by which statisticians go about approaching and analyzing quantitative observational data, thus preparing students for future work in capstone classes. Students in the course are taught the basic concepts of statistical learning and #8212;inference vs.prediction, supervised vs. unsupervised learning, regression vs. classification, etc. and #8212;and will reinforce this knowledge by applying, e.g., linear regression, random forest, principal components analysis, and/or hierarchical clustering and more to datasets provided by the instructor. Students will also practice disseminating the results of their analyses via oral presentations and posters. Analyses will be carried out using the R programming language.
Prerequisites: 36-200 or 36-247 or 36-220 or 36-207 or 70-207
Course Website: http://coursecatalog.web.cmu.edu/schools-colleges/dietrichcollegeofhumanitiesandsocialsciences/depar
- 36-294 Independent Study Mini
- All Semesters
Independent study mini course.
- 36-297 Early Undergraduate Research
- Fall and Spring: 6 units
This course is designed to give early undergraduate students (those who have not yet taken 36-401) experience navigating real data science research problems. Small groups of students are matched with clients and do supervised research for a semester. From an academic perspective, the course presents an opportunity for students to gain skills in, e.g., data acquisition and cleaning, exploratory data analysis, and basic statistical modeling; which skills are practiced is project-dependent. Additionally, the course will help students develop the professional skills necessary for successfully navigating team-based project delivery roles. Programming will be performed in R and/or Python; previous programming experience is not required.
- 36-300 Statistics & Data Science Internship
- All Semesters
The Department of Statistics and amp; Data Science considers experiential learning as an integral part of our program. One such option is through an internship. If a student has an internship, they dont have to register for this class unless they want it listed on their official transcripts. This process should be used by international students interested in Curricular Practical Training (CPT) and should also be authorized by the Office of International Education (OIE). More information regarding CPT is available on OIE's website. This course will be taken as Pass/Fail, and students will be charged tuition for 3 units. There is an approval process in order to register for this course. Please contact your advisor the Department of Statistics and amp; Data Science for more details.
- 36-301 Documenting Human Rights
- Intermittent: 9 units
This course will teach students about the origins of modern human rights and the evolution of methods to document the extent to which these rights are being upheld or violated. The need to understand and document human rights issues is at the center of the most pressing current events. From threats to democracy and civil rights to work holding perpetrators of mass harm accountable in legal proceedings to efforts to quantify and advance economic, social, cultural, and environmental rights, making human rights violations visible is fundamental to achieving a more just world. We will begin with an overview of the history of human rights, the main philosophical and political debates in the field, and the most relevant organizations, institutions, and agreements. We will then delve into specific cases that highlight methodological opportunities and challenges, including: the identification of mass atrocity victims, the disappeared, and missing migrants; efforts to estimate civilian casualties in war; the documentation of police brutality and other human rights violations with smartphones; as well as the use of satellite imagery and drone footage for the documentation of genocide, environmental rights, and war crimes. We will critically assess the technical challenges that arise in each context and how the human rights and scientific communities have responded. After reviewing these cases, we will conclude by reflection on why the documentation of human rights actually matters and what happens to evidence once it is gathered. Students will then take what they've learned and do two multidisciplinary group projects, one involving the document of a rights violation in Western Pennsylvania and the other involving an international situation. Assignments include an essay, a data analysis assignment, and a group project that include a written component, quantitative and/or qualitative data analysis, and a presentation.
- 36-303 Sampling, Survey and Society
- Spring: 9 units
This course will revolve around the role of sampling and sample surveys in the context of U.S. society and its institutions. We will examine the evolution of survey taking in the United States in the context of its economic, social and political uses. This will eventually lead to discussions about the accuracy and relevance of survey responses, especially in light of various kinds of nonsampling error. Students will be required to design, implement and analyze a survey sample.
Prerequisites: 36-200 or 36-220 or 70-207
- 36-309 Experimental Design for Behavioral & Social Sciences
- Fall and Summer: 9 units
This course focuses on the statistical aspects of the design and analysis stages of planned experiments. The design stage focuses on determining how experimental factors are allocated, the sample size necessary to achieve adequate statistical power, and how subjects/variables are measured. The analysis stage focuses on how data are collected and which statistical models are most appropriate to answer the research questions of interest. Although students will have to do some computer programming to implement these statistical techniques, the most important aspect of the course will be on interpreting analyses' results (e.g., whether a given analysis is appropriate, to what extent that analysis can answer research questions of interest, and the broader implications of an analysis within the context of the experiment). In addition to a weekly lecture, students will attend a computer lab once a week to get guidance and hands-on practice implementing statistical techniques we learn in class.
Prerequisites: 36-220 or 36-218 or 70-207 or 36-247 or 15-260 or 36-236 or 36-226 or 36-200 or 36-326
Course Website: http://www.stat.cmu.edu/academics/courselist
- 36-311 Statistical Analysis of Networks
- Intermittent: 9 units
Networks are omnipresent in modern data science,arising in social systems, biology, technology, and beyond. In this course,students will get an introduction to network science, mainly focusing on social and biological networks. We will begin with some empirical examples, an overview of concepts for measuring and describing networks, and discuss network visualization. Traditional statistical models are often inadequate for network data due to complex dependence structure. We will introduce random graph models and statistical network models developed to characterize network structure and growth. We will also cover methods for statistical inference and prediction on network-linked data and, if time permits, special topics such as dynamic networks and multilayer networks.
Prerequisites: 36-226 or 36-236
- 36-313 Statistics of Inequality and Discrimination
- Intermittent: 9 units
Many social questions about inequality, injustice and unfairness are, in part, questions about evidence, data, and statistics. This class lays out the statistical methods which let us answer questions like "Does this employer discriminate against members of that group?", "Is this standardized test biased against that group?", "Is this decision-making algorithm biased, and what does that even mean?" and "Did this policy which was supposed to reduce this inequality actually help?" We will also look at inequality within groups, and at different ideas about how to explain inequalities between groups. The class will interweave discussion of concrete social issues with the relevant statistical concepts.
Prerequisite: 36-202
- 36-315 Statistical Graphics and Visualization
- All Semesters: 9 units
Graphical displays of quantitative information take on many forms as they help us understand both data and models. This course will serve to introduce the student to the most common forms of graphical displays and their uses and misuses. Students will learn both how to create these displays and how to understand them. Each student will be required to engage in a project using graphical methods to understand data collected from a real scientific or engineering experiment. In addition to two weekly lectures there will be lab sessions where the students learn to use software to aid in the production of appropriate graphical displays.
Prerequisites: 36-220 or 70-207 or 36-200
- 36-318 Introduction to Causal Inference
- Intermittent: 9 units
Many social science and scientific inquiries can be framed as causal questions. Does a new cancer treatment cause a reduction in mortality? Do financial grants cause students to do better in college? Does a new public policy cause an increase in voter turnout? When tackling these questions, we frequently come across the phrase "correlation does not imply causation." If that's the case, then what does imply causation? In this course, we will discuss causal inference methods for measuring causal effects of different interventions (e.g., drug treatments, financial grants, and public policies). First, we will discuss how experiments and #8212;-where interventions are randomized among subjects and #8212;-can imply causation when an appropriate experimental design and statistical analysis is used. Then, we will discuss how observational studies and #8212;-where interventions are not randomized and #8212;-can also imply causation when approaches like propensity score methods, matching, and doubly robust estimation are employed. Finally, we will discuss instrumental variables and regression discontinuity designs and #8212;-which are frequently used in medicine and public policy for establishing causal inferences. Throughout we will use R to conduct causal analyses. A working knowledge of regression is encouraged, but regression will also be discussed and taught during much of the course.
Prerequisites: 36-218 Min. grade C or 36-235 Min. grade C or 15-259 Min. grade C or 21-325 Min. grade C or 36-219 Min. grade C or 21-425 Min. grade C or 36-225 Min. grade C
- 36-319 Statistics and Machine Learning for the Physical Sciences
- All Semesters: 9 units
In this course, we will learn about new methodological research at the intersection of statistics, machine learning, and AI with applications in the physical sciences (including meteorology, remote sensing, astronomy, and particle physics). Students will see examples of how core statistical ideas (confidence sets, hypothesis testing, two-sample testing, and uncertainty quantification) play a key role in cutting-edge research in predictive inference, probabilistic forecasting, anomaly or signal detection, and likelihood-free (aka simulator-based) inference. Homework assignments will include assigned reading, theory problems, and some simple computer exercises. Students will work with real-world data during TA-led computer labs using Python/PyTorch and Jupyter notebooks to apply the skills and knowledge acquired throughout the program. (Prior exposure to Python programming is not required.)
Prerequisites: 36-235 or 36-219 or 36-218 or 15-259 or 36-220 or 36-225
- 36-320 Statistics & Data Science Internship
- Fall and Spring
TBD.
- 36-326 Mathematical Statistics (Honors)
- Spring: 9 units
This course is a rigorous introduction to the mathematical theory of statistics. A good working knowledge of calculus and probability theory is required. Topics include maximum likelihood estimation, confidence intervals, hypothesis testing, Bayesian methods, and regression. A grade of C or better is required in order to advance to 36-401, 36-402 or any 36-46x course. Not open to students who have received credit for 36-625. Prerequisites: 15-359 or 21-325 or 36-217 or 36-225 with a grade of A AND advisor approval. Students interested in the course should add themselves to the waitlist pending review.
Prerequisites: 36-218 Min. grade A or 36-217 Min. grade A or 36-225 Min. grade A or 15-359 Min. grade A or 21-325 Min. grade A
- 36-350 Statistical Computing
- All Semesters: 9 units
Statistical Computing: an introduction to computing targeted at statistics majors with (perhaps) minimal programming knowledge. The main topics are core ideas of programming (functions, objects, data structures, flow control, input and output, debugging, logical design and abstraction). The class will be taught primarily in the R language, with Python also being utilized beginning in Spring 2025. No previous programming experience in R is required but exposure to it in previous classes is assumed; students who have not been exposed to Python in the courses 15-110/15-112 or their equivalents are advised to work through, e.g., the free CMU Open Learning Initiative (OLI) Python course prior to the first week of class.
Prerequisites: (36-225 Min. grade C or 21-325 Min. grade C or 36-218 Min. grade C or 36-217 Min. grade C or 36-235 Min. grade C or 15-259 Min. grade C or 36-219 Min. grade C) and (15-112 or 15-110 or 02-120)
- 36-390 Study Abroad Experience in Statistics and Data Science
- Summer: 9 units
Statistics and Data Science at the Monteverde Institute in Costa Rica. This is a five-week study abroad experience in which students will directly engage with, and will process, visualize, and/or analyze data collected by, researchers at the institute. Students will also have the opportunity to participate in data collection, as appropriate. The mission of the institute is to promote sustainable practices that benefit both the local community and local wildlife, and the data that students can examine include, but are not limited to, ecological data on bats, birds, reforestation, and stream beds, as well as data arising from community surveys. This course does not require prior knowledge of, or exposure to, data processing, visualization, or analysis techniques beyond what is covered in the prerequisite classes, and necessary techniques and methods will be introduced and discussed in daily classes. Project goals will be modified for students with more advanced backgrounds (e.g., students who have completed 36-401 and 36-402). The 2024 class is limited to six students overall.
- 36-394 Independent Study Mini
- All Semesters
Independent study.
- 36-396 Tartan Athletics Analytics
- Intermittent: 9 units
The Tartan Athletics Analytics course gives students hands-on experience applying statistics and data science methodologies to real-world datasets, and communicating results to stakeholders in the Carnegie Mellon Athletics Department (aka Tartan Athletics). Students will gain skills in approaching real world problems, critical thinking, advanced statistical analysis, scientific writing, collaboration with clients, communicating results, and meeting expectations with respect to deliverables and timelines. The projects will change and rotate each semester. The course size is limited, and students with skill sets and interests matching expected projects will be given priority. We will also take into consideration whether or not a student has had a recent prior data science experience with the goal of providing experiences to a broad group of qualified students. Students do not need to be experts in sports analytics or have extensive knowledge in sports to be considered.
- 36-400 Overview of Statistical Learning and Modeling
- Fall: 9 units
This course is a high-level introduction both to fundamental concepts of probability and statistics and to the ways by which statisticians go about approaching and analyzing data. The course will cover data processing, exploratory data analysis, parameter estimation and hypothesis testing, clustering, and common regression and classification models. Students will carry out work using the R and Python programming languages. This course is open only to students not majoring in Stat and amp; DS who have taken the prerequisite courses.
Prerequisites: 36-200 and (36-290 or 36-309 or 36-202)
- 36-401 Modern Regression
- Fall: 9 units
This course is an introduction to the real world of statistics and data analysis. We will explore real data sets, examine various models for the data, assess the validity of their assumptions, and determine which conclusions we can make (if any). Data analysis is a bit of an art; there may be several valid approaches. We will strongly emphasize the importance of critical thinking about the data and the question of interest. Our overall goal is to use a basic set of modeling tools to explore and analyze data and to present the results in a scientific report. A grade of C is required to move on to 36-402 or any 36-46x course. *The spring course is primarily for SCS students.*
Prerequisites: (36-326 Min. grade C or 36-218 Min. grade B or 15-259 Min. grade B or 36-236 Min. grade C or 36-226 Min. grade C) and (21-240 or 21-242 or 21-241 or 18-202)
- 36-402 Advanced Methods for Data Analysis
- Spring: 9 units
This course introduces modern methods of data analysis, building on the theory and application of linear models from 36-401. Topics include nonlinear regression, nonparametric smoothing, density estimation, generalized linear and generalized additive models, simulation and predictive model-checking, cross-validation, bootstrap uncertainty estimation, multivariate methods including factor analysis and mixture models, and graphical models and causal inference. Students will analyze real-world data from a range of fields, coding small programs and writing reports.
Prerequisite: 36-401 Min. grade C
- 36-410 Introduction to Probability Modeling
- Spring: 9 units
An introductory-level course in stochastic processes. Topics typically include Poisson processes, Markov chains, birth and death processes, random walks, recurrent events, and renewal theory. Examples are drawn from reliability theory, queuing theory, inventory theory, and various applications in the social and physical sciences.
Prerequisites: 36-218 or 36-235 or 15-259 or 21-325 or 36-225 or 21-425 or 36-217 or 36-219
- 36-460 Special Topics: Sports Analytics
- Spring: 9 units
This course introduces students to fundamental topics in sports analytics and the relevant statistical methods for tackling problems in this growing area. The first half of the course will cover foundational topics in sports analytics including building models for the expected value of game states and multilevel modeling for player and team evaluation. The second half of the course focuses on Bayesian thinking with hierarchical models to estimate and quantify the uncertainty around player / team ratings across multiple sports, including static and dynamic techniques. Remaining time of the course will introduce students to working with complex player-tracking data and relevant spatio-temporal methods. All methods in the course are motivated by real sports problems that a statistician / data scientist working in sports analytics encounters. The focus is on understanding the foundations of the considered methods and introducing software for implementation. Students will develop their own sports analytics project using techniques covered in the course for their final assessment.
Prerequisite: 36-401 Min. grade C
- 36-461 Special Topics: Statistical Methods in Epidemiology
- Intermittent: 9 units
Epidemiology is concerned with understanding factors that cause, prevent, and reduce diseases by studying associations between disease outcomes and their suspected determinants in human populations. Epidemiologic research requires an understanding of statistical methods and design. Epidemiologic data is typically discrete, i.e., data that arise whenever counts are made instead of measurements. In this course, methods for the analysis of categorical data are discussed with the purpose of learning how to apply them to data. The central statistical themes are building models, assessing fit and interpreting results. There is a special emphasis on generating and evaluating evidence from observational studies. Case studies and examples will be primarily from the public health sciences.
Prerequisites: (36-226 Min. grade C or 36-326 Min. grade C or 36-236 Min. grade C) and (21-241 or 21-242 or 21-240)
Course Website: http://coursecatalog.web.cmu.edu/schools-colleges/dietrichcollegeofhumanitiesandsocialsciences/depar
- 36-462 Special Topics: Statistical Machine Learning
- Intermittent: 9 units
Data mining is the science of discovering patterns and learning structure in large data sets. Covered topics include clustering, dimension reduction, regression, classification, and decision trees.
Prerequisite: 36-401 Min. grade C
Course Website: http://www.stat.cmu.edu/academics/courselist
- 36-463 Special Topics: Multilevel and Hierarchical Models
- Intermittent: 9 units
Multilevel and hierarchical models are among the most broadly applied "sophisticated" statistical models, especially in the social and biological sciences. They apply to situations in which the data "cluster" naturally into groups of units that are more related to each other than they are the rest of the data. In the first part of the course we will review linear and generalized linear models. In the second part we will see how to generalize these to multilevel and hierarchical models and relate them to other areas of statistics, and in the third part of the course we will learn how Bayesian statistical methods can help us to build, estimate and diagnose problems with these models using a variety of data sets and examples.
Prerequisite: 36-401 Min. grade C
Course Website: http://www.stat.cmu.edu/academics/courselist
- 36-464 Special Topics: Psychometrics: A Statistical Modeling Approach
- Intermittent: 9 units
Much of the social, educational, policy, and professional worlds involve measuring the skills, abilities, attitudes, decision-making, etc. of people and #8212; from SAT's and GRE's for school, to 360-evaluations in business. This is the field of modern psychometrics, and it involves (at least) two kinds of craft: designing good sets of questions, and designing and fitting statistical models that extract the information we want from the responses to those questions. In this course we will touch on both kinds of craft, but we will concentrate on the second: what do statistical models for psychometric data look like, and how can we design, fit, and use them in practice? We will look at these models from a variety of statistical perspectives, but we will concentrate on the applied Bayesian point of view.
Prerequisite: 36-401 Min. grade C
Course Website: http://www.stat.cmu.edu/academics/courselist
- 36-465 Special Topics: Conceptual Foundations of Statistical Learning
- Intermittent: 9 units
This class is an introduction to the foundations of statistical learning theory, and its uses in designing and analyzing machine-learning systems. Statistical learning theory studies how to fit predictive models to training data, usually by solving an optimization problem, in such a way that the model will predict well, on average, on new data. The course will focus on the key concepts and theoretical tools, at a level accessible to students who have taken 36-401 and its pre-requisites. The course will also illustrate those concepts and tools by applying them to carefully selected kinds of machine learning systems (such as kernel machines). Students wanting exposure to a broad range of algorithms and applications would be better served by 36-462/662 ("Data Mining"). This class is for those who want a deeper understanding of the principles underlying all machine learning methods.
Prerequisite: 36-401 Min. grade C
- 36-466 Special Topics: Statistical Methods in Finance
- Intermittent: 9 units
Statistical methods are fundamental to modern finance. As financial data grow in volume and complexity, statistical models are crucial to quantifying the risks and potential rewards of various financial products, for both those buying and selling. This course introduces core statistical techniques used in finance while reinforcing key concepts from prior statistics coursework. Topics will include, but are not limited to, the following: model calibration to historical data, uncertainty quantification, benefits and limitations of geometric Brownian motion, Poisson process and jump models for asset prices, interest rate and yield curve models, CAPM and factor models, portfolio optimization, and financial time series models. Students will also learn data exploration tools essential for quantitative analysis. Python will be used throughout the course since it is the standard package used in finance. No prior experience with Python is required.
Prerequisites: (36-236 Min. grade C or 36-226 Min. grade C or 36-326 Min. grade C) and (21-242 or 21-240 or 21-241)
- 36-467 Special Topics: Data over Space & Time
- Intermittent: 9 units
This course is an introduction to the opportunities and challenges of analyzing data from processes unfolding over space and time. It will cover basic descriptive statistics for spatial and temporal patterns; linear methods for interpolating, extrapolating, and smoothing spatio-temporal data; basic nonlinear modeling; and statistical inference with dependent observations. Class work will combine practical exercises in R, a little mathematics on the underlying theory, and case studies analyzing real problems from various fields (economics, history, meteorology, ecology, etc.). Depending on available time and class interest, additional topics may include: statistics of Markov and hidden-Markov (state-space) models; statistics of point processes; simulation and simulation-based inference; agent-based modeling; dynamical systems theory.
Prerequisite: 36-401 Min. grade C
Course Website: http://coursecatalog.web.cmu.edu/schools-colleges/dietrichcollegeofhumanitiesandsocialsciences/depar
- 36-468 Special Topics: Text Analysis
- Intermittent: 9 units
The analysis of language is concerned with how variables relate to people (their gender, age, and location, for example), how variables relate to use (such as writing in different academic disciplines), and how variables change over time. While we are surrounded by data that might potentially shed light on many of these questions, working with real-world linguistic data can present some unique challenges in sampling, in the distribution of features, and in their high dimensionality. In this course, we work through some of these issues, paying particular attention to the aligning of the statistical questions we want to investigate with the choice of statistical models, as well as focusing on the interpretation of results. Analysis will be carried out in R and students will develop a suite of tools as they work through their course projects.
Prerequisites: (36-326 Min. grade C or 36-236 Min. grade C or 36-226 Min. grade C) and (21-241 or 21-242 or 21-240)
- 36-469 Special Topics: Statistical Genomics and High Dimensional Inference
- Intermittent: 9 units
The field of computational and statistical genomics focuses on developing and applying computationally efficient and statistically robust methods to sort through increasingly rich and massive genome wide data sets to identify complex genetic patterns, gene interactions, and disease associations. Because the genome is vast, analytical approaches require high dimensional statistical approaches such as multiple testing, dimension reduction techniques, regularization and high dimensional regression analysis, best linear unbiased prediction models, networks and deep learning. In this course, we will motivate these topics using data obtained from the human genetic and genomic literature. No prior knowledge in biology is required.
Prerequisite: 36-401 Min. grade C
- 36-470 Special Topics: Statistical Methods in Health Sciences
- Intermittent: 9 units
As the volume of health and clinical data continues to expand, the integration of statistical and machine learning methods becomes increasingly important for enhancing healthcare efficiency. However, there are challenges in modeling health data, for example, annotated data is often limited or subject to incompleteness. In this course, we will introduce statistical methods that address these challenges, including survival analysis, latent variable models, clustering, semi-supervised learning, and so on. An emphasis will put on understanding methodological foundations and how to appropriately apply methods to health data. Through homework assignments, labs, paper presentations, and a final project, students will gain hands-on-experience in applying statistical methods to solve problems arising from health sciences.
Prerequisite: 36-401 Min. grade C
- 36-471 Special Topics: Time Series
- Fall: 9 units
This course covers time series analysis from fundamentals to advanced models in both time and frequency domains. The focus is on practical execution and interpretation of time series analyses with realistic real-world data.
Prerequisite: 36-401
- 36-472 Special Topics: Computational Statistical Methods in Life Sciences
- Intermittent: 9 units
The life sciences are rapidly becoming data-driven fields due to technological advancements in genomics, neuroscience, epidemiology, and other areas. Statistical methods are essential for analyzing complex datasets that arise in these disciplines, such as genomic sequences, neuroimaging data, and public health records. This course will introduce statistical techniques for analyzing data-driven life science problems, emphasizing computational aspects from a Bayesian perspective.
Prerequisite: 36-401 Min. grade C
- 36-473 Special Topics: Statistical Principles of Generative AI
- Intermittent: 9 units
Generative artificial intelligence systems are based on statistical models of text and images. The systems are very new, but they rely on well-established ideas about modeling and inference, some of them from the early 1900s. This course will introduce students to the statistical basis of large language models and image generation models, emphasizing high-level principles over implementation details. It will also examine whether, in light of those statistical foundations, we should think of generative AI as a step towards "artificial general intelligence", or instead as a "cultural technology".
Prerequisite: 36-402
- 36-474 Special Topics: Statistical Thinking for the AI Era
- Intermittent: 9 units
TBD - New spring 2027
- 36-490 Undergraduate Research
- Fall and Spring: 9 units
This course is designed to give undergraduate students experience using statistics in real research problems. Small groups of students will be matched with clients and do supervised research for a semester. Students will gain skills in approaching a research problem, critical thinking, statistical analysis, scientific writing, and conveying and defending their results to an audience.
- 36-493 Sports Analytics Capstone
- Intermittent: 9 units
This course is designed to give undergraduate students experience applying statistics and amp; data science methodology to research problems in sports analytics. Small groups of students will be matched with clients in the Carnegie Mellon Athletics Department and do supervised projects for a semester. Students will gain skills in approaching a real world problem, critical thinking, advanced statistical analysis, scientific writing, collaboration with clients, communicating results, and meeting expectations with respect to deliverables and timelines. The projects will change and rotate each semester. The course size is limited, and students will submit an application including their project preferences. Students with skill sets matching project needs will be given priority. We will also take into consideration whether or not a student has had a recent prior data science experience with the goal of providing experiences to a broad group of qualified students. Students do not need to be experts in sports analytics or have extensive knowledge in sports.
- 36-497 Corporate Capstone Project
- Fall and Spring: 9 units
This course is designed to give undergraduate students experience applying statistics and amp; data science methodology to real industry projects. Small groups of students will be matched with industry clients and do supervised projects for a semester. Students will gain skills in approaching a real world problem, critical thinking, advanced statistical analysis, scientific writing, collaborating in an industry setting, communicating results, and meeting expectations with respect to deliverables and timelines. The industry clients will change and rotate each semester; available projects will be advertised prior to registration. The course size is limited, and students will submit an application including their project preferences. Students with skill sets matching project needs will be given priority. We will also take into consideration whether or not a student has had a recent prior corporate capstone experience with the goal of providing experiences to a broad group of qualified students.
- 36-498 Corporate Capstone II
- Fall and Spring
This course allows students to continue work on projects begun as part of 36-497, Corporate Capstone Project. Enrollment is at the discretion of the external advisor for the 36-497 project and the Department of Statistics and amp; Data Science.
- 36-620 Advanced Methods I
- Intermittent: 6 units
TBD
- 36-621 Advanced Methods II
- Intermittent: 6 units
TBD
- 36-631 Foundations of Causal Inference
- Intermittent: 6 units
This course will provide an introduction to the fundamentals of causal inference. Causal inference is concerned with whether and how one can go beyond statistical associations to draw causal conclusions from observational data. Topics will include: counterfactuals (potential outcomes and graphs), identification and estimation of average treatment effects in experiments and observational studies, nonparametric bounds, sensitivity analysis, instrumental variables, effect modification, and longitudinal studies.
- 36-632 Modern Causal Inference
- Intermittent: 6 units
This course will provide an in-depth look at modern causal inference. Topics will include: optimal treatment regimes, mediation, principal stratification, stochastic interventions, accounting for complex confounding and exposures, and methods for efficient nonparametric estimation. Some background in mathematical statistics is advised.
- 36-700 Probability and Mathematical Statistics
- Fall: 12 units
This is a one-semester course covering the basics of statistics. We will first provide a quick introduction to probability theory, and then cover fundamental topics in mathematical statistics such as point estimation, hypothesis testing, asymptotic theory, and Bayesian inference. If time permits, we will also cover more advanced and useful topics including nonparametric inference, regression and classification. Prerequisites: one- and two-variable calculus and matrix algebra. Graduate students in degree-seeking programs are given priority.
- 36-711 High Dimensional Probability and Applications
- All Semesters: 6 units
In this course, we will introduce non-asymptotic methods in high-dimensional probability that find common use in applications across statistics, computer science, data science, and engineering. Topics include tail bounds for i.i.d. sums and martingale differences, concentration inequalities for non-linear functions, matrix concentration, and suprema of stochastic processes.
- 36-712 Hypothesis Testing: A Minimax Perspective
- All Semesters: 6 units
Hypothesis testing is a fundamental task in statistics: given data, we want to decide which of two competing explanations (hypotheses) generated it. As statistical models become higher-dimensional and more complex, determining whether reliable testing is possible (and designing procedures that achieve optimal performance) becomes increasingly challenging. This course will be an introduction to minimax hypothesis testing, which aims to characterize the smallest separation between hypotheses that permits reliable testing. Our primary focus will be on the interplay between optimal tests and matching information-theoretic lower bounds, as well as on understanding when testing is easier than learning the underlying distribution. This course will be theory-oriented.
Prerequisites: (10-601 or 15-781) and (36-705 or 36-725)
- 36-738 Statistical Optimal Transport I
- All Semesters: 6 units
"This course will explore some of the contact points between statistics and optimal transport (OT). The OT framework gives rise to a useful set of ideas and methods which have found numerous applications across machine learning and statistics. Our primary focus will be on understanding how well we can estimate various objects in the OT framework (Wasserstein distances, OT maps, entropic OT) in a statistical minimax setup. Our secondary goal will be to study ideas at the intersection of high-dimensional probability and OT (concentration inequalities, gradient flows, sampling). Along the way we will introduce many ideas from convex analysis, non-parametric statistics and minimax theory."
- 36-739 Statistical Optimal Transport II
- All Semesters: 6 units
"This course will explore some of the contact points between statistics and optimal transport (OT). The OT framework gives rise to a useful set of ideas and methods which have found numerous applications across machine learning and statistics. Our primary focus will be on understanding how well we can estimate various objects in the OT framework (Wasserstein distances, OT maps, entropic OT) in a statistical minimax setup. Our secondary goal will be to study ideas at the intersection of high-dimensional probability and OT (concentration inequalities, gradient flows, sampling). Along the way we will introduce many ideas from convex analysis, non-parametric statistics and minimax theory."
- 36-768 Advanced Theory I
- Intermittent: 6 units
TBD
Faculty
SIVARAMAN BALAKRISHNAN, Professor – Ph.D., Carnegie Mellon University; Carnegie Mellon, 2015–
ELI BEN-MICHAEL, Assistant Professor, Joint with Heinz College – Ph.D., University of California; Carnegie Mellon, 2022–
ZACHARY BRANSON, Associate Teaching Professor; Assistant Director for the Undergraduate Program – Ph.D., Harvard University; Carnegie Mellon, 2019–
DAVID CHOI, Associate Professor of Statistics and Information Systems – Ph.D., Stanford University; Carnegie Mellon, 2004–
PETER E. FREEMAN, Associate Teaching Professor; Director of the Undergraduate Program – Ph.D., University of Chicago; Carnegie Mellon, 2004–
CHRISTOPHER R. GENOVESE, Professor – Ph.D., University of California; Carnegie Mellon, 1994–
JOEL B. GREENHOUSE, Professor – Ph.D., University of Michigan; Carnegie Mellon, 1982–
AMELIA HAVILAND, E.J. Barone Professor of Health Systems Management – Ph.D., Carnegie Mellon University; Carnegie Mellon, 2003–
JIASHUN JIN, Professor – Ph.D., Stanford University; Carnegie Mellon, 2003–
ROBERT E. KASS, Maurice Falk Professor of Statistics & Computational Neuroscience – Ph.D., University of Chicago; Carnegie Mellon, 1981–
EDWARD KENNEDY, Associate Professor – Ph.D., University of Pennsylvania; Carnegie Mellon, 2016–
ARUN KUCHIBHOTLA, Associate Professor – Ph.D., University of Pennsylvania; Carnegie Mellon, 2020–
MIKAEL KUUSELA, Associate Professor – Ph.D., Ecole Polytechnique Federale de Lausanne; Carnegie Mellon, 2018–
ANN LEE, Professor, Co-Director of PhD Program – Ph.D., Brown University; Carnegie Mellon, 2005–
JING LEI, Professor – Ph.D., University of California; Carnegie Mellon, 2011–
GONZALO E. MENA, Assistant Professor – Ph.D., Columbia University; Carnegie Mellon, 2023–
DANIEL NAGIN, Teresa and H. John Heinz III Professor of Public Policy – Ph.D., Carnegie Mellon University; Carnegie Mellon, 1976–
NYNKE NIEZINK, Associate Professor – Ph.D., University of Groningen; Carnegie Mellon, 2017–
REBECCA NUGENT, Department Head Stephen E. and Joyce Fienberg Professor of Statistics & Data Science – Ph.D., University of Washington; Carnegie Mellon, 2006–
TAEYONG PARK, Assistant Teaching Professor, CMU-Qatar – Ph.D., Washington University in St. Louis; Carnegie Mellon, 2018–
ANKIT PENSIA, Assistant Professor – Ph.D., University of Wisconsin, Madison; Carnegie Mellon, 2023–
AADITYA RAMDAS, Assistant Professor – Ph.D., Carnegie Mellon University; Carnegie Mellon, 2018–
ALEX REINHART, Associate Teaching Professor – Ph.D., Carnegie Mellon University; Carnegie Mellon, 2018–
KATHRYN ROEDER, UPMC Professor of Statistics and Life Sciences – Ph.D., Pennsylvania State University; Carnegie Mellon, 1994–
CHAD M. SCHAFER, Professor – Ph.D., University of California, Berkeley; Carnegie Mellon, 2004–
TEDDY SEIDENFELD, Herbert A. Simon Professor of Philosophy and Statistics – Ph.D., Columbia University; Carnegie Mellon, 1985–
COSMA SHALIZI, Associate Professor – Ph.D., University of Wisconsin, Madison; Carnegie Mellon, 2005–
YANDI SHEN, Assistant Professor – Ph.D., University of Washington;
WEIJING TANG, Assistant Professor – Ph.D., University of Michigan; Carnegie Mellon, 2023–
WILL TOWNES, Assistant Professor – Ph.D., Harvard University; Carnegie Mellon, 2022–
VALERIE VENTURA, Professor, Co-Director of PhD Program – Ph.D., University of Oxford; Carnegie Mellon, 1997–
ISABELLA VERDINELLI, Professor in Residence – Ph.D., Carnegie Mellon University; Carnegie Mellon, 1991–
LARRY WASSERMAN, UPMC University Professor of Statistics – Ph.D., University of Toronto; Carnegie Mellon, 1988–
RON YURKO, Assistant Teaching Professor – Ph.D., Carnegie Mellon University; Carnegie Mellon, 2022–
Emeriti Faculty
GEORGE T. DUNCAN, Professor of Statistics and Public Policy – Ph.D., University of Minnesota; Carnegie Mellon, 1974–
WILLIAM F. EDDY, John C. Warner Professor of Statistics – Ph.D., Yale Universit; Carnegie Mellon, 1976–
BRIAN JUNKER, Professor – Ph.D., University of Illinois; Carnegie Mellon, 1990–
JOSEPH B. KADANE, Leonard J. Savage Professor of Statistics and Social Sciences – Ph.D., Stanford University; Carnegie Mellon, 1969–
JOHN P. LEHOCZKY, Thomas Lord Professor of Statistics – Ph.D., Stanford University; Carnegie Mellon, 1969–
MARK J. SCHERVISH, Professor – Ph.D., University of Illinois; Carnegie Mellon, 1979–
DALENE STANGL, Teaching Professor – Ph.D, Carnegie Mellon University; Carnegie Mellon, 2017–
Special Faculty
F. SPENCER KOERNER, Lecturer – Ph.D., Carnegie Mellon University; Carnegie Mellon, 2022–
JAMIE MCGOVERN, Director: Master’s in Applied Data Science Program – B.A., Rice University; Carnegie Mellon, 2020–
GORDON WEINBERG, Senior Lecturer – M.A., University of Pittsburgh; Carnegie Mellon, 2004–
Affiliated Faculty
ANTHONY BROCKWELL – Ph.D., Melbourne University; Carnegie Mellon, 1999–
PHILIPP BURCKHARDT – Ph.D., Carnegie Mellon University; Carnegie Mellon, 2022–
BERNIE DEVLIN – Ph.D., Pennsylvania State University; Carnegie Mellon, 1994–
SAM VENTURA – Ph.D., Carnegie Mellon University; Carnegie Mellon, 2015–
