Freshly Printed - allow 10 days lead
Couldn't load pickup availability
Probability and Statistics for Data Science
A self-contained introduction to probability and statistics for data science with examples involving real-world datasets.
Carlos Fernandez-Granda (Author)
9781009180085, Cambridge University Press
Hardback, published 3 July 2025
624 pages
25.4 x 17.8 x 3.3 cm, 1.426 kg
'If you're mathematically inclined and want to master the foundations of data science in one go, this book is for you. It covers a broad range of essential modern topics - including nonparametric methods, causal inference, latent variable models, Bayesian approaches, and a thorough introduction to machine learning - all illustrated with an abundance of figures and real-world data examples. Highly recommended.' David Rosenberg, Office of the CTO, Bloomberg
This self-contained guide introduces two pillars of data science, probability theory, and statistics, side by side, in order to illuminate the connections between statistical techniques and the probabilistic concepts they are based on. The topics covered in the book include random variables, nonparametric and parametric models, correlation, estimation of population parameters, hypothesis testing, principal component analysis, and both linear and nonlinear methods for regression and classification. Examples throughout the book draw from real-world datasets to demonstrate concepts in practice and confront readers with fundamental challenges in data science, such as overfitting, the curse of dimensionality, and causal inference. Code in Python reproducing these examples is available on the book's website, along with videos, slides, and solutions to exercises. This accessible book is ideal for undergraduate and graduate students, data science practitioners, and others interested in the theoretical concepts underlying data science methods.
Preface
Book Website
Introduction and Overview
1. Probability
2. Discrete variables
3. Continuous variables
4. Multiple discrete variables
5. Multiple continuous variables
6. Discrete and continuous variables
7. Averaging
8. Correlation
9. Estimation of population parameters
10. Hypothesis testing
11. Principal component analysis and low-rank models
12. Regression and classification
A. Datasets
References
Index.
Subject Areas: Probability & statistics [PBT]
