Books
Welcome to the books section of my portfolio. Here, I list the books I have studied to expand my professional skills.
Software Engineering for Data Scientists
by Catherine Nelson, O’Reilly, USA, published 2024
Skills: code quality, performance and data structures, object-oriented programming, error handling and logging, formatting and linting, documentation, testing, version control, APIs, automation and deployment;
Overview
This book bridges the gap between writing code in notebooks for exploratory
analysis and writing robust, maintainable, and reusable software. It covers
the practices that software engineers take for granted but data scientists
rarely learn formally, such as testing, documentation, refactoring, version
control, and deployment. These are exactly the practices I apply when
building my software packages.
Causal Inference in Statistics: A Primer
by Judea Pearl, Madelyn Glymour, and Nicholas P. Jewell, USA, published 2016
Skills: structural causal models, causal graphs, d-separation, do-operator, backdoor and front-door criteria, mediation analysis, causal inference in linear systems, counterfactuals;
Overview
This book is the technical companion to The Book of Why: a concise
introduction to the formal methods of causal inference, with exercises in
every chapter. It covers how to encode assumptions in causal graphs, how to
decide which variables to adjust for, how to compute the effects of
interventions from observational data, and how to reason about
counterfactuals. My particular focus was the sections on linear systems,
which show how these methods translate to linear models and regression
coefficients, the models I work with most in my research.
The Book of Why
by Judea Pearl and Dana Mackenzie, USA, published 2018
Skills: causal inference, causal diagrams, confounding, interventions, counterfactuals;
Overview
In this book, Turing Award winner Judea Pearl explains the “causal
revolution” for a broad audience: why classical statistics long avoided
questions of cause and effect, and how causal diagrams and the do-operator
now make it possible to answer them from data. It introduces the “ladder of
causation”, which distinguishes three levels of reasoning: association
(seeing), intervention (doing), and counterfactuals (imagining). It gave me
the conceptual big picture of causal inference, which the formal methods
then build on.
Modeling Mindsets
by Christoph Molnar, Germany, published 2022
Skills: Frequentist and Bayesian statistics, causal inference, supervised unsupervised reinforcement machine learning, deep learning;
Overview
This book discussed the idea of viewing modeling not in terms of the algorithms/model classes, like linear regression, but in terms of mindsets. Depending on the mindset, different conclusions can be drawn from the model and different outcomes achieved. For example, when linear regression is used with the Frequentist mindset, the focus is on hypothesis testing with the help of the parameter values, and the outcomes are p-values and confidence intervals. On the other hand, when the same algorithm is used with a supervised machine learning mindset, the focus is on the prediction performance on holdout test data.
Machine Learning Yearning
by Andrew Ng, USA, published 2018
Skills: machine learning, error analysis, project structuring, performance benchmarking, data strategy development;
Overview
This ebook is like a guidebook for planning and improving machine learning
projects. Instead of focusing on the math or code, it teaches you how to
think about building AI systems—like choosing the right data, figuring out
what’s going wrong, and deciding what to fix first. It’s about making smart
decisions to help your machine learning models work better in the real world.
Machine Learning Engineering
by Andriy Burkov, Canada, published 2020
Skills: machine learning, feature engineering, modeling, deployment and monitoring;
Overview
This ebook focuses on navigating machine learning projects, from best practices
of how to split the data, how to get great features, to choosing appropriate
models for the given problems, and more. This book gave me very valuable
information of what can be called the art of machine learning.
Introduction to Statistics
by David Lane, Rice University, USA, published 2003
Skills: distributions, probability, estimation, hypothesis testing, power, regression, transformations, chi square, effect size, research design, etc.;
Overview
I spent many hours deeply engaging with this book, which covers all fundamental
topics in statistics, from graphing and summarizing distributions to
probability, hypothesis testing, regression, and ANOVA. Working through each
chapter strengthened my intuition for statistical concepts, giving me a solid
foundation for data analysis, experimental design, and statistical inference.
Learn Data Mining Through Excel
by Hong Zhou, University of Saint Joseph, USA, published 2023
Skills: machine learning algorithms;
Overview
This ebook provided a hands-on approach to data mining by implementing
algorithms
manually in Excel, forcing me to break down each step and truly understand their
mechanics. Unlike automated tools, Excel makes every calculation visible,
reinforcing intuition for methods like linear and logistic regression, decision
trees, neural networks, k-NN, Naïve Bayes, sentence sentiment analysis, and
more.