Books

Welcome to the books section of my portfolio. Here, I list the books I have studied to expand my professional skills.



Software Engineering for Data Scientists

by Catherine Nelson, O’Reilly, USA, published 2024

Skills: code quality, performance and data structures, object-oriented programming, error handling and logging, formatting and linting, documentation, testing, version control, APIs, automation and deployment;

Link to the book

Overview
This book bridges the gap between writing code in notebooks for exploratory analysis and writing robust, maintainable, and reusable software. It covers the practices that software engineers take for granted but data scientists rarely learn formally, such as testing, documentation, refactoring, version control, and deployment. These are exactly the practices I apply when building my software packages.



Causal Inference in Statistics: A Primer

by Judea Pearl, Madelyn Glymour, and Nicholas P. Jewell, USA, published 2016

Skills: structural causal models, causal graphs, d-separation, do-operator, backdoor and front-door criteria, mediation analysis, causal inference in linear systems, counterfactuals;

Link to the book

Overview
This book is the technical companion to The Book of Why: a concise introduction to the formal methods of causal inference, with exercises in every chapter. It covers how to encode assumptions in causal graphs, how to decide which variables to adjust for, how to compute the effects of interventions from observational data, and how to reason about counterfactuals. My particular focus was the sections on linear systems, which show how these methods translate to linear models and regression coefficients, the models I work with most in my research.



The Book of Why

by Judea Pearl and Dana Mackenzie, USA, published 2018

Skills: causal inference, causal diagrams, confounding, interventions, counterfactuals;

Link to the book

Overview
In this book, Turing Award winner Judea Pearl explains the “causal revolution” for a broad audience: why classical statistics long avoided questions of cause and effect, and how causal diagrams and the do-operator now make it possible to answer them from data. It introduces the “ladder of causation”, which distinguishes three levels of reasoning: association (seeing), intervention (doing), and counterfactuals (imagining). It gave me the conceptual big picture of causal inference, which the formal methods then build on.



Modeling Mindsets

by Christoph Molnar, Germany, published 2022

Skills: Frequentist and Bayesian statistics, causal inference, supervised unsupervised reinforcement machine learning, deep learning;

Link to the ebook

Overview
This book discussed the idea of viewing modeling not in terms of the algorithms/model classes, like linear regression, but in terms of mindsets. Depending on the mindset, different conclusions can be drawn from the model and different outcomes achieved. For example, when linear regression is used with the Frequentist mindset, the focus is on hypothesis testing with the help of the parameter values, and the outcomes are p-values and confidence intervals. On the other hand, when the same algorithm is used with a supervised machine learning mindset, the focus is on the prediction performance on holdout test data.



Machine Learning Yearning

by Andrew Ng, USA, published 2018

Skills: machine learning, error analysis, project structuring, performance benchmarking, data strategy development;

Link to the ebook

Overview
This ebook is like a guidebook for planning and improving machine learning
projects. Instead of focusing on the math or code, it teaches you how to
think about building AI systems—like choosing the right data, figuring out what’s going wrong, and deciding what to fix first. It’s about making smart decisions to help your machine learning models work better in the real world.



Machine Learning Engineering

by Andriy Burkov, Canada, published 2020

Skills: machine learning, feature engineering, modeling, deployment and monitoring;

Link to the ebook

Overview
This ebook focuses on navigating machine learning projects, from best practices of how to split the data, how to get great features, to choosing appropriate models for the given problems, and more. This book gave me very valuable information of what can be called the art of machine learning.



Introduction to Statistics

by David Lane, Rice University, USA, published 2003

Skills: distributions, probability, estimation, hypothesis testing, power, regression, transformations, chi square, effect size, research design, etc.;

Link to the ebook

Overview
I spent many hours deeply engaging with this book, which covers all fundamental topics in statistics, from graphing and summarizing distributions to probability, hypothesis testing, regression, and ANOVA. Working through each chapter strengthened my intuition for statistical concepts, giving me a solid foundation for data analysis, experimental design, and statistical inference.



Learn Data Mining Through Excel

by Hong Zhou, University of Saint Joseph, USA, published 2023

Skills: machine learning algorithms;

Link to the ebook

Overview
This ebook provided a hands-on approach to data mining by implementing algorithms manually in Excel, forcing me to break down each step and truly understand their mechanics. Unlike automated tools, Excel makes every calculation visible, reinforcing intuition for methods like linear and logistic regression, decision trees, neural networks, k-NN, Naïve Bayes, sentence sentiment analysis, and more.