Statistics: An introduction using R

By Daphna Harel
2026

Description

The R language of statistical computing has an interesting history. It evolved from the S language, which was first developed at the AT&T Bell Laboratories by Rick Becker, John Chambers and Allan Wilks. Their idea was to provide a software tool for professional statisticians who wanted to combine state-of-the-art graphics with powerful model-fitting capability. S is made up of three components. First and foremost, it is a powerful tool for statistical modelling. It enables you to specify and t statistical models to your data, assess the goodness of fit and display the estimates, standard errors and predicted values derived from the model. It provides you with the means to dene and manipulate your data, but the way you go about the job of modelling is not predetermined, and the user is left with maximum control over the model-fitting process. Second, S can be used for data exploration, in tabulating and sorting data, in drawing scatter plots to look for trends in your data, or to check visually for the presence of outliers. Third, it can be used as a sophisticated calculator to evaluate complex arithmetic expressions, and a very flexible and general object-orientated programming language to perform more extensive data manipulation. One of its great strengths is in the way in which it deals with vectors (lists of numbers). These may be combined in general expressions, involving arithmetic, relational and transformational operators such as sums, greater-than tests, logarithms or probability integrals. This book is an introduction to the essentials of statistical analysis for students who have little or no background in mathematics or statistics. The audience includes first and second year undergraduate students in science, engineering, medicine and economics, along with post experience and other mature students who want to relearn their statistics, or to switch to the powerful new language of R. For many students, statistics is the least favourite course of their entire time at university. The approach adopted here involves virtually no statistical theory. Instead, the assumptions of the various statistical models are discussed at length, and the practice of exposing statistical models to rigorous criticism is encouraged. This book provides an elementary-level introduction to R, targeting both non-statistician scientists in various fields and students of statistics. The main mode of presentation is via code examples with liberal commenting of the code and the output, from the computational as well as the statistical viewpoint. Brief sections introduce the statistical methods before they are used.

About Author

Daphna Harel is Associate Professor of Applied Statistics at New York University, USA. She is known for her innovative pedagogical approach to the teaching of statistics, from the introductory undergraduate to the advanced graduate level. She earned her BSc and PhD from the McGill University Department of Mathematics and Statistics, Canada. Her research has been supported by federal agencies and foundations, such as the National Institutes for Health, the Canadian Institutes for Health Research, and the Spencer Foundation. As a highly productive researcher, she has published numerous peer-reviewed articles across statistics, as well as several domain areas.

Table of Content

Preface
Chapter 1. Introduction to Statistics and R
Chapter 2. Organizing and Summarizing Data
Chapter 3. Data Visualization in R
Chapter 4. Probability Basics
Chapter 5. Probability Distributions
Chapter 6. Discrete Probability Distributions
Chapter 7. Continuous Probability Distributions
Chapter 8. Sampling and Sampling Distributions
Chapter 9. Estimation and Confidence Intervals
Chapter 10. Comparing Two Groups
Bibliography
Index