Would you read a statistics textbook?

Would you read a statistics textbook?

9 pointsby usernametaken296 comments

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • In my dotage I have become interested in radiation epidemiology and, of course, the issue of 'statistical power' comes up repeatedly. As noted, most physicists know only as much statistics as is absolutely necessary for their work. It would be wonderful to introduce statistics by a set of scenarios (possibly with references to relevant parts of the text itself) drawn from, say, 10 distinct fields. To my dismay, I found that chatGPT did a remarkably decent job of telling me what had been missing in my readings. So think of the project as: what value is added above what a reader could find after 30 minutes of AI delving.
  • There are a number of fairly good books on statistics. One of the best is How To Lie With Statistics by Darrell Huff which was written in 1954. There is also the Cartoon Guide to Statistics by Larry Gonick and Woolcott Smith. The problem most statistics books is that they don't focus on what you can do with statistics. It is important if you are a gambler or want to understand artificial intelligence. Most statistics books are super dry and full of bad examples. I would focus on books like The Theory of Gambling and Statistical Logic. It becomes much more interesting when it is applied.
  • Did you see that Andrew Gelman and others just published Bayesian Workflows? https://avehtari.github.io/Bayesian-Workflow/ It sounds similar to what you're after, away from describing the logic of models and the maths, instead it's about (quote from intro) 'There are all sorts of tacit knowledge in applied statistics that do not make it into published papers and textbooks. The present book is intended to put some of these ideas out in the open'

    Where would your book fit into this?

  • A long time ago, I took a "statistics for engineers" class in order to graduate. I slept through most of the classes. It sucked, and 70 percent of it was just "distribution of the week." It did not help that homework was optional for half of it.

    At some point in my professional career I started reading non-fiction books and even bought a used statistics textbook for 10 bucks on abebooks. I didn't end up actually reading it until 12 years later during the COVID lockdown. I ended up shooting for 10 pages a day, 7 days a week. If those 10 pages included review exercises, it would be a long night.

    Could just be the right book at the right time, but this one really helped me understand stuff beyond the normal HS math stuff, like RMS-error, calculating correlation, the difference between standard error and standard deviation, the relationship between sample size and standard error, t-tests, and chi-squared. Working as an SRE/release engineer, this stuff really helped me overcome a lot of _bad_ canary data analysis my predecessors had constructed.

    That book was the 3rd edition of Statistics by Freedman et al.[1] One thing I want to complement was getting the pedagogy right. Most chapters have strong narrative hooks, several "check your knowledge" problems, review exercises, and post chapter bullet points to assist with spaced repetition. There's even a series of "special" review exercises covering entire sections of the book, ie exams.

    For the HN crowd I should also probably note that the book is almost entirely non-bayesian and not intended to prepare readers for further coursework. You will not learn normal phraseology like "IID," "random variable" or "kernel".

    [1]: https://www.amazon.com/dp/B00SLB5Q72?lv=shuf&channelId=520&p...

  • There are a number of good books to explain the main concepts of statistics without getting bogged down in the technical details. As a grad student at a public university in the USA, I worked as a teaching assistant for the undergraduate course Introduction to Statistics. The course used SPSS software to enable the students to perform statistical analysis on data sets. Although we offered a good textbook (listed below), many of the students struggled with the main concepts, so my colleagues and I added a optional supplemental reading list to the syllabus. The following books are the ones we recommended to the students. I have read all these books, and they are pretty good at explaining statistics to anyone interested in learning the material at a conceptual level; anyone can read and understand these books, regardless of their educational background.

    Ian Ayres (2007) Super Crunchers: Why Thinking by Numbers is the New Way to Be Smart.

    Joseph Healey (2020) Statistics: A Tool for Social Research [This is the textbook used in the class, which I highly recommend].

    John Allen Paulos (1988) Innumeracy: Mathematical Illiteracy and Its Consequences.

    Nate Silver (2012) The Signal and the Noise: Why So Many Predictions Fail, but Some Don't.

    Nassim Nicholas Taleb (2005) Fooled by Randomness, second edition.

    Nassim Nicholas Taleb (2010) The Black Swan: The Impact of the Highly Improbable, second edition.

  • Practical advice:

    - put a lot of 'teaching' into the book. Don't write so the learner memorizes; instead, motivate each discussion so that the learner understands, and gains what Wirth calls _coverage_.

    - ensure there are no errors; this will turn off learners.

    - keep a light tone, be yourself not a stuffed shirt. A good model for such writing is Nassim Taleb's non-technical books.

    - never seem to be in a hurry. Many good books were ruined because midway the author lost interest and just hurried up.

    - don't be too verbose, and neither too concise. Every sentence should do required work; learn how to write (if needed) and practice daily.

    - ensure the book uses good, comfortable fonts, and good typography. Must learn this. Many good books ruined due to lack of fluency effect.

    If you conscientiously do these things, there will always be readers for your book.

    Good luck!

  • Having also studied statistics in university (undergrad), something I kept running into is that you can't really unlock the intuition for many concepts without taking more advanced courses. For example, degrees of freedom shows up as early as AP Statistics, but even a non-rigorous visual explanation of it leans on linear algebra, which most students don't see until much later.

    I think more resources like seeing-theory would be great since stats books are almost universally dry (Blitzstein being a notable exception), but I'm not sure how easily more advanced concepts lend themselves to visual explanation in a way that's digestible for a non-stats person.

  • I feel this boils down my learning journey as well. You start unraveling a very good intuition about the underlying concepts MUCH much later, but partly because those intuitions themselves are never conveyed and are supposed to be learned from the proofs, and are an indirect product of learning.
  • For degrees of freedom specifically, check out these videos: https://www.youtube.com/playlist?list=PLmtbsGjqSdcbF261LACoV...

    It's still complicated but the visualizations help a lot!

  • I'm a nerd and do a lot of stats for my job and I would not read a statistics textbook.

    I read a lot of informational things, but math / stats / software has always felt like an area where a book is just the wrong format.

    If I were you I'd make an interactive website like SQLZoo or a video series like StatQuest.

    Those are educational formats that really clicked with me for whatever reason.

  • How would it overlap or differ from 'Statistical Rethinking'?

    This is widely regarded as the most accessible intro textbook to Bayesian statistics.

    https://xcelab.net/rm/

  • Maybe my expectations were too high given all the online praise, but I've been working through it for the past two weeks and I've been underwhelmed:

    - The language is often very vague and imprecise, so it's more difficult than it feels it should be

    - Concepts are sometimes introduced "at random" in ways that only really make sense in hindsight. So as you're reading you're left scratching your head as to why something was brought up.

    - There are constant philosophical and historical digressions that seem to hold deeper meaning, but maybe once you know the topic already.

    - Similarly, constantly talking about people that take issue with the method. People not liking the method is a constant theme (they seem really butthurt about this?). But the craziest part is this all done before you even really understand what the method is!!

    - The editor must have placed some strict requirement of saying "Bayesian" at least five times per page.

    - No index. Useless table of context. But lots of end-notes you feel compelled to flip to constantly

    Overall it feels like a textbook written to impress other statistics professors - and as an outlet for the author to air some frustrations with how people do statistics (which may be completely valid!)

    The overall structure and objectives seem solid for the most part. It's just a lot of the details aren't great. The problems (so far) have been good. The examples in the text are fun and compelling, but you have to do your own legwork to actually pick through all the prose and tie the pieces together - to figure how it fits together mathematically. Fortunately AI helps as a tutor

  • + Krushke’s Doing Bayesian Data Analysis, and for the very basics there is Downey’s Think Bayes. I guess there might be a gap in the literature for a different approach, heck, I’d read it, but there is some very good material already out there.
  • I've got a Bachelors and two Masters degrees in CS/Math, but yet I feel Probability and Statistics is my greatest weakness. I just cannot grok it / build an intuition for it, and believe me, I've tried. My biggest gripe with Prob/Stats textbooks is that it's very hard to explain things without needing to rely on measure theory.

    Maybe probability and statistics are a skill issue on my behalf, but what I absolutely loathe is the absolute lack of standardisation when it comes to notation in measure theory. Every textbook does it differently. All of them assume that their notation is the one everyone uses. Nobody bothers to explain _what_ the notation means. If you ask me, every bit of new notation should be introduced with a sentence or two on "how to read this symbol in your head" - especially when there are indices, subscripts and superscripts involved. It's especially terrible for measure theory because there's so much "implicit" information you're supposed to gather from the context - but in a way I understand it, because if every bit of notation of absolute and complete, I imagine it would be quite hard to type up.

    Anyways, my rant on measure theory notation aside - I would absolutely read yet another Prob/Stats textbook. But unfortunately I will also drop it really quickly if the author doesn't show me any "notation-sympathy" :)

  • My personal opinion is that statistics textbooks usually come from a prescriptive perspective, and that makes it challenging for the reader to get visceral intuition for what is actually going on. Any reader would be far better off just visualizing the damn distribution / samples and using reasonable judgement, instead of implicitly assuming a Gaussians distribution and blindly memorizing tests / formulae. Making the distributions explicit allows us to model them and get an intuition for what the samples are telling us. I would whole-heartedly recommend the Model based machine learning book to anyone (online version is free) https://mbmlbook.com/
  • My larger issue, any time I have tried to learn statistics, is how fast the notation moves. You end up flipping back pages and pages just to double-check a definition that was given once and is now being extended syntactically. It's infuriating.
  • How do you make 'reasonable judgements'? How do you tell whether someone else made reasonable judgements? How do you judge other people's intuition?

    Modelling distributions explicitly sounds nice, yes.

    by eru
  • No idea about statistics, but in most physict courses in my university, they recomend 3 books:

    1) The main book, that has a complete explanation and is well ordered. It's for learning.

    2) Tha Landau book, that is super short and hard. It's only to check you didn't miss any important formula or topic.

    3) There Feynman book, that is anassorted colection of fairytales for physicist. It's a pleasure to read it but you must already read 1 to understand it.

    4) The Shaum book, that is almost a colection of exercices. Some people hate it. Some people love it. I like it as a companion to theother books.

    I guess you are complaining that 1 is boring and want to write 3. It's a good idea, but it's harder than expected.

  • Similar to an old idea I had about how every programming language needs three books:

    1. Basic introduction.

    2. Reference tome, which has absolutely everything.

    3. Cookbook with style advice for the more advanced student, which assumes you've read 1 and can look up various details in 2.

    These days, 2 would be a wiki and 1 would likely be a bunch of pages on that wiki, but it's still good if you have someone sit down and write 3.

    by msla