chapter nine

9 Crimes against data: Statistics done wrong

 

In this chapter

  • Common statistics mistakes and why they happen
  • Cognitive and data biases and how to avoid them
  • Identifying and avoiding p-hacking, which involves cherry-picking models and data

Data, statistics, analytics, and machine learning can be easily misused. It is easy to approach data with the wrong objective and to take advantage of the flexibility that modeling can offer. This is often not malicious, but rather a result of human nature. We want things to work, and we want to troubleshoot, but this desire can quickly slip into making the abstract ideas of statistics bend to our will.

This chapter discusses how to identify and prevent common mistakes in statistics, including p-value hacking.

The importance of reading past headlines

It’s common for news stories to lack critical context in their headlines. A catchy headline may get attention, but it can bury important details in the article that marginalize the headline at best and contradict it at worst. In 2024, I encountered an article from the New York Post titled “Celebrity-loved Diet Linked to Higher Risk of Heart Disease Death” (https://mng.bz/4nyV).

Confounding variables

Data bias

Selection bias

Recall bias

Anchor bias

Confirmation bias

Survivorship bias

The operating domain

Data coverage and the S-curve

Data drift

P-hacking

Texas sharpshooter fallacy

Torturing data and models

Why people p-hack

Summary