P-Hacking in Statistics

Carrying out a Hypothesis Test without any errors is not quite straightforward. A few things have to be considered & done right in order to maximize the chances of a correct result.

Consider 10 versions of the vaccine for Corona Virus. Hypothesis Tests are carried out for each of them by comparing the time taken by patient to recover if he took a vaccine VS when he didn’t take a vaccine. After carrying out tests for all the 10 vaccines, consider you get high P values for 9 of them and a low value of 0.02 for one of the vaccines. Can you just conclude that this vaccine is effective as 0.02<0.05? No. If you do so, this might be a false positive. You just P-Hacked.:bulb:

P Hacking is the errors in judgments and getting fooled due to misuse of the statistical analysis.
Solution?
One way is to keep repeating the tests until you have sufficient results. But that poses a problem again. Higher the number of tests, more will be the total false positives (Almost 5% of tests). One solution is False Discovery Rate. FDR takes in p-values of all the experiments and adjusts them for a more confident result. The 0.02 in our case might be adjusted to 0.06.

Values such as 0.06 which are close to 0.05 might still be a false positive!

#statistics #datascience