Showing posts with label NHST. Show all posts
Showing posts with label NHST. Show all posts

Saturday, March 22, 2014

Statistical Cognition

As consultants, members of DaSAL are frequently exposed to new statistical techniques that drive our continual learning. Unsurprisingly, we often find learning new techniques and approaches to be a challenging process. Perhaps more surprising, however, are findings that even trained researchers find it challenging to understand the fundamental statistical techniques that are used in almost all psychological research, and may be overconfident in their true level of understanding.

Shedding light on this phenomenon is research on statistical cognition, the study of how people understand and statistical concepts and the presentation of statistical analyses. Some statistical cognition research has examined how people interpret findings from NHST and estimation like confidence intervals. While researchers can often make more accurate interpretations of data using confidence intervals than significance testing (Coulson et al., 2012) researchers’ understanding of confidence intervals is far from  perfect. In fact, several findings have shown that researchers misunderstand confidence intervals (Belia et al., 2005; Hoekstra et al., 2014). Reading about common misconceptions of fundamental statistics to our field highlights the need to review our own understanding of these “basic” concepts even as we develop our knowledge of increasingly complex statistical procedures. Below are some interesting articles that provide insight into some common shortcomings of our statistical cognition.

Belia, S., Fidler, F., Williams, J., & Cumming, G. (2005). Researchers misunderstand confidence intervals and standard error bars. Psychological Methods, 10, 389-396.

Beyth-Marom, R., Fidler, F., & Cumming, G. (2008). Statistical cognition: Towards evidence-based practice in statistics and statistics education. Statistics Education Research Journal., 7, 20-39

Coulson, M., Healey, M., Fidler, F., & Cumming, G. (2010). Confidence intervals permit, but do not guarantee, better inference than statistical significance testing. Frontiers in Quantitative Psychology and Measurement, 1:26. doi:10.3389/fpsyg.2010.00026.

Hoekstra, R., Morey, R. D., Rouder, J. N., Wagenmakers, E.-J. (2014). Robust misinterpretation of confidence intervals. Psychonomic Bulletin & Review, 1-8.

Tuesday, February 18, 2014

Replication In Psychological Research

Eminent psychologist Daniel Kahneman warned of a “looming train wreck” while others have declared that psychology is in the throes of a replicability crisis (Pashler & Harris, 2012). The questions surrounding replication in psychology are both philosophical and statistical in nature, and are prompting concerted effort to change the incentives in academia and standards of publications. Indeed, the Society for Personality and Social Psychology recently released new standards for research and publications and journals like Perspectives in Psychological Science have published articles about the role of replications. Given the increased interest in replication, we provided a review of some of the key points of discussion below.

What’s all the fuss about?
While replication has always been an important aspect psychological research (as with any science), the recent fervor was energized by the uncovering of several major cases of fraud in social psychology. In particular, the case of Diederik Stapel, who committed years of scientific fraud, raised questions about our tendency to dismiss the importance of direct replications and discard failures to replicate as uninformative.

These cases of fraud compounded other attempts to call attention to the irreproducibility of research, including Ioannidis‘ (2005) exposition on research in biology. While psychology may not be any worse off than other “harder” sciences, problems of reproducibility have implications for attitudes about the utility and credibility of scientific research more generally.

Why is it difficult to replicate research?

Data “cleaning” and Unethical Practices
In some cases, effects may be difficult to replicate because researchers manipulated data in ways that were not fully described in the literature. For example, researchers may find an effect when outliers are deleted and no effect when everyone is included in analyses. If the treatment of outliers is not mentioned in the original research article, there is no way for other researchers to identify this as an important factor in finding the effect even if the choice to delete outliers is valid. In the worst cases, the effects might not replicate because of fraudulent practices on the part of the researcher, as in Stapel case.

The Nature of our Statistics
Most research in psychology uses the Null Hypothesis Significance Testing (NHST) method of inferential statistics. One serious downside of NHST is the tendency to think in terms of a dichotomy in which effects exist (p < .05) or effects don’t exist (p > .05), ignoring the inherent uncertainty in psychological effects. It is tempting to use p-values as indicators of the reliability of an effect, but the problems with this kind of thinking are nicely illustrated in the Dance of the p-values video from psychologist Geoff Cumming. Understanding the variability in p-values suggests that unreliability in “significant” statistical tests is not surprising, and psychologists benefit from thinking in terms of estimation (e.g., confidence intervals) when trying to evaluate replication attempts.

Sample Size
Another pervasive problem in psychological research is our generally small sample sizes. Many of our studies continue to be very underpowered given the typically small-to-moderate effect sizes we find in our research (Nelson et al., 2013). As such, we are less likely to replicate findings in the literature. However, another problem with small sample sizes is the implication for our estimates of effect sizes. Estimates of effect size are going to be more unreliable with small samples, which can lead to overestimation of the true size of an effect in the population. As such, it might be much more difficult to replicate a finding than a reported effect size (if there is one) would suggest.

What is a good replication?
The best replications begin with transparency and collaboration between both the replicator and original authors. With this in mind, Brandt et al. (2014) also outlined important considerations that make for good replications. These authors argue that a good replication has five main ingredients, including the following:

1) Carefully defining the effects and methods that the researcher intends to replication 
2) Following exactly the methods of the original study 
3) Having high statistical power 
4) Making complete details about the replication available  
5) Evaluating replication results and comparing them critically


The article goes into greater detail in relation to each of these “ingredients,” but the question of how to evaluate replications deserves special attention. Brandt et al. (2014) recommend evaluating the replication of effects in two ways: 1) reporting the size, direction and confidence interval of the target effect (tells us whether the effect is different from the null) and 2) testing whether the effect is different from the original effect. Another approach to evaluate the success of replications is to apply a meta-analytic aggregation if the replication and original study effects. There are many other approaches to evaluating replications (Simohnson, 2013), but it is clear that evaluating the significance of results is insufficient.

What can I do if I am interested in conducting replications?
If you are interested in conducting replications, the Open Science Framework’s Reproducibility Project  is seeking partners to conduct replications of studies found in 2008 issues of Journal of Personality and Social Psychology, Psychological Science, and Journal of Experimental Psychology: Learning, Memory, and Cognition. In doing so, the OSF aims to learn more about the overall reproducibility in the psychology literature. The OSF also provides assistance with carrying out the replications and provides workflow resources that can make it easier for others to replication your own research. 

References
Brandt, M. J., IJzerman, H., Dijksterhuis, A., Farach, F. J., Geller, J., Giner-Sorolla, R., ... & Van't Veer, A. (2014). The replication recipe: What makes for a convincing replication?. Journal of Experimental Social Psychology, 50, 217-224

Ioannidis, J. P. (2005). Why most published research findings are false. PLoS medicine, 2(8), e124.

Pashler, H., & Harris, C. R. (2012). Is the replicability crisis overblown? Three arguments examined. Perspectives on Psychological Science, 7(6), 531-536.

Simonsohn, U. (2013). Evaluating replication results. Available at SSRN: http://ssrn.com/
abstract=2259879

Monday, October 7, 2013

How Can Researchers Claim Support for the Null Hypothesis?

Traditional null hypothesis significance testing (NHST) does not allow researchers to claim support for the null hypothesis given a finding of non-significance (i.e., p > .05).  This is due to the fact that NHST methods only signify the probability of a set of data given the null model (D | H0), and it does necessarily follow that the null model is probable given the data (H0 | D). Rather, researchers using NHST methods are only able to state that the null “was not able to be rejected.” This, of course, belies the spirit of scientific pursuit and goes against the desires of most researchers, who wish to make substantive claims based on their research findings. If this is the case, how might researchers proceed?

Unlike NHST methods, Bayesian methods allow researchers to approximate the probability of a hypothesis given the data (H0 | D) using a model comparison approach. Specifically, Bayes Factors are used to compute the ratio of two models (e.g., the null and the alternative hypothesis) to determine which is more supported by the statistical evidence. This allows researchers to make claims in support of either the null or alternative model. Consequently, Bayesian methods offer a useful alternative to NHST that is more in line with the goals of the scientific enterprise, which seek to support the validity of competing hypotheses. However, Bayesian methods are often computationally complex and many researchers in the field of psychology may not have the skills necessary appropriately employ them. Yet, easily computable alternatives exist.

In addition to reviewing the drawbacks of NHST in greater detail, Masson (2011) offers a reasonable method whereby researchers can appropriate the benefits of Bayesian methods in a way that requires little computational complexity. Instead of computing Bayes Factors in a traditional manner, he builds on a method specified by Wagenmakers (2007) that computes the “Bayes Information Criterion” or BIC. The BIC uses results computed from traditional NHST tests (such as the sums-of-squares values generated by ANOVA’s) to easily produce a model comparison ratio that approximates the Bayes Factor. Researchers are urged to read Masson (2011) in order to understand the full methodological processes. Additionally, those wishing for a useful categorization scheme to describe the magnitude of the resulting ratio should refer to Raftery (1995). Finally, researchers interested in reading more about the drawbacks of NHST and Bayesian methods in general will find articles by Wagenmakers (2007), Gallistel (2009), and Kruscke (2010, 2013) helpful.

In conclusion, researchers often seek to demonstrate that some experimental manipulations have no demonstrable effect on an outcome variable. Moreover, even when a researcher has no a prior motivation to demonstrate a null effect, it is scientifically responsible to test the extent to which a null model is more or less supported relative to a specified alternative model, rather than simply making assumptions about its validity. This benefits the quality of research that comes out of the psychological research community and provides important information for other researchers that may preserve valuable resources, including time and money. In addition, although most psychological journals still advocate the reporting of NHST results, it is easy enough for researchers to compute BICs and report them alongside more traditional tests.

References

Gallistel, C. R. (2009). The importance of proving the null. Psychological Review, 116,
439-453.

Kruschke, J. K. (2010). Bayesian data analysis. Wiley Interdisciplinary Reviews:
Cognitive Science, 1, 658-676.

Kruschke, J. K. (2013). Bayesian estimation supersedes the t test. Journal of
Experimental Psychology: General, 142, 573-603.

Masson, E. J. (2011). A tutorial on a practical Bayesian alternative to null-hypothesis
significance testing. Behavior Research Methods, 43, 679-690.

Raftery, A. E. (1995). Bayesian model selection in social research. In P. V. Marsden
(Ed.), Sociological methodology 1995 (pp. 111-196). Cambridge: Blackwell.

Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p
values. Psychonomic Bulletin & Review, 14, 779-804.