Showing posts with label recommended reading. Show all posts
Showing posts with label recommended reading. Show all posts

Saturday, March 22, 2014

Statistical Cognition

As consultants, members of DaSAL are frequently exposed to new statistical techniques that drive our continual learning. Unsurprisingly, we often find learning new techniques and approaches to be a challenging process. Perhaps more surprising, however, are findings that even trained researchers find it challenging to understand the fundamental statistical techniques that are used in almost all psychological research, and may be overconfident in their true level of understanding.

Shedding light on this phenomenon is research on statistical cognition, the study of how people understand and statistical concepts and the presentation of statistical analyses. Some statistical cognition research has examined how people interpret findings from NHST and estimation like confidence intervals. While researchers can often make more accurate interpretations of data using confidence intervals than significance testing (Coulson et al., 2012) researchers’ understanding of confidence intervals is far from  perfect. In fact, several findings have shown that researchers misunderstand confidence intervals (Belia et al., 2005; Hoekstra et al., 2014). Reading about common misconceptions of fundamental statistics to our field highlights the need to review our own understanding of these “basic” concepts even as we develop our knowledge of increasingly complex statistical procedures. Below are some interesting articles that provide insight into some common shortcomings of our statistical cognition.

Belia, S., Fidler, F., Williams, J., & Cumming, G. (2005). Researchers misunderstand confidence intervals and standard error bars. Psychological Methods, 10, 389-396.

Beyth-Marom, R., Fidler, F., & Cumming, G. (2008). Statistical cognition: Towards evidence-based practice in statistics and statistics education. Statistics Education Research Journal., 7, 20-39

Coulson, M., Healey, M., Fidler, F., & Cumming, G. (2010). Confidence intervals permit, but do not guarantee, better inference than statistical significance testing. Frontiers in Quantitative Psychology and Measurement, 1:26. doi:10.3389/fpsyg.2010.00026.

Hoekstra, R., Morey, R. D., Rouder, J. N., Wagenmakers, E.-J. (2014). Robust misinterpretation of confidence intervals. Psychonomic Bulletin & Review, 1-8.

Monday, January 20, 2014

Open-Ended Questions in Survey Research

Controversy initially erupted over the use of open-ended (or free response) and close-ended questions in surveys and interviews (Converse, 1984) in the 1940’s. While close-ended questions rose to dominance due in large part to the expense of processing and analyzing free response data (Converse, 1984), interest in mixed methods in survey research has re-emerged in the last several years. Driving this resurgence is the value open-ended questions can add to the interpretation of responses; including open-ended prompts in surveys populated with otherwise close-ended questions provides insight into the considerations, concerns, and thought processes of survey participants. Specifically, open-ended comments are forwarded as  useful in understanding the replies to closed questions (Driscoll, Appiah-Yeboah, Salib, & Rupert, 2007; Garcia et al., 2004), providing more depth to the topics discussed in the survey (Garcia et al., 2004), identifying new research issues (Garcia et al., 2004), obtaining feedback on the research process (Garcia et al., 2004), and avoiding bias that may result from suggesting responses to participants (Reja, Manfreda, Hlebec, & Vehovar, 2003).

Despite the many advantages of using open-ended comments in survey research, such questions introduce a unique set of challenges and concerns. Most notably, open-ended questions require extensive coding and are often associated with higher levels of non-response (Reja et al., 2003). How, then, can open-ended comments be employed most productively in research efforts? We provide an outline below of best practices in the use of open-ended questions as suggested by research in this area.

  1. Target open-ended question prompts toward specific topics. Responses to general (e.g., “if you have any additional comments, please provide them below”) prompts are more likely to vary in relevance and scope (Evans et al., 2005; Garcia et al., 2004) and may not provide the level of detail or kind of information desired.
  2. To increase response rate on open-ended questions, include targeted questions throughout the survey rather than only using a general open-ended question at the end of the survey. Open-ended questions asked at the end of a survey may elicit shorter answers than open-ended questions asked earlier in a survey (Galesic & Bosnjak, 2009).
  3. Assess response bias to open-ended comments. Certain people, including those with education (Garcia et al., 2004), more interest in the survey topic (Geer, 1988; 1991), higher perceptions of survey value (Rogelberg, Fisher, Maynard, Hakel, & Horvath, 2001), and more negative experiences (Evans et al., 2005; Garcia et al., 2004; Poncheri, Lindberg, Thompson, & Surface, 2008) may be more likely to respond to open-ended questions. Consequently, responses to open-ended comments may not be representative of all respondents’ opinions.
  4. To reduce negativity bias in – and potentially boost response rate to – responses to open-ended questions, provide more detailed, motivating instructions in open-ended item stems (Smyth, Dillman, Christian, & McBride, 2009). Including explanations or instructions in the stem (e.g., emphasizing the importance of open-ended responses to the project) may improve open-ended item response length, elaboration on themes, and item response rate (Smyth et al., 2009). However, it is unclear whether these instructions would substantively affect the negativity of responses in addition to the response rate.
Using open-ended questions in survey research can illuminate responses to close-ended questions, providing researchers with richer data. To minimize the potential weaknesses of open-ended questions, the literature suggests that researchers should use targeted open-ended questions with motivating instructions throughout their surveys rather than only using a general open-ended question at the end of the survey. Additionally, researchers should assess how generalizable the content of open-ended responses is by comparing the characteristics and attitudes of participants who responded and did not respond to open-ended questions.

References
Converse, J. M. (1984). Strong arguments and weak evidence: The open/closed questioning controversy of the 1940’s. Public Opinion Quarterly, 48, 267-282.
Driscoll, D. L., Appiah-Yeboah, A., Salib, P., & Rupert, D. J. (2007). Merging qualitative and quantitative data in mixed methods research: How to an why not. Ecological and Environmental Anthropology, 3, 19-28.
Evans, J., Lambert, T. W., & Goldacre, M. J. (2005). Postal surveys of doctors’ careers: Who writes comments and what do they write about? Quality & Quantity, 39, 217-239.
Galesic, M. & Bosnjak, M. (2009). Effects of questionnaire length on participation and indicators of response quality in a web survey. Public Opinion Quarterly, 73, 349-360.
Garcia, J., Evans, J., & Reshaw, M. (2004). “Is there anything else you would like to tell us” – Methodological issues in the use of free-text comments from postal surveys. Quality & Quantity, 38, 113-125.
Geer, J. G. (1991). Do open-ended questions measure “salient” issues? Public Opinion Quarterly, 55, 360-370.
Geer, J. G. (1988). What do open-ended questions measure? Public Opinion Quarterly, 52, 365-371.
Poncheri, R. M., Lingberg, J. T., Thompson, L. F. & Surface, E. A. (2008). A comment on employee surveys: Negativity bias in open-ended responses. Organizational Research Methods, 11, 614-630.
Reja, U., Manfreda, K. L., Hlebec, V., & Vehovar, V. (2003). Open-ended vs. close-ended questions in web questionnaires. Developments in Applied Statistics, 19, 159-177.
Rogelberg, S., G., Fisher, G. G., Maynard, D. C., Hakel, M. D., & Hovath, M. (2001). Attitudes toward surveys: Development of a measure and its relationship to respondent behavior. Organizational Research Methods, 4, 3-25.
Smyth, J. D., Dillman, D. A., Christian, L. M., & McBride, M. (2009). Open-ended questions in web surveys: Can increasing the size of answer boxes and providing extra verbal instructions improve response quality? Public Opinion Quarterly, 73, 325-337.

Saturday, November 9, 2013

Building on Beta: Indicators of Predictor Importance



Researchers are rarely interested in knowing only whether a variable exhibits a significant relationship with another variable. The more interesting research questions, and answers, are often about the relative importance of a variable in predicting an outcome. Indicators of relative importance may rank order predictors’ contributions to an overall regression effect or partition the variance explained by individual predictors into unique and shared variance. Relative importance of predictors is important for theoretical development because it encourages parsimony, but also for practicality as researchers and practitioners are faced with limited resources.

A brief review of psychology publications that report multiple regression (MR) analyses would show that beta coefficients are clearly the most popular indicator of relative importance reported and interpreted. However, an exclusive reliance on beta coefficients to interpret MR results, and predictor importance in particular, is often misguided and we will briefly review additional indicators that can provide useful information.

Beta Coefficients
In multiple regression, beta coefficients (i.e., standardized regression coefficients) represent the expected change in the dependent variable (in standard deviation units) per standard deviation increase in the independent variable, holding all other independent variables constant. When independent variables are perfectly uncorrelated, one can estimate variable importance by squaring the beta weights. However, this becomes problematic when independent variables are correlated, as is often the case. In these situations, a given beta weight can reflect the explained variance it shares with other variables in the model, and it becomes difficult to disentangle a variable’s unique importance from the “extra credit” it gets from shared variance. As such, beta weights should be relied on as an easily computed but preliminary indicator of a predictor’s contribution to a regression coefficient.

Pratt Product Measure
The Pratt product measure is a relatively simple method of partitioning a model’s explained variance (R2) into non-overlapping segments. The Pratt measure multiplies a variable’s zero-order correlation with a dependent variable by its beta weight. The correlation measure captures the direct effect of the predictor in isolation from other predictors while the beta weight is an indicator of a predictor’s contribution accounting for all other predictors in the model.  If a given beta weight is inflated by shared variance, multiplying the value by its corresponding (smaller) zero order correlation will correct for the beta’s inflation. Summing the Pratt product measures for all the variables in a model generally results in the R2 except in some cases where a negative product is yielded (often a signal of suppression), allowing for straightforward partitioning and ranking of variables.

Commonality Coefficients
Commonality analysis partitions R2 into variance that is unique to each independent variable (unique effects; e.g., X1…X3) and variance that is shared by all possible subsets of predictors (common effects; e.g., X1X2, X2X3, etc.). Partitioning variance in this way produces non-overlapping values that can be compared easily. Common effects, in particular, provide a great deal of information about the overlap between independent variables and how this overlap contributes to predicting the dependent variable. However, common effects can be difficult to interpret, especially as the number of variables increases in a model and commonalities reflect the combination of more than two variables.

Relative Weights
Relative weight analysis (RWA) is a slightly more complex method of partitioning R2 between independent variables. When predictors in a MR are correlated, RWA uses principal components analysis to generate principal components that are the most highly correlated with the dependent variable while being uncorrelated with one another. The dependent variable is regressed on these components in one analysis and a second analysis is conducted in which the original independent variables are regressed on the components. A given relative weight is equal to the product of the squared regression coefficient from the first analysis and the squared regression coefficient from the second analysis. Dividing relative weights by R2 then allows for ranking of individual predictors’ contributions. Essentially, relative weights are an indicator of a predictor’s contribution to explaining the dependent variable as a joint function of how highly related independent variables are to the principal components and how highly related the principal components are to the dependent variable. As such, RWA partitions R2 while minimizing the influence of multicollinearity among predictors, a major strength of the method.

Dominance Analysis
Dominance analysis compares the unique variance contributed by pairs of independent variables across all possible subsets of predictors. To assess complete dominance, the unique effect of a given independent variable when entered last in multiple regression equation is compared to another independent variable across all subsets of a model. Complete dominance occurs when the unique effect is greater across all the models. Conditional dominance is determined by calculating the averages of all independent variables’ contributions to all possible model subsets and comparing those averages. Conditional dominance is shown when an independent variable contributes more to predicting the dependent variable, on average across all possible models, compared to another independent variable.  

References
Azen, R., & Budescu, D. V. (2003). The dominance analysis approach for comparing predictors in multiple regression. Psychological Methods, 8, 129-148.

Johnson, J. W. (2000). A heuristic method for estimating the relative weight of predictor variables in multiple regression. Multivariate Behavioral Research, 35, 1-19. 

Nimon, K. F. & Oswald, F. L. (In Press). Understanding the results of multiple linear regression: Beyond standardized regression coefficients. Organizational Research Methods.  

Nimon, K. F., & Reio, T. (2011). Regression commonality analysis: A technique for quantitative theory building. Human Resource Development Review, 10, 329-340. 
 


Monday, October 7, 2013

How Can Researchers Claim Support for the Null Hypothesis?

Traditional null hypothesis significance testing (NHST) does not allow researchers to claim support for the null hypothesis given a finding of non-significance (i.e., p > .05).  This is due to the fact that NHST methods only signify the probability of a set of data given the null model (D | H0), and it does necessarily follow that the null model is probable given the data (H0 | D). Rather, researchers using NHST methods are only able to state that the null “was not able to be rejected.” This, of course, belies the spirit of scientific pursuit and goes against the desires of most researchers, who wish to make substantive claims based on their research findings. If this is the case, how might researchers proceed?

Unlike NHST methods, Bayesian methods allow researchers to approximate the probability of a hypothesis given the data (H0 | D) using a model comparison approach. Specifically, Bayes Factors are used to compute the ratio of two models (e.g., the null and the alternative hypothesis) to determine which is more supported by the statistical evidence. This allows researchers to make claims in support of either the null or alternative model. Consequently, Bayesian methods offer a useful alternative to NHST that is more in line with the goals of the scientific enterprise, which seek to support the validity of competing hypotheses. However, Bayesian methods are often computationally complex and many researchers in the field of psychology may not have the skills necessary appropriately employ them. Yet, easily computable alternatives exist.

In addition to reviewing the drawbacks of NHST in greater detail, Masson (2011) offers a reasonable method whereby researchers can appropriate the benefits of Bayesian methods in a way that requires little computational complexity. Instead of computing Bayes Factors in a traditional manner, he builds on a method specified by Wagenmakers (2007) that computes the “Bayes Information Criterion” or BIC. The BIC uses results computed from traditional NHST tests (such as the sums-of-squares values generated by ANOVA’s) to easily produce a model comparison ratio that approximates the Bayes Factor. Researchers are urged to read Masson (2011) in order to understand the full methodological processes. Additionally, those wishing for a useful categorization scheme to describe the magnitude of the resulting ratio should refer to Raftery (1995). Finally, researchers interested in reading more about the drawbacks of NHST and Bayesian methods in general will find articles by Wagenmakers (2007), Gallistel (2009), and Kruscke (2010, 2013) helpful.

In conclusion, researchers often seek to demonstrate that some experimental manipulations have no demonstrable effect on an outcome variable. Moreover, even when a researcher has no a prior motivation to demonstrate a null effect, it is scientifically responsible to test the extent to which a null model is more or less supported relative to a specified alternative model, rather than simply making assumptions about its validity. This benefits the quality of research that comes out of the psychological research community and provides important information for other researchers that may preserve valuable resources, including time and money. In addition, although most psychological journals still advocate the reporting of NHST results, it is easy enough for researchers to compute BICs and report them alongside more traditional tests.

References

Gallistel, C. R. (2009). The importance of proving the null. Psychological Review, 116,
439-453.

Kruschke, J. K. (2010). Bayesian data analysis. Wiley Interdisciplinary Reviews:
Cognitive Science, 1, 658-676.

Kruschke, J. K. (2013). Bayesian estimation supersedes the t test. Journal of
Experimental Psychology: General, 142, 573-603.

Masson, E. J. (2011). A tutorial on a practical Bayesian alternative to null-hypothesis
significance testing. Behavior Research Methods, 43, 679-690.

Raftery, A. E. (1995). Bayesian model selection in social research. In P. V. Marsden
(Ed.), Sociological methodology 1995 (pp. 111-196). Cambridge: Blackwell.

Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p
values. Psychonomic Bulletin & Review, 14, 779-804.