Chi-squared test for the relationship between two categorical variables overview

This page offers structured overviews of one or more selected methods. Add additional methods for comparisons (max. of 3) by clicking on the dropdown button in the right-hand column. To practice with a specific method click the button at the bottom row of the table

Chi-squared test for the relationship between two categorical variables	Kruskal-Wallis test	Binomial test for a single proportion $z$ test for a single proportion $z$ test for the difference between two proportions Goodness of fit test Chi-squared test for the relationship between two categorical variables One sample $z$ test for the mean One sample $t$ test for the mean Paired sample $t$ test Two sample $z$ test Two sample $t$ test - equal variances not assumed Two sample $t$ test - equal variances assumed One way ANOVA Two way ANOVA Pearson correlation Regression (OLS)Logistic regression Mann-Whitney-Wilcoxon test Kruskal-Wallis test Sign test McNemar's test Cochran's Q test Marginal Homogeneity test / Stuart-Maxwell test Friedman test Wilcoxon signed-rank test One sample Wilcoxon signed-rank test Spearman's rho
Independent /column variable	Independent/grouping variable
One categorical with $I$ independent groups ($I \geqslant 2$)	One categorical with $I$ independent groups ($I \geqslant 2$)
Dependent /row variable	Dependent variable
One categorical with $J$ independent groups ($J \geqslant 2$)	One of ordinal level
Null hypothesis	Null hypothesis
H₀: there is no association between the row and column variable More precisely, if there are $I$ independent random samples of size $n_i$ from each of $I$ populations, defined by the independent variable: H₀: the distribution of the dependent variable is the same in each of the $I$ populations If there is one random sample of size $N$ from the total population: H₀: the row and column variables are independent	If the dependent variable is measured on a continuous scale and the shape of the distribution of the dependent variable is the same in all $I$ populations: H₀: the population medians for the $I$ groups are equal Else: Formulation 1: H₀: the population scores in any of the $I$ groups are not systematically higher or lower than the population scores in any of the other groups Formulation 2: H₀: P(an observation from population $g$ exceeds an observation from population $h$) = P(an observation from population $h$ exceeds an observation from population $g$), for each pair of groups. Several different formulations of the null hypothesis can be found in the literature, and we do not agree with all of them. Make sure you (also) learn the one that is given in your text book or by your teacher.
Alternative hypothesis	Alternative hypothesis
H₁: there is an association between the row and column variable More precisely, if there are $I$ independent random samples of size $n_i$ from each of $I$ populations, defined by the independent variable: H₁: the distribution of the dependent variable is not the same in all of the $I$ populations If there is one random sample of size $N$ from the total population: H₁: the row and column variables are dependent	If the dependent variable is measured on a continuous scale and the shape of the distribution of the dependent variable is the same in all $I$ populations: H₁: not all of the population medians for the $I$ groups are equal Else: Formulation 1: H₁: the poplation scores in some groups are systematically higher or lower than the population scores in other groups Formulation 2: H₁: for at least one pair of groups: P(an observation from population $g$ exceeds an observation from population $h$) $\neq$ P(an observation from population $h$ exceeds an observation from population $g$)
Assumptions	Assumptions
Sample size is large enough for $X^2$ to be approximately chi-squared distributed under the null hypothesis. Rule of thumb: 2 $\times$ 2 table: all four expected cell counts are 5 or more Larger than 2 $\times$ 2 tables: average of the expected cell counts is 5 or more, smallest expected cell count is 1 or more There are $I$ independent simple random samples from each of $I$ populations defined by the independent variable, or there is one simple random sample from the total population	Group 1 sample is a simple random sample (SRS) from population 1, group 2 sample is an independent SRS from population 2, $\ldots$, group $I$ sample is an independent SRS from population $I$. That is, within and between groups, observations are independent of one another
Test statistic	Test statistic
$X^2 = \sum{\frac{(\mbox{observed cell count} - \mbox{expected cell count})^2}{\mbox{expected cell count}}}$ Here for each cell, the expected cell count = $\dfrac{\mbox{row total} \times \mbox{column total}}{\mbox{total sample size}}$, the observed cell count is the observed sample count in that same cell, and the sum is over all $I \times J$ cells.	$H = \dfrac{12}{N (N + 1)} \sum \dfrac{R^2_i}{n_i} - 3(N + 1)$ Here $N$ is the total sample size, $R_i$ is the sum of ranks in group $i$, and $n_i$ is the sample size of group $i$. Remember that multiplication precedes addition, so first compute $\frac{12}{N (N + 1)} \times \sum \frac{R^2_i}{n_i}$ and then subtract $3(N + 1)$. Note: if ties are present in the data, the formula for $H$ is more complicated.
Sampling distribution of $X^2$ if H₀ were true	Sampling distribution of $H$ if H₀ were true
Approximately the chi-squared distribution with $(I - 1) \times (J - 1)$ degrees of freedom	For large samples, approximately the chi-squared distribution with $I - 1$ degrees of freedom. For small samples, the exact distribution of $H$ should be used.
Significant?	Significant?
Check if $X^2$ observed in sample is equal to or larger than critical value $X^{2*}$ or Find $p$ value corresponding to observed $X^2$ and check if it is equal to or smaller than $\alpha$	For large samples, the table with critical $X^2$ values can be used. If we denote $X^2 = H$: Check if $X^2$ observed in sample is equal to or larger than critical value $X^{2*}$ or Find $p$ value corresponding to observed $X^2$ and check if it is equal to or smaller than $\alpha$
Example context	Example context
Is there an association between economic class and gender? Is the distribution of economic class different between men and women?	Do people from different religions tend to score differently on social economic status?
SPSS	SPSS
Analyze > Descriptive Statistics > Crosstabs... Put one of your two categorical variables in the box below Row(s), and the other categorical variable in the box below Column(s) Click the Statistics... button, and click on the square in front of Chi-square Continue and click OK	Analyze > Nonparametric Tests > Legacy Dialogs > K Independent Samples... Put your dependent variable in the box below Test Variable List and your independent (grouping) variable in the box below Grouping Variable Click on the Define Range... button. If you can't click on it, first click on the grouping variable so its background turns yellow Fill in the smallest value you have used to indicate your groups in the box next to Minimum, and the largest value you have used to indicate your groups in the box next to Maximum Continue and click OK
Jamovi	Jamovi
Frequencies > Independent Samples - $\chi^2$ test of association Put one of your two categorical variables in the box below Rows, and the other categorical variable in the box below Columns	ANOVA > One Way ANOVA - Kruskal-Wallis Put your dependent variable in the box below Dependent Variables and your independent (grouping) variable in the box below Grouping Variable
Practice questions	Practice questions

Chi-squared test for the relationship between two categorical variables - overview