When a Regression Pretends to Read Minds

2026-09-14 · 5,815 words · Singular Grit Substack · View on Substack

A response to Bailey Connor on identity-protective cognition, the limits of cultural explanations, and what happens when the conclusions outrun the mathematics.

Keywords: identity-protective cognition; cultural cognition; risk perception; causal inference; statistical power; regression interactions; measurement validity; missing data; intellectual independence; scientific criticism.

Bailey Connor asked for my assessment of Kahan and colleagues’ Culture and Identity-Protective Cognition: Explaining the White-Male Effect in Risk Perception. It is a worthwhile question, not least because the paper offers something perpetually attractive: an explanation of why other people disagree with us. Such explanations have an excellent social life. Their evidential credentials sometimes receive rather less attention.

The central distinction is straightforward. Discovering that opinions cluster around cultural commitments is one achievement. Establishing that a particular defensive psychological mechanism produced those opinions is another. Establishing that the opinions are false is a third. A paper may make progress on the first without completing the second, and without even conducting the inquiry required for the third. These distinctions are not academic ornament. They determine what the research actually entitles anyone to say.

My assessment is that the study reports interesting associations, but repeatedly gives them more explanatory authority than its design and statistical tests warrant. There are also directly checkable inconsistencies in the manuscript’s coding instructions and reporting. Those deserve correction, not theatrical accusations. A mislabelled table is not evidence of dishonesty; neither is an elegant theory a licence to overlook one.

Scope matters. This article audits the supplied SSRN manuscript revised on 14 June 2007. Page references below are to its printed pagination, not the journal’s page numbers. Calculations use the figures reported in that manuscript; this is not a replication using respondent-level data. The journal article is identified in the references, but I do not assume that every manuscript defect survived into its final typeset version. Later methodological sources clarify the issues; they are not evidence that a particular modern software package should have been used in 2004 (Kahan et al., 2007, manuscript title page).

A worthwhile observation, promoted beyond its rank

The study surveyed 1,844 American adults by telephone between June and September 2004. It measured cultural orientations along hierarchy–egalitarianism and individualism–communitarianism dimensions, and related them to environmental, gun-related, and abortion-related risk responses. The sample included an African-American oversample of 242 respondents. Political ideology, party affiliation, religious affiliation, and several demographic characteristics also entered the analyses (Kahan et al., 2007, manuscript pp. 9, 13–16, 43).

On the authors’ reported interpretation, the findings are more interesting than the proposition that white men are uniformly fearless. Hierarchical and individualistic orientations predict lower concern about environmental risks and a more favourable assessment of guns. For abortion, hierarchy predicts greater concern, whereas individualism predicts less. The direction depends on the activity. The authors also explicitly deny that identity-protective cognition is peculiar to white men (Kahan et al., 2007, manuscript pp. 30–32).

That is a useful corrective to demographic categories masquerading as explanations. But cultural categories can perform the same trick. Replacing “white male” with “hierarchical individualist” does not, by itself, reveal a process of reasoning. It supplies a different classification. The scientific question is what additional observations establish the mechanism attributed to the classified person.

The paper proposes that people selectively accept or dismiss risk claims to protect the identities and social roles associated with their cultural commitments. That mechanism is plausible. It is also a hypothesis, rather than a property automatically conferred on every association between a worldview questionnaire and another questionnaire (Kahan et al., 2007, manuscript pp. 6–8). A correlation may arrive impeccably dressed. It remains a correlation.

The missing mechanism

The decisive gap is between measuring a distribution of beliefs and identifying how those beliefs were formed. The survey observes worldview responses and risk responses. It does not directly observe a sequence in which a specified identity threat changes a respondent’s evaluation of controlled evidence. Nor does it compare belief updating under experimentally varied identity conditions. The authors acknowledge that their evidence is indirect and that alternative or supplementary mechanisms remain possible (Kahan et al., 2007, manuscript pp. 9, 32–33).

Let C represent a measured cultural orientation, M identity-protective processing, and Y a risk response. The proposed explanation has this structure:

CMY.

But an association between C and Y is also compatible with some unmeasured experience or information environment, U, affecting both. Here is a deliberately simple mathematical counterexample, not an estimate of what happened in this survey:

C = aU + η; Y = bU + ε.

With mutually independent, mean-zero U, η, and ε:

Cov(C, Y) = ab Var(U).

The correlation can be substantial although this model contains no identity-protective mechanism and no direct causal effect of C on Y. The example does not prove the alternative explanation. It proves that the observed association does not uniquely select the preferred one. Causal inference requires assumptions or research designs that distinguish such possibilities; adding more covariates is not a substitute for explaining those assumptions (Hernán & Robins, 2020).

The manuscript attempts to sharpen the inference by arguing that neither expected-utility reasoning nor conventional cognitive-bias accounts predict the observed cultural differences: there is, it suggests, no reason to expect cultural groups to differ in information access or bounded rationality. But the absence of a proposed reason is not a measurement of equal information. The relevant equality is asserted, not demonstrated (Kahan et al., 2007, manuscript p. 9).

Even equal information would not settle every issue. Two people could agree on a probability of harm and disagree about the importance of the consequences or the benefits forgone. In the simplest illustration, expected loss is pL: the same probability, p, combined with different valuations of loss, L, produces different levels of concern. This matters especially when an outcome measures anxiety or a net policy assessment rather than a numerical probability.

None of this renders observational research worthless. It defines the claim that this observational design can carry. The paper supplies evidence compatible with identity protection. It does not isolate that mechanism from all the alternatives capable of producing the same pattern. Compatibility is an invitation to investigate, not a certificate of exclusive ownership.

A biography is not a rebuttal

There is a further distinction that no regression can abolish: the origin of a belief is not identical to its truth. Someone may defend a proposition for poor reasons and nevertheless defend a true proposition. Someone else may approach a question with admirable intentions and reach a false conclusion. The relevant inquiry into accuracy still requires evidence about the proposition itself.

To use a suspected motive as a substitute for examining an argument is to confuse explanation with refutation. “Your group benefits from that conclusion” may identify a reason to scrutinise the reasoning. It is not the missing scrutiny. The same standard applies to the researcher. An account of an author’s professional incentives would not, by itself, refute the author’s statistical results.

This is also why the paper’s “individualism” must not be confused with independent judgment. Its scale measures attitudes towards government, markets, personal responsibility, and collective provision. It does not directly test whether respondents arrived at those attitudes by examining evidence or by copying their companions (Kahan et al., 2007, manuscript pp. 45–46). A person can recite a defence of free markets as mechanically as another recites a demand for regulation. A borrowed opinion does not become independent because its vocabulary includes freedom.

Intellectual independence is a discipline, not a membership card. It requires accepting an unwelcome fact when the evidence warrants it and rejecting a congenial claim when the evidence does not. Anyone who abandons that discipline to preserve a group’s approval has surrendered judgment. Anyone who diagnoses that surrender without sufficient evidence has merely found a more elaborate way to avoid exercising their own.

Before reading minds, decide which way the scale runs

The first reproducibility problem requires no subtle philosophy. Appendix B gives the response coding as 1 for strong agreement through 4 for strong disagreement. It instructs the reader to reverse the egalitarian items in the hierarchy scale. Yet the methods say that higher final scores indicate greater hierarchy (Kahan et al., 2007, manuscript pp. 13, 44–45).

Follow those instructions literally. A respondent who strongly agrees with a hierarchical statement receives 1. A respondent who strongly disagrees with an egalitarian statement receives 4, which becomes 1 under the standard reversal, xreversed = 5 − x. A consistently hierarchical respondent therefore receives a low score, not the high score described in the methods. The individualism scale has the corresponding problem when communitarian items are reversed. Environmental-risk responses also require an additional change of direction to match the stated interpretation (Kahan et al., 2007, manuscript pp. 13–14, 44–46).

An omitted final reversal could reconcile the descriptions. So could an error in the appendix’s response labels. A consistent reversal of a completed scale can be a harmless change of coordinates when the analysis and interpretation are adjusted accordingly. The manuscript therefore does not establish that respondents were actually miscoded. It establishes that its printed instructions do not reconstruct the stated variables without an undocumented step. That is a reproducibility defect with a clear remedy: publish the scoring code.

The item count is similarly unsettled. The methods describe 32 worldview items, while the appendix lists 14 hierarchy–egalitarianism items and 17 individualism–communitarianism items: 31 in total. The difference could be a mistaken count or a missing item. Without clarification, the reader cannot know whether the documented instrument is the complete instrument used to construct the reported scales (Kahan et al., 2007, manuscript pp. 13, 44–46).

Note. Based on Kahan et al. (2007), supplied manuscript pp. 13–15, 21–26, and 44–46. These are documentation and reporting findings, not allegations about intent.

The political coefficients warrant particular attention because this is not a matter of a missing comma. The text says conservatism predicts viewing guns as safer and Democratic affiliation predicts viewing them as more dangerous. Table 3 prints the opposite signs, given the stated direction of the dependent variable. Controls cannot explain a contradiction between a regression table and prose purporting to describe that same table (Kahan et al., 2007, manuscript pp. 14–15, 23–24).

Such defects need not destroy the underlying findings. They do mean that the reader should not be asked to choose whichever interpretation most flatters the argument. The questionnaire, the code, and the prose must refer to the same analysis. A theory of selective interpretation should be especially reluctant to require selective interpretation of its own tables.

The null hypothesis has not confessed

The manuscript’s power discussion makes the more serious statistical move. It argues that the sample can detect small effects at a significance criterion of .01 and concludes that nonsignificant findings cannot properly be attributed to Type II error. The paper subsequently treats nonsignificant demographic coefficients as confirming that particular differences have been explained entirely (Kahan et al., 2007, manuscript pp. 16, 23, 30).

Power is not immunity from error. It is the probability of rejecting a specified null under a specified alternative, design, and analysis. Even a test with 95% power will fail to reject approximately 5% of the time under that alternative. Its power can be much lower against smaller effects. Nor does adequate power for one association automatically establish adequate power for every interaction and demographic contrast (Cohen, 1988).

The logical distinction is elementary:

p > α does not imply β = 0.

A nonsignificant result does not demonstrate that the null is true. To support a claim that a remaining difference is negligible, the analyst needs to define a substantively meaningful margin, δ, and test whether effects outside that margin can be excluded. Equivalence testing addresses that question. Merely failing to reject an exact zero addresses a different one (Lakens, 2017).

Figure 1 gives an explicit illustration. Assume independent observations from a setting where the usual Fisher-transformation approximation for a correlation applies. With N = 1,844 and a two-sided .01 test, the approximate power is 95.8% when the population correlation is .10, but only 33.4% when it is .05. These are not estimates of the paper’s regression power. They show why the existence of high power against one alternative cannot exclude missed effects against another.

The calculation is reproducible without access to the survey. Let μ = √(N − 3) atanh(ρ), let Φ be the standard normal cumulative distribution function, and let c = Φ−1(.995). Approximate power is 1 − Φ(c − μ) + Φ(−c − μ). The assumptions are explicit because a numerical illustration should not acquire a misleading air of empirical discovery.

There is a related error to avoid: a coefficient being significant in one model and nonsignificant in another does not itself establish a statistically significant difference between those coefficients. Gelman and Stern (2006) explain this distinction. The paper does estimate explicit interactions, so it would be inaccurate to say it never tests differences. The criticism concerns the stronger claims attached to the remaining nonsignificant demographic terms.

The consequence is substantial. “We did not detect a remaining difference in this specification” cannot be upgraded into “this subgroup accounts for the difference in its entirety” without further evidence. The word entirely is doing work that the statistical test has not performed. An absent asterisk is not a certificate of nonexistence.

Variance does not become additive because the adjectives do

The tables report semipartial correlations rather than ordinary regression slopes. In a conventional least-squares analysis of a single dataset, the square of a predictor’s semipartial correlation equals the decrease in R2 when that predictor is removed from the full model. It describes a unique contribution conditional on the other predictors (StataCorp, n.d.).

For hierarchy, H, and individualism, I, the separate contributions are:

srH2 = R2full − R2without H;

srI2 = R2full − R2without I.

The joint contribution is a different comparison:

ΔR2H,I = R2full − R2without both H and I.

In general, the sum of the first two quantities is not the third. The sum can be reported as a chosen summary of unique contributions. It should not be silently equated with the additional variance accounted for by entering the two variables together. Correlated predictors and suppression make the distinction consequential.

Note. Author calculations from the rounded values in Kahan et al. (2007), manuscript Tables 2–4, pp. 21, 24, and 28. The manuscript uses multiple imputation; its pooling details are insufficient to reconstruct the exact underlying quantities. The displayed differences are approximate, not new respondent-level estimates.

The manuscript’s claim that cultural measures explain approximately 34 times as much gun-risk variance as education is closely reproduced by (.218² + .204²)/.051² = 34.27. The arithmetic is intelligible. The interpretive difficulty is presenting a sum of individual unique contributions as though its meaning were interchangeable with the joint block increment. The latter is approximately .13 in the reported models, not .0891 (Kahan et al., 2007, manuscript pp. 24–25).

Notice the direction of the discrepancy. The joint increment exceeds the sum for environmental and gun responses, but is smaller for abortion responses. The criticism is not that a single arithmetic device uniformly inflates culture’s importance. It is that different statistical quantities need different names and interpretations. Precision matters even when the correction makes a favoured effect larger.

The reported relative improvements in fit are broadly arithmetically sound: moving from .23 to .28 is a 21.7% relative increase, .24 to .37 a 54.2% increase, and .21 to .23 a 9.5% increase. But these are increases of 5, 13, and 2 percentage points in sample outcome variance accounted for. They are not percentages of opinion caused by a psychological mechanism. A percentage is not a causal argument with a smaller typeface.

The disappearing effect and the disappearing controls

The subgroup argument also depends on models that change more than the addition of a supposedly decisive variable. In the final environmental specification, the hierarchical-white-male indicator appears while the continuous hierarchy measure and gender–worldview interactions disappear; R2 falls from .28 to .26. The final gun specification substitutes two white-male subgroup indicators for the continuous worldview terms and interactions; fit falls from .39 to .27 (Kahan et al., 2007, manuscript Tables 2–3, pp. 21, 24).

A simpler model is not inherently illegitimate. Nor is a lower R2 proof that it should be rejected. But this is not a clean experiment in adding one explanatory control while holding the rest of the specification constant. The disappearance of a demographic asterisk cannot be credited entirely to the new subgroup indicator when other terms have also been removed.

The restriction becomes clear if the indicator is written as D = 1 for white men above the hierarchy median and 0 otherwise. A model containing D but excluding a general hierarchy term gives that threshold a role specifically for white men. It does not estimate the corresponding threshold effect for other groups. This may be a useful restricted model, but it cannot independently demonstrate the truth of restrictions that it has imposed. The remedy is to compare it with a sufficiently general specification and test the relevant contrasts.

Interactions create a second interpretive problem. Consider an illustrative model for gender and worldview:

E[Y | F,H,I] = a + bFF + bHH + bII + bFHFH + bFIFI.

The female–male difference at particular worldview values is bF + bFHH + bFII. Testing the standalone bF does not test whether gender differences vanish throughout the worldview range. Unless the scales were centred, which the manuscript does not document, that standalone coefficient refers to both worldview scores being zero, outside their stated 1–4 range. With centring, it still refers to a specified reference point rather than every respondent (Brambor et al., 2006; Kahan et al., 2007, manuscript pp. 13, 15, 21–23).

The paper’s printed semipartial correlations cannot simply be substituted for those regression slopes. Verification needs the unstandardised coefficients and their covariance matrix. For a contrast expressed as c′β, its uncertainty depends on c′Var(β)c, including covariances. A collection of isolated significance stars does not provide that information.

There is a specific overstatement in the gun discussion. It interprets the interaction findings as showing that individualism produces greater risk scepticism among white men than among white women and minority groups. Yet the displayed Female × Individualism term is .002 and nonsignificant. With whites as the racial reference group, that term is the relevant female–male individualism-slope contrast among whites. A significant three-way interaction elsewhere does not supply the missing pairwise result (Kahan et al., 2007, manuscript pp. 24–25).

The appropriate conclusion is narrower: some conditional relationships differ, and the precise differences require explicit estimation. “An interaction exists” does not mean “every comparison in the preferred narrative has been established.” Statistical significance is not a transferable credential.

What the scales actually measure

The gun index deliberately combines judgments about accidents, defensive benefits, crime, and emotional concern about insufficient or excessive regulation. Its stated interpretation is a net assessment of whether gun ownership reduces public safety. That may be useful for studying public controversy. It is not the same thing as a calibrated numerical estimate of danger (Kahan et al., 2007, manuscript pp. 14–15, 46).

Consider just the two opposed concern items. On a conceptual recoding where larger numbers mean greater concern, call fear of insufficient regulation A and fear of excessive regulation B. Combining the first with the reversed second produces [A + (5 − B)]/2. A person scoring 4 on both obtains 2.5. A person scoring 1 on both also obtains 2.5.

Those people differ dramatically in overall concern, yet make identical contributions from that pair to the net-position score. This illustration concerns the pair, not the entire six-item index. It does not invalidate a net-position measure. It shows why movement along that measure should not be casually translated into movement from fear to fearlessness. A balance between competing concerns is not their total magnitude.

The abortion outcome is narrower: agreement with the statement that women obtaining abortions put their health in danger. It specifies no comparator, numerical probability, time horizon, procedure, or definition of harm. One respondent could understand danger as any nonzero risk; another could understand it as substantial risk relative to an alternative. The answers can differ without the respondents making different assessments of precisely the same proposition (Kahan et al., 2007, manuscript pp. 15, 47).

The environmental scale averages concern about general pollution, global warming, and living near a nuclear plant. A relationship with that average does not establish the same relationship with each component. Item-level analyses would show whether the association is general or disproportionately driven by a particular hazard (Kahan et al., 2007, manuscript pp. 14, 46).

There is a related modelling choice. Treating four ordered response categories as numerical outcomes assumes useful meaning in their assigned spacing and in the specified conditional-mean relationship. That is not automatically wrong. It calls for sensitivity checks, particularly for the single abortion item, using models that respect the ordering without requiring the same interval interpretation. This is a proposed robustness check, not a demonstrated failure of the reported estimates.

The reported worldview alpha coefficients, .77 and .81, do not settle construct validity. Internal consistency does not establish that a scale measures identity rather than related political attitudes, or that its items express one unambiguous dimension (Kahan et al., 2007, manuscript p. 13; Sijtsma, 2009). A questionnaire can be consistently measuring something other than the mechanism its title invites us to imagine.

Cross-group interpretation requires another check: whether the measures function comparably across the groups being compared. The hierarchy items include judgments about racial discrimination, women’s rights, family roles, and redistribution. The same numerical score need not automatically represent an identical underlying orientation across respondents with different experiences. Measurement-invariance analysis addresses the conditions under which such comparisons are defensible; the manuscript does not report such an analysis (Meredith, 1993; Kahan et al., 2007, manuscript pp. 44–45).

Finally, dividing respondents at sample medians creates categories; it does not demonstrate naturally occurring cultural clusters. People on opposite sides of a cutoff can have nearly identical scores. The paper uses continuous scales in its main regressions, so the entire analysis should not be condemned as dichotomised. The concern applies particularly to its named subgroups and final dummy-variable specifications (MacCallum et al., 2002; Kahan et al., 2007, manuscript pp. 13–14). A median is a convenient knife, not a discovery that humanity has a seam.

The sample, the missing observations, and the missing documentation

The African-American oversample is defensible: it improves the opportunity to study a group that would otherwise have fewer observations. But it changes the composition of the sample. Appendix A reports 407 Black respondents. That is approximately 22.1% of 1,844; removing the specified oversample leaves 165 out of 1,602, approximately 10.3%. These are calculations from the manuscript’s counts, not comparisons with an external population benchmark (Kahan et al., 2007, manuscript p. 43).

The paper explicitly excludes the oversample from its white-men-versus-everyone-else cultural-group comparisons, to avoid overweighting African-American responses in the latter category. Its regression tables nevertheless report the full sample. The weighting and design-adjustment procedures are not sufficiently documented to reconstruct the treatment of all reported statistics (Kahan et al., 2007, manuscript pp. 18–19, 21, 24, 28).

This is not a basis for declaring every unweighted coefficient biased. Weighting depends on the target quantity, selection mechanism, and model. Population means, conditional regression relationships, and sample fit statistics are not interchangeable targets. The appropriate demand is that the target be stated and the weighting decision justified, with sensitivity analyses where relevant (Solon et al., 2015).

The response rate is reported as 42%, with a 59% cooperation rate. A low response rate alone does not prove bias: what matters is how participation relates to the quantities being estimated. But representativeness is not established simply by calling a sample broadly representative. A nonresponse analysis would help determine whether missing participants plausibly differ in ways relevant to worldview–risk relationships (Groves & Peytcheva, 2008; Kahan et al., 2007, manuscript p. 43).

Missing questionnaire data present a separate issue. The manuscript reports chained-equations multiple imputation, five completed datasets, and pooling under established procedures. It does not provide variable-level missingness rates, detailed imputation models, convergence diagnostics, or a clear account of how the interaction structure was handled (Kahan et al., 2007, manuscript p. 16).

Multiple imputation is not fabrication. It is a principled way to represent uncertainty under specified assumptions. Its validity depends on those assumptions and the models used. Interactions in the substantive analysis need compatible treatment in the imputation process; otherwise, the relationships of interest can be distorted. Later methodological work explains this problem in detail (Bartlett et al., 2015; White et al., 2011).

Nothing in the manuscript proves that the authors implemented imputation incorrectly. Nor is five imputations automatically inadequate. The problem is that the amount of missing information and the procedure’s stability are undocumented, so adequacy cannot be independently assessed. The honest conclusion is an unresolved replication question, not an invented numerical estimate of the resulting bias.

Finally, the total sample size is not the effective information available for every contrast. There are 153 Black men and 254 Black women before further subdivision by worldview, and the cultural-group table supplies no individual cell counts or standard errors. Strong assurance based on 1,844 respondents cannot simply be inherited by every three-way interaction. The relevant evidence includes subgroup distributions, overlap, measurement precision, and uncertainty around the actual contrast (Kahan et al., 2007, manuscript pp. 19, 43).

A theory must risk losing

The gun interaction specification contains twelve displayed interaction terms, and other analyses add further comparisons. The discussion treats some results with p = .07 as borderline support. That does not make them meaningless, but it makes selective emphasis consequential. As an illustration only, twelve independent tests of true null hypotheses at .05 would have probability 1 − .9512, approximately 46%, of producing at least one false positive. The actual tests are correlated, so this is not the paper’s false-positive probability (Kahan et al., 2007, manuscript pp. 24–25, 29).

The point is not to erase every strong association with a ritual invocation of multiple testing. It is to distinguish robust general patterns from a detailed subgroup story assembled partly from marginal results. The paper states hypotheses in advance of presenting results, but the manuscript alone does not provide a time-stamped analysis plan. A reanalysis should disclose the testing family, distinguish confirmatory contrasts from exploration, and report how multiplicity affects the narrower claims.

The abortion findings also contain an acknowledged anomaly: African-American respondents remain more concerned after the relevant cultural controls. The authors offer a possible identity-protective explanation involving respectability and responses to racial stereotypes, expressly labelling it conjectural (Kahan et al., 2007, manuscript pp. 33–34).

That acknowledgement deserves credit. A new explanation prompted by an anomaly can be scientifically useful. It cannot count as an additional successful prediction of the original test. Otherwise, disappearance of a demographic association confirms the mechanism, persistence confirms a newly elaborated mechanism, and the theory becomes incapable of embarrassment. A theory that can explain every possible result after lunch has not necessarily predicted anything before breakfast.

From explanation to political permission

The normative implications require the same discipline. The paper considers whether some risk evaluations might be rejected because they express supposedly inappropriate hierarchical or individualistic norms. It also presents the opposite concern: regulation might impose an egalitarian or communitarian orthodoxy. The authors explicitly take no position between these possibilities (Kahan et al., 2007, manuscript p. 35).

My objection is to any inference that treats the proposed psychological origin of a preference as sufficient permission to override a person’s freedom. A descriptive association cannot, by itself, establish a principle of rights or legitimate coercion. Those questions require their own arguments. An official’s confidence that a citizen has the wrong cultural attachments does not make the official’s judgment self-validating.

The symmetry is essential. A business executive cannot dismiss evidence of harm merely because regulation is unwelcome. A regulator cannot dismiss a sound objection merely because the objector values markets. Intellectual independence demands allegiance to facts, not to the social respectability of whichever faction currently occupies the meeting room.

It would also be unfair to accuse this paper of denying objective truth. Its communication proposals aim to make empirically sound information easier to accept by reducing identity-based resistance. That is a defensible research objective. The survey itself does not establish the effectiveness of the proposed interventions, and successful persuasion would still need to be distinguished from improved accuracy (Kahan et al., 2007, manuscript pp. 35–36).

For the question suggested by the title When Loyalty Follows the Office, identity protection is therefore one candidate explanation, not a diagnosis delivered by this study. A person may recognise an office-holder’s authority while rejecting that person’s factual claims. Public compliance can also differ from private belief. Investigating loyalty requires separating those possibilities. An office can confer responsibility; it cannot confer infallibility. That applies as readily to an academic office as to any other.

What would repair the argument?

A proper repair begins with reproducibility rather than rhetoric: the original questionnaire, complete coding key, item-level data under appropriate privacy protections, sampling information, missingness indicators, and executable scoring, imputation, and analysis code. These materials would resolve whether the direction-of-coding problem is merely documentary, which item count is correct, and which table signs correspond to the actual analysis.

The next step is to reconstruct the estimands. Keep the continuous worldview measures; state the comparisons each model is intended to estimate; publish unstandardised coefficients and covariance matrices; and calculate demographic contrasts over observed worldview ranges. Test the restrictions implied by subgroup-only explanations rather than relying on coefficients that lose significance when the specification changes. Claims that an effect is negligible require justified equivalence bounds.

The measurement analysis should separate factual likelihood judgments from emotional concern, moral evaluation, and perceived benefits. Report item-level results alongside indices, examine comparability across demographic groups, and check sensitivity to alternative models for ordered responses. Weighting, nonresponse, and imputation should be assessed against explicit targets and assumptions rather than treated as invisible technical preliminaries.

Finally, the psychological mechanism needs a design that can discriminate it from alternatives. A prospective experiment might present common evidence while randomly varying an identity-relevant framing or affirmation, measuring baseline beliefs, comprehension, subsequent updating, and accuracy on questions with a defensible evidential benchmark. It would need to distinguish identity protection from changes in source credibility, interpretation, or information content. Randomisation of a message is not, by itself, proof of why that message works.

Those are proposed studies and checks, not results already obtained. Some may strengthen the paper’s account; others may narrow it. That is precisely their value. A research programme should be organised around observations that could change the conclusion, not around increasingly sophisticated ways to preserve it.

Table 3. What the problems change—and what they do not establishIssueConsequence for the argumentLimit of this auditCausal mechanism not isolatedAssociations support a hypothesis, not a unique explanation of belief formation.Does not show that identity protection is absent.Nonsignificance treated as complete explanationClaims of no remaining demographic effect require different evidence.Does not establish a particular residual effect size.Scoring and reporting inconsistenciesDirections and reconstruction require correction or clarification.Does not establish actual miscoding or misconduct.Different variance summaries conflatedRelative-importance claims need an explicit statistical definition.Does not imply uniform exaggeration across outcomes.Changed controls and insufficient subgroup contrastsExclusivity claims are not secured by disappearing significance stars.Does not prove every interaction is spurious.Measurement, sampling, and imputation gapsRobustness and generalisability require further analysis and documentation.The direction and magnitude of any resulting bias are unknown.

Note. This table summarises the manuscript audit and the mathematical distinctions developed above. It is not a set of new empirical findings.

The answer to Bailey

My answer is neither that the paper should be worshipped nor that it should be discarded. It reports associations worth investigating and a mechanism worth testing. It also contains documentation defects, an invalid inference from nonsignificance to absence, inadequately distinguished variance calculations, and conclusions about complete subgroup explanation that exceed the displayed evidence.

The most defensible surviving claim is that the study’s cultural-attitude measures are associated with its self-reported risk measures, with some demographic variation in those relationships. That is a substantive finding as reported. It is not equivalent to proving that identity protection uniquely produced the beliefs, that a particular subgroup explains every remaining demographic difference, or that the respondents were therefore mistaken about the risks.

The discipline required here is not hostility to psychology. It is refusal to let psychology become an exemption from logic. People can protect an identity instead of examining a fact. Researchers can also protect an explanation instead of testing its limits. Neither possibility licenses a diagnosis in the absence of evidence.

A person who sacrifices judgment to belonging has surrendered independence. A scholar who substitutes a story about belonging for an examination of the argument has not restored it. The obligation on both sides remains the same: establish what was measured, calculate what the numbers mean, and stop where the evidence stops.

A correlation is not a confession. The regression has not read anyone’s mind.

References

Bartlett, J. W., Seaman, S. R., White, I. R., & Carpenter, J. R. (2015). Multiple imputation of covariates by fully conditional specification: Accommodating the substantive model. Statistical Methods in Medical Research, 24(4), 462–487. https://doi.org/10.1177/0962280214521348

Brambor, T., Clark, W. R., & Golder, M. (2006). Understanding interaction models: Improving empirical analyses. Political Analysis, 14(1), 63–82. https://doi.org/10.1093/pan/mpi014

Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates. https://doi.org/10.4324/9780203771587

Gelman, A., & Stern, H. (2006). The difference between “significant” and “not significant” is not itself statistically significant. The American Statistician, 60(4), 328–331. https://doi.org/10.1198/000313006X152649

Groves, R. M., & Peytcheva, E. (2008). The impact of nonresponse rates on nonresponse bias: A meta-analysis. Public Opinion Quarterly, 72(2), 167–189. https://doi.org/10.1093/poq/nfn011

Hernán, M. A., & Robins, J. M. (2020). Causal inference: What if. Chapman & Hall/CRC. https://miguelhernan.org/whatifbook

Kahan, D. M., Braman, D., Gastil, J., Slovic, P., & Mertz, C. K. (2007). Culture and identity-protective cognition: Explaining the white-male effect in risk perception. Journal of Empirical Legal Studies, 4(3), 465–505. https://doi.org/10.1111/j.1740-1461.2007.00097.x

Version audited: The supplied manuscript revised on 14 June 2007, available through SSRN at https://ssrn.com/abstract=995634. All manuscript-specific page references, table checks, and reconstructions in this article concern that version.

Lakens, D. (2017). Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science, 8(4), 355–362. https://doi.org/10.1177/1948550617697177

MacCallum, R. C., Zhang, S., Preacher, K. J., & Rucker, D. D. (2002). On the practice of dichotomization of quantitative variables. Psychological Methods, 7(1), 19–40. https://doi.org/10.1037/1082-989X.7.1.19

Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance. Psychometrika, 58(4), 525–543. https://doi.org/10.1007/BF02294825

Sijtsma, K. (2009). On the use, the misuse, and the very limited usefulness of Cronbach’s alpha. Psychometrika, 74(1), 107–120. https://doi.org/10.1007/s11336-008-9101-0

Solon, G., Haider, S. J., & Wooldridge, J. M. (2015). What are we weighting for? Journal of Human Resources, 50(2), 301–316. https://doi.org/10.3368/jhr.50.2.301

StataCorp. (n.d.). pcorr—Partial and semipartial correlation coefficients [Software documentation]. https://www.stata.com/manuals/rpcorr.pdf

White, I. R., Royston, P., & Wood, A. M. (2011). Multiple imputation using chained equations: Issues and guidance for practice. Statistics in Medicine, 30(4), 377–399. https://doi.org/10.1002/sim.4067


← Back to Substack Archive