The Governance of Belief
On knowing what you believe and changing the right part by the right amount.
There is a sentence in my archive I would not write today. It took the behaviour of one cohort in one semester and spent it as a verdict on where higher education was heading. I remember the confidence. What I cannot reconstruct is any accounting that followed when the confidence failed. Which part of the claim died? Which part survived? How sure am I now of whatever survived? Nothing so orderly took place. The belief faded the way most beliefs go: unexamined and never formally retired.
Changing your mind and knowing what changed are different achievements, and everything in this piece follows from the difference.
We praise critical thinking constantly and define it rarely. When institutions get specific, what they describe is mostly analysis: evaluating arguments and weighing sources. These are real skills and worth teaching. But notice what they share. They are performed on other people’s reasoning. The student dissects an argument somebody else made, about a claim somebody else holds, with consequences somebody else bears. Their own beliefs sit outside the exercise the whole time, ungoverned.
Here is a definition that puts them back in. Critical thinking is the disciplined, self-correcting governance of belief and action under uncertainty. In plainer words: knowing what you believe, why you believe it, how sure you are, what would count against it, and changing the right part by the right amount when the evidence warrants it.
Governance is the load-bearing word, and I mean it in the unglamorous committee sense. Anyone who has sat through an audit and risk meeting knows what governance looks like when the object is money: thresholds, review dates, delegations, a record of who decided what on what basis. Universities govern budgets, buildings, research data and reputational risk with elaborate care. Beliefs run on trust. There is no review date on a conviction.
Five clauses, then. Each sounds obvious. Each names a capacity the evidence says people mostly lack. The sentence is easy to nod along to, and the nodding is the problem.
Knowing what you believe sounds like the free one. It is the strangest of the five.
In 2005, researchers at Lund University showed people pairs of photographs and asked them to pick the more attractive face. Using sleight of hand, the experimenters then handed back the rejected photograph and asked each person to explain their choice. Most swaps went unnoticed. People gave fluent, confident reasons for a choice they had not made: I liked the earrings, said of a face a card trick had picked for them. The effect has since been reproduced with the taste of jam and with political opinions.
Psychologists had suspected something like this for decades. In 1977 Richard Nisbett and Timothy Wilson reviewed years of experiments and concluded that when people report on their own mental processes, they are often not inspecting anything. They are composing: producing a plausible story about what someone like them would think, then mistaking the story for a memory.
So the first clause is already in trouble. Knowing what you believe feels like reading your own filing system. It behaves more like drafting a statement on your own behalf, and statements can be drafted badly. The honest version of “I believe this” carries a small rider: to the best of my knowledge of myself, which is patchier than I assume.
Why you believe it hides two questions inside one word. There is the history: the causes that got the belief into you. And there are the grounds: the reasons that would justify holding it now. These come apart more often than is comfortable. Jonathan Haidt’s studies of moral judgement found that the verdict tends to arrive first and the reasons afterwards, the way a press secretary briefs the room once the decision has been made elsewhere. Ziva Kunda showed that when we want a conclusion, we search memory the way a lawyer searches case law: for support. Ask someone why they believe something and the fluent answer you receive may be the brief, not the journey.
The tradition that takes grounds seriously runs back through Hume, who wrote that a wise man proportions his belief to the evidence. Its fiercest statement belongs to W. K. Clifford, a Victorian mathematician who opened an 1877 essay with a shipowner. The man knew his emigrant ship was old and often repaired. Doubts about her seaworthiness surfaced, and he worked on himself until they dissolved. She had weathered so many voyages. Providence would hardly abandon the families aboard. The refit could wait. He achieved sincerity, watched her leave port and collected the insurance money when she went down mid-ocean. Clifford’s verdict was that the sincerity changes nothing, because the man had no right to believe on the evidence in front of him. The belief had passengers.
William James answered a generation later that some questions are live, forced and momentous, that they will not wait for the evidence to arrive, and that on those questions the heart may decide. The argument between them has run for a century and a half, and this essay does not need to settle it. It needs only the premise both men shared: beliefs have grounds, grounds can be adequate or not, and you are answerable for the difference.
How sure you are turns out to have a scoreboard, which surprises people. The measure is called calibration, and the plain version is this: of all the things you say you are seventy per cent sure of, roughly seventy per cent should turn out true. Decades of research find people pervasively overconfident, and worst on hard questions.
Then somebody checked whether experts do better. Philip Tetlock spent twenty years collecting predictions from political and economic specialists and found the average expert performing at a level he compared, in a joke that has outlived the book, to a dart-throwing chimpanzee. The follow-up carried the better news. In forecasting tournaments run for American intelligence agencies, a small fraction of ordinary volunteers proved reliably, measurably better, and their habits turned out to be teachable. They gave estimates in single percentage points and meant the grain: when their forecasts were rounded to the nearest ten, their accuracy scores got worse, because the difference between sixty-three and sixty-seven was information. And they moved in small, frequent steps rather than rare lurches, adjusting a few points on each new scrap of evidence.
Sureness, on this account, is an estimate that can be trained and scored. There is even a golf score for it, the Brier score, which punishes you in proportion to how far your stated probability sat from what happened. Almost nobody’s education included an hour of this.
What would count against it is Karl Popper’s question, and it is the clause that separates holding a belief from being held by one. A belief maintained well comes with tripwires specified in advance: the observations that would trigger review.
Two different alarms need telling apart. You believe the meeting is at ten because the calendar says so. A colleague tells you it moved to eleven: that is evidence you are wrong. Then you notice the calendar has been displaying yesterday’s schedule all week: that is evidence your reason was never any good, and it is the sneakier alarm, because it does not tell you when the meeting is. It only tells you that you no longer know.
We do not set these tripwires naturally. Peter Wason’s card experiments in the 1960s showed that people asked to test a rule reach for the cards that could confirm it and leave untouched the one that could break it. The corrective with the best evidence behind it is almost embarrassingly simple. In 1984 Charles Lord and colleagues found that telling people to be fair and unbiased changed little, while instructing them to consider the opposite, to ask how they would rate this same study had its results pointed the other way, measurably reduced bias. Jonathan Baron built a research programme around the underlying disposition, actively open-minded thinking: the practised habit of searching for reasons you might be wrong. His own caveat is the honest one. People endorse the habit on questionnaires and abandon it on the beliefs that carry their identity.
Changing the right part by the right amount is the clause doing the most work, and it contains two separate problems: which part, and how much.
Which part first. Beliefs do not stand alone. W. V. O. Quine’s image was a web: experience touches only the edges, and when reality contradicts you, it contradicts the whole arrangement at once, without specifying which strand to cut. Suppose the usage figures for an expensive library database fall by a third. Several beliefs are implicated together: that researchers value the resource, that the counting is accurate, that the discovery system still routes people to it, that the field it serves remains active on this campus. The falling number contradicts the conjunction and is silent about the members. Something must give. The evidence will not tell you what.
Quine’s advice was minimum mutilation: change as little as you can and still fit the facts. Sensible, and also the loophole, because you can always protect the belief you love by amputating something cheaper. The philosopher Imre Lakatos supplied a test for when that protection is honest work and when it is rot. Watch what the repair does next. If the adjusted story predicts something new that could be checked, the belief is earning its survival. If the adjustments only ever explain why the old conclusion should stand, the thing being defended has stopped being answerable to anything.
You can watch the degenerate version run in any organisation. A company orders everyone back to the office, citing productivity. The productivity evidence turns out to be contested, and the mandate stays, now citing collaboration. Then culture. Then the mentoring of juniors. Each move can be defended on its own. The pattern is a rule whose justification migrates whenever it is threatened, which is how a policy becomes unkillable. The Lakatos question cuts through it: would the surviving reason, on its own, justify this rule at this cost, and what future evidence could ever narrow or end it?
Then how much. Here the findings are old and awkward. In the 1960s Ward Edwards ran experiments with bags of poker chips and found that people revise in the right direction but at a fraction of the warranted rate. By his estimate it took between two and five observations to produce one observation’s worth of movement. The modern refinement is sharper and stranger. A 2025 study in the Quarterly Journal of Economics, pooling a large body of updating experiments, found that we overreact to weak evidence and underreact to strong evidence. We move too far on a rumour and not far enough on a result.
And zero is a legitimate amount. A weather vane moves with every gust and a fanatic with none, and neither is thinking. What the clause names is discrimination: moving when the evidence warrants, by the amount it warrants, and holding still against noise. Some of the best updates are refusals, made for reasons you could state.
One word of the definition is still unexplained: action. Belief and action run at different thresholds. Belief should track the evidence; action must also track the stakes, and what it would cost to be wrong. A cheap, reversible step can be rational at modest confidence. An expensive step, or one imposed on other people, should demand more, and an irreversible one most of all. You can hold a claim at fifty-fifty and still run a small trial, provided it stays a trial, with an end date and something that would count as failure. What fifty-fifty cannot license is a permanent rule wearing the costume of settled science.
That is the machinery, and its relation to a more fashionable virtue needs stating precisely. Intellectual humility is having a research boom, complete with measurement scales and training interventions. The disposition is real, and I want much much more of it in public life. But a disposition is a readiness, and readiness is cheap. Humility says I may be wrong. The five clauses establish whether you are, by how much, and what follows. Without them, humility degrades into performance: the leader who allows that mistakes were made, sounds appropriately chastened and changes nothing. The machinery running looks different. It specifies which claim failed and states the confidence that remains. It revises the public rationale and reviews the rule the belief had authorised. A mind can be humble while everything it governs stays exactly where it was.
Which brings me to education, my industry, and the reason this definition unsettles me.
An essay handed in at midnight is a photograph of thought at the instant the deadline froze it. Two students can submit the same conclusion. One began there and spent the week armouring it. The other crossed the question three times and came back with better reasons. The finished product cannot tell them apart. Nor can the tests. As far as I can establish, none of the widely used standardised critical thinking assessments introduces new evidence partway through and scores what you do with it. I hold that claim at about ninety per cent. It is a claim about an absence, and absences resist proof. What the tests do contain is easier to state: analysis of supplied arguments, evaluation of supplied inferences. The fifth clause, the update, is the one thing the format never sees. The mark goes to whatever position was left standing when the clock ran out.
Two findings make this gap expensive rather than merely untidy. Keith Stanovich and colleagues measured myside bias, the tendency to rate evidence more kindly when it flatters your side, and found it shows almost no relationship with intelligence. Nearly every other bias in the literature shrinks as cognitive ability rises. This one does not. Intelligence upgrades the lawyer. And Hugo Mercier and Dan Sperber have argued that reasoning evolved for argument in the first place: it is built to produce justifications and to scrutinise other people’s, which is why your reasoning flatters you and is sharp about mine. Two consequences follow. Cleverness will not save us. And the machinery runs best in company, against people licensed to disagree.
There is one irony I cannot leave out, given where this post lives. The training loop inside a large language model is proportional error correction: every internal weight nudged in proportion to its contribution to the mistake, billions of times over. Changing the right part by the right amount is the engineering. We built the discipline into the machines and left it out of the curriculum.
I began with a sentence I would no longer defend, and I admitted that no accounting followed. Let me at least start one here. The clause that failed was the amount. I moved a long way on a weak signal, one semester, one cohort, because the pattern was elegant and I wanted it. The scale of the claim should have been a classroom and I made it a sector. What survives, at lower confidence, is the narrower observation underneath. That is the audit. It took a paragraph, and it cost something to write, which tells you why they are rare.
Because published beliefs are beliefs with readers. Clifford’s shipowner is every one of us whose convictions harden into rules, budgets, bans and marking schemes that other people live under. Once your answer governs someone else’s day, the update stops being private hygiene. It becomes an account owed to the people downstream: which claim failed, what confidence remains, whether the surviving reason would justify the rule on its own, what evidence would narrow or end it.
This essay is itself an answer handed in at a deadline. Somewhere in it is the sentence I will not defend in two years. The definition tells me what I will owe that sentence when I find it. Whether I pay is not settled by knowing.



I love this definition of critical thinking and its exactly what I’ve been trying to teach students to do, without the pithy articulation. I’m going to use this to scaffold on of my classes next semester. Thank you!
And Ta for the zoominable encrypted intricate apt visual.