Block I · Foundations of Knowledge & Reasoning · Day 011 / 180
Heuristics, Biases & Rationality
Your mind takes shortcuts. The question that started a fifty-year war: are they bugs, or are they brilliant?
Meet Linda. She is 31, single, outspoken, and very bright. She majored in philosophy. As a student she cared deeply about discrimination and social justice, and she joined anti-nuclear demonstrations. Now, quickly, which is more probable? A. Linda is a bank teller. B. Linda is a bank teller and is active in the feminist movement.
Most people feel the pull of B. It fits. It tells a story. When Amos Tversky and Daniel Kahneman ran this in 1983, more than 80% of respondents chose B, including doctoral students in decision science who had studied probability formally. But look again. Every feminist bank teller is a bank teller. The group in B sits entirely inside the group in A. A conjunction can never be more probable than one of its parts. Answer B is not just wrong; it violates a law of logic you already know. You may have just watched your own mind break a rule it agrees with. That crack has a name, the conjunction fallacy, and prying it open is today’s work.
Where we are
For ten days we have built the machinery of good reasoning: what counts as knowing (Day 1), how science filters claims (Day 2), the three engines of inference (Day 3), Bayes as the law of belief-updating and the base-rate trap (Day 4), and why every model is a useful lie (Day 10). All of that was normative: how a mind ought to reason. Today we turn the microscope on the actual instrument. Do human beings live up to the standard? And when we fall short, is that a defect to fix, or the fingerprint of a mind doing something cleverer than logic?
The program
The shortcuts we think with
In 1974, two Israeli psychologists published a paper in Science that quietly rerouted a field. Tversky and Kahneman argued that people do not estimate probabilities by doing probability. Instead we reach for a handful of heuristics: fast mental shortcuts that usually work and occasionally fail in predictable ways. Three did most of the heavy lifting.
Representativeness. We judge how likely something is by how much it resembles our mental prototype. This is exactly what snares Linda: she resembles the stereotype of a feminist so strongly that “feminist bank teller” feels like a better fit than plain “bank teller,” and the feeling of fit quietly overrides the arithmetic of sets. Representativeness also makes us ignore base rates, the trap from Day 4: told that “Steve is meek and tidy and likes order,” people guess librarian over farmer, forgetting that there are vastly more farmers than librarians to begin with.
Availability. We judge how common something is by how easily examples come to mind. After a plane crash leads the news, flying feels lethal, even though the very newsworthiness of the crash is evidence of how rare it is. Vivid, recent, emotional events are overweighted because they are easy to recall.
Anchoring and adjustment. When we estimate an unknown number, we latch onto whatever value is nearby and adjust from it, usually not far enough. Tversky and Kahneman spun a rigged wheel of fortune, let it stop on 10 or 65, then asked people what percentage of African nations are in the UN. Those who saw 65 guessed far higher than those who saw 10. The wheel was obviously random and had nothing to do with Africa. It anchored them anyway.
The unsettling part is not that we err. It is that these errors are as stubborn as optical illusions. In the Muller-Lyer illusion, two lines of equal length look unequal because of little arrowheads on their ends, and knowing the trick does not make them look equal. Kahneman’s claim was that many cognitive biases are like that: not ignorance, but the automatic output of machinery you cannot simply switch off. Understanding the fallacy does not dissolve the feeling. Kahneman won the 2002 Nobel in Economics for this line of work. Tversky had died in 1996, and the prize is not awarded posthumously.
Interactive · watch the illusion work
The Linda Machine
First, make your choice. Then flip between the two lenses your mind can use on the same question: the one that feels right, resemblance, and the one that is right, nested sets.
Example · Linda’s sets
The figure shows the essential result: the vivid description makes the smaller set feel like a better match, but “bank teller and feminist activist” remains a subset of “bank teller.”
| Lens | What it sees | Verdict |
|---|---|---|
| Resemblance | Linda fits the feminist-activist story better than the plain teller story. | B feels more plausible. |
| Set logic | Every feminist bank teller is already inside the set of bank tellers. | B cannot be more probable than A. |
The model — and its cracks
Fast and slow, handled carefully
To organize all this, Kahneman popularized a two-character drama in his 2011 bestseller Thinking, Fast and Slow. System 1 is fast, automatic, effortless, intuitive: it reads the word on a billboard, senses hostility in a voice, and blurts “feminist” at Linda before you have noticed. System 2 is slow, deliberate, effortful: it is what you use to fill out a tax form or check whether B is really a subset of A. Biases, on this telling, happen when busy System 2 fails to audit the quick answer.
It is a wonderful teaching device. It is also, taken literally, probably wrong, and the field knows it. Three cracks are worth naming, because a good scientist keeps the ladder and kicks away the scaffolding.
There may not be two of anything. The “two systems” were always more metaphor than anatomy. Even leading proponents retreated: Jonathan Evans and Keith Stanovich, in a careful 2013 paper, abandoned the idea of two systems in favor of two types of processing, a weaker and fuzzier claim. In 2018 David Melnikoff and John Bargh went further, calling the dual-process typology “a convenient and seductive myth” that “lacks empirical support” and “systematically thwart[s] scientific progress.” The tidy binary, conscious/unconscious and effortful/automatic, does not cluster into two neat bundles.
The large, general mental-fuel model did not hold up. A famous companion idea, ego depletion, held that effortful self-control runs on a limited energy reserve: resist the cookies now, and you will cave to the next temptation because your willpower muscle is tired. It launched a thousand studies. In 2016, a preregistered replication across 23 laboratories estimated a standardized effect of about 0.04, statistically indistinguishable from zero. Later preregistered multilab studies did not restore the original large, domain-general claim: a 12-lab study estimated a small effect of about 0.10, while a 36-site confirmatory test estimated a nonsignificant effect of about 0.06. The evidence therefore does not support a large, general-purpose willpower fuel; whether some small sequential-task effect remains under particular protocols is still unsettled. This is the Day 2 replication crisis reaching into today’s topic.
So hold the model loosely. Fast versus slow remains a useful shorthand, and we will keep using it, but treat it the way Day 10 taught us to treat any model: a lossy map, not the territory. The deeper action has moved elsewhere.
The debate
The Great Rationality War
Here is where the field split into two camps that have argued, productively and pointedly, for four decades.
On one side stands the heuristics-and-biases tradition of Kahneman and Tversky. Its message, roughly: human intuition is riddled with systematic error. Biases are cognitive illusions, real measurable failures against the gold standard of logic and probability. The practical upshot is meliorist: since our minds mislead us, we should build corrections such as training, checklists, expert systems, and later, nudges.
On the other side stands Gerd Gigerenzer and the ABC Research Group in Berlin, who reply: not so fast. The “errors,” they argue, are often not errors at all, and the experiments are frequently rigged, unintentionally, against the humans. They make two distinct moves, and it is worth keeping them apart.
Move one: change the format, and the fallacy fades
Gigerenzer’s first argument is about information format, a direct callback to Day 7. The same content can carry very different amounts of usable structure depending on how it is encoded. Take the base-rate problems that make people, and doctors, look hopeless. In probability format:
A disease affects 1% of women. A test detects it 90% of the time, but also gives a false positive 9% of the time. A woman tests positive. How likely is it that she is actually sick?
Most people, including many physicians in study after study, answer around 80-90%. The correct answer is about 9%. But watch what happens when you hand the mind the same facts as natural frequencies instead of abstract percentages: Of 1,000 women, 10 have the disease and 9 of them test positive. Of the 990 healthy women, about 89 also test positive. So of roughly 98 women who test positive, only 9 are sick. Suddenly the answer, 9 out of 98, about 1 in 11, is almost visible. You can count it. A 2017 meta-analysis found that switching to natural frequencies roughly quintupled correct Bayesian answers, from about 4% to about 24% of people. Gigerenzer’s point is that the base-rate fallacy is not just a fixed flaw in human wiring. It is partly an artifact of feeding the mind single-event probabilities, a format it is poorly built to digest. Change the representation and much of the irrationality dissolves.
Interactive · same facts, two formats
The Natural-Frequency Grid
1,000 women, one imperfect test. Flip between the format that confuses, probabilities, and the format the mind can see, counted people. Then drag the base rate and watch the real chance of being sick swing.
Of everyone who tests positive, the share who are actually sick is ≈ 9%.
Example · Natural-Frequency Grid
The figure uses the classic 1% case in natural frequencies, so the Bayesian answer can be counted.
| Group | Count out of 1,000 | Positive tests |
|---|---|---|
| Sick women | 10 | 9 true positives, 1 missed case |
| Healthy women | 990 | 89 false positives, 901 true negatives |
| All positive tests | 98 | Only 9 are true cases, so the posterior chance is about 9%. |
Move two: less can be more
Gigerenzer’s second, bolder argument is that simple heuristics are not merely excusable. They are sometimes better than the fancy, information-hungry methods statisticians prefer. The banner is fast-and-frugal heuristics: little rules that ignore most of the available information and still make excellent decisions.
The showpiece is the gaze heuristic. How does an outfielder catch a fly ball? Not by measuring velocity, computing a parabola, and sprinting to where the ball will land. That is intractable in real time with wind and spin. Instead the fielder fixes their gaze on the ball and runs so that the angle of gaze stays constant. Follow that one rule and you arrive where the ball comes down, no calculus required. The heuristic works not despite ignoring information but because it does: it throws away everything except the single cue that matters.
Or the recognition heuristic: asked which of two cities is larger, if you recognize one and not the other, bet on the one you recognize. Absurdly crude, and yet, because recognition tends to track size, it can beat elaborate models. In a celebrated finding, this even produced a “less-is-more effect”: people who recognized fewer of the cities sometimes scored better, because they could use the heuristic while those who recognized everything had to fall back on shakier knowledge.
Why would ignoring information ever help? The deep answer, which Gigerenzer and Henry Brighton laid out in a 2009 paper titled Homo Heuristicus, is the bias-variance tradeoff, an idea we will meet again in machine learning on Day 136. A complex model with many parameters fits the data it trained on beautifully, but it also fits the noise, so it lurches around on new data. A simple heuristic cannot fit noise because it barely fits anything. It may be biased, but it is stable, and in a noisy, small-sample world, stability often wins.
The reframe
Simon’s scissors
Underneath the whole fight is a question about the yardstick. When Kahneman says people are irrational, he means: measured against the laws of logic and probability. When Gigerenzer says they are smart, he means: measured against success in the actual environment they live in. These are different rulers, and much of the war is a disagreement about which ruler is legitimate.
The person who saw this first was Herbert Simon: economist, cognitive scientist, AI pioneer, and winner of the 1978 Nobel. Simon spent the 1950s arguing that the fantasy of a perfectly rational agent who optimizes over all options is a nonstarter for any real creature. Real minds have limited time, memory, and attention. So instead of optimizing, we satisfice: we search until we find an option that is good enough and stop. Simon called this bounded rationality, and he gave it an image that quietly settles the debate:
“Human rational behavior is shaped by a scissors whose two blades are the structure of task environments and the computational capabilities of the actor.”
A pair of scissors cuts only when both blades engage. You cannot judge a blade on its own. Kahneman and Tversky studied mostly the mind blade, and, holding it up against the ruler of logic, found it wanting. Gigerenzer insists you must look at both blades together: the mind and the world it is cutting through. Once you do, yesterday’s “bias” often reveals itself as a tool matched to its environment. Neither camp is simply wrong. They are inspecting different blades of the same scissors.
Diagram · one tool, two verdicts
Which blade are you judging with?
The same fast-and-frugal heuristic, say the gaze heuristic, gets opposite grades depending on which blade you hold it against.
Example · Simon’s Scissors
The figure keeps both blades visible: behavior is produced by the contact between cognitive processes and the task environment.
| Judged against | What a heuristic looks like | Example verdict |
|---|---|---|
| Logic alone | It ignores most of the information. | Irrational: the rule is too crude. |
| Mind plus environment | It exploits the structure of the task world. | Smart: the gaze heuristic solves the real-time catching problem cheaply. |
The frontier · 2026
Three live edges, with the hype filter on
Every day in this course ends at the research frontier, each claim tagged for how much weight it can bear. The rationality wars did not end in a treaty; they matured into sharper empirical questions.
Which biases survived the replication crisis?
The heuristics-and-biases literature was not spared the reckoning of Day 2. But the wreckage sorted itself into a revealing pattern. The perceptual-style, one-shot cognitive illusions held up; the fragile, many-moving-parts social effects often did not.
| Finding | Evidence status | What survived |
|---|---|---|
| Conjunction fallacy | Established | The Linda effect replicates readily across decades, cultures, and framings, and large language models can fall for it too. |
| Anchoring | Established | Among the sturdier effects in psychology; now paired with computational explanations. |
| Ego depletion | Contested / small | Preregistered multilab estimates range from about d = 0.04 to 0.10. The classic large fuel model is unsupported; residual effects appear small and protocol-sensitive. |
| Many social-priming effects | Contested | Some concept-priming results failed to replicate; Kahneman publicly warned the field. |
Social priming is the claim that a cue—words, images, or a social stereotype—activates an associated concept and shifts a later judgment or behavior, often without awareness. The classic example said that reading words related to old age made people walk more slowly afterward. That kind of multi-step, context-sensitive effect is exactly why the social-priming literature has been more fragile than perceptual illusions: several headline findings failed independent replication, while broader priming effects remain mixed.
The lesson is not that “biases are not real.” It is that the robust ones tend to be the ones with a mechanism, which sets up the most interesting development of all.
The peace treaty: biases as an optimally lazy mind
The most important recent idea in this area is a genuine attempt to dissolve the war rather than win it, and it lives squarely on our computation thread. It is called resource-rational analysis, developed by Falk Lieder and Tom Griffiths in a 2020 target article in Behavioral and Brain Sciences. Take Simon’s insight seriously and make it mathematical: assume the mind is trying to be accurate, but thinking costs a resource: time, memory, computation. Then ask what strategy an agent optimally uses given that budget. Many classic biases begin to look like cheap approximations that trade a little accuracy for a lot of speed.
The cleanest worked example is anchoring. Suppose the mind estimates an uncertain quantity the way a statistician might: by drawing samples from a probability distribution and averaging them. But each sample costs effort, so an optimal agent takes only a few and stops. If it starts sampling from a nearby value, an anchor, and stops before the samples have wandered far, its estimate stays near the anchor. That is under-adjustment: the anchoring bias, derived as the correct behavior for a mind that values its own time. The flaw is what optimal can look like when thinking is not free.
How much weight can this bear? The framework is promising and increasingly productive. It has re-derived anchoring, certain memory effects, and even a rational reinterpretation of why a mind might look like it has two systems. But keep the hype filter on. Because resource-rationality has free parameters, such as the cost, prior, and algorithm, critics warn that it risks becoming unfalsifiable, able to rationalize any behavior after the fact. That is the Day 2 demarcation worry wearing new clothes: a theory that can explain everything explains nothing.
The machines inherited our shortcuts
Here is a twist that would have delighted Simon. When researchers gave large language models classic cognitive-psychology tasks, the models looked unnervingly like us. In a 2023 PNAS paper, Marcel Binz and Eric Schulz gave GPT-3 the Linda problem, and it committed the conjunction fallacy. It also showed anchoring and framing effects. Later work found that stronger models make fewer intuitive slips, but the biases do not vanish. A system trained to predict the next word in a vast ocean of human text apparently soaks up not just our knowledge but some of our reasoning reflexes.
Resist the overclaim. That an LLM reproduces the conjunction fallacy does not establish that it reasons like a human. Whether these systems perform genuine reasoning or sophisticated pattern-completion is exactly the debate reserved for Day 139. For now: the mirroring is a real, replicated, fascinating hint; the interpretation is wide open. And it hands the AI blocks, Days 138-145, a sharp question: when a machine gives a fluent, confident, wrong answer, is that a flaw in the machine, or a mirror held up to us?
Open questions
What’s genuinely unsettled
Fifty years into the study of how minds actually reason, the unresolved list is still long:
- Is there one correct standard of rationality at all? Or is “rational” always relative to a goal and an environment?
- One process or many? Is the mind a single engine doing approximate inference under a budget, or are there genuinely distinct modes? The two-systems picture is wounded, but no agreed replacement has won.
- Are biases to be corrected or respected? If a bias is the optimal output of a bounded mind, debiasing it might make you worse in the real world while better on a logic test.
- Does resource-rationality explain, or merely redescribe? Can it be made falsifiable, or will it always find a cost function that excuses whatever people did?
- And the question the AI block inherits: when a system reproduces our fallacies, has it learned our reasoning, or only our residue?
The day in three sentences
- Big idea
- Heuristics can be useful shortcuts and predictable traps; the yardstick decides which story we tell.
- Best analogy
- Simon’s scissors: mind and environment cut together.
- Live controversy
- Resource-rational analysis tries to reconcile bias research with ecological rationality.
Threads today › computation · evolution · information formats · the unsupported large willpower-fuel model · light emergence.
Tomorrow → Day 12
Networks
Today we studied a single mind and the environment it is tuned to. Tomorrow we zoom out to the connections between minds and things: six degrees of separation, hubs and power laws, how ideas and epidemics spread, and the controversy over whether real-world networks are truly scale-free.
Sources
Sources & further reading
- Tversky, A. & Kahneman, D. (1974). “Judgment under Uncertainty: Heuristics and Biases.” Science 185(4157): 1124-1131.
- Tversky, A. & Kahneman, D. (1983). “Extensional versus Intuitive Reasoning: The Conjunction Fallacy in Probability Judgment.” Psychological Review 90(4): 293-315. doi:10.1037/0033-295X.90.4.293. doi.org/10.1037/0033-295X.90.4.293
- Simon, H. A. (1955). “A Behavioral Model of Rational Choice.” Quarterly Journal of Economics 69(1): 99-118.
- Simon, H. A. (1956). “Rational Choice and the Structure of the Environment.” Psychological Review 63(2): 129-138.
- Simon, H. A. (1990). “Invariants of Human Behavior.” Annual Review of Psychology 41: 1-19.
- Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
- Evans, J. St. B. T. & Stanovich, K. E. (2013). “Dual-Process Theories of Higher Cognition: Advancing the Debate.” Perspectives on Psychological Science 8(3): 223-241.
- Melnikoff, D. E. & Bargh, J. A. (2018). “The Mythical Number Two.” Trends in Cognitive Sciences 22(4): 280-293.
- Hagger, M. S. et al. (2016). “A Multilab Preregistered Replication of the Ego-Depletion Effect.” Perspectives on Psychological Science 11(4): 546-573. doi.org/10.1177/1745691616652873 Dang, J. et al. (2021). “A Multilab Replication of the Ego Depletion Effect.” Social Psychological and Personality Science 12(1): 14-24. doi.org/10.1177/1948550619887702 Vohs, K. D. et al. (2021). “A Multisite Preregistered Paradigmatic Test of the Ego-Depletion Effect.” Psychological Science 32(10): 1566-1581. doi.org/10.1177/0956797621989733
- Gigerenzer, G. & Hoffrage, U. (1995). “How to Improve Bayesian Reasoning Without Instruction: Frequency Formats.” Psychological Review 102(4): 684-704.
- McDowell, M. & Jacobs, P. (2017). “Meta-analysis of the Effect of Natural Frequencies on Bayesian Reasoning.” Psychological Bulletin 143(12): 1273-1312.
- Gigerenzer, G. & Goldstein, D. G. (1996). “Reasoning the Fast and Frugal Way: Models of Bounded Rationality.” Psychological Review 103(4): 650-669.
- Gigerenzer, G. & Brighton, H. (2009). “Homo Heuristicus: Why Biased Minds Make Better Inferences.” Topics in Cognitive Science 1(1): 107-143.
- Lieder, F. & Griffiths, T. L. (2020). “Resource-rational Analysis: Understanding Human Cognition as the Optimal Use of Limited Computational Resources.” Behavioral and Brain Sciences 43: e1. doi:10.1017/S0140525X1900061X. doi.org/10.1017/S0140525X1900061X
- Lieder, F., Griffiths, T. L., Huys, Q. J. M. & Goodman, N. D. (2018). “The Anchoring Bias Reflects Rational Use of Cognitive Resources.” Psychonomic Bulletin & Review 25(1): 322-349.
- Hertwig, R. & Grune-Yanoff, T. (2017). “Nudging and Boosting: Steering or Empowering Good Decisions.” Perspectives on Psychological Science 12(6): 973-986.
- Mertens, S., Herberz, M., Hahnel, U. J. J. & Brosch, T. (2022). “The Effectiveness of Nudging: A Meta-analysis of Choice Architecture Interventions.” PNAS 119(1): e2107346118.
- Maier, M., Bartos, F., Stanley, T. D., Shanks, D. R., Harris, A. J. L. & Wagenmakers, E.-J. (2022). “No Evidence for Nudging After Adjusting for Publication Bias.” PNAS 119(31): e2200300119. With Szaszi et al. (2022), e2200732119. doi.org/10.1073/pnas.2200300119
- Binz, M. & Schulz, E. (2023). “Using Cognitive Psychology to Understand GPT-3.” PNAS 120(6): e2218523120. doi.org/10.1073/pnas.2218523120
Deep dive appendixThe Deep CabinetOptional extension.
The main descent traced one clean corridor: heuristics, the rationality war, Simon’s scissors, and the resource-rational peace treaty. That was the right route for the day. But the cabinet behind it is deeper: choice under risk, overconfidence, confirmation, reflection tests, noise, and the strange cases where less information wins with real money on the table.
What's down here
The main lesson asked whether shortcuts are bugs, adaptive tools, or the shape of a resource-limited mind. This appendix adds the evidence and texture we walked past: prospect theory, the three faces of overconfidence, Wason’s laboratory traps, the bat-and-ball test, dual-process revisions, the bias/noise split, and the profit-and-loss version of “less is more.”
Room Ichoice under risk
Prospect theory: a different value function
The 1974 heuristics paper was about judgment. Kahneman and Tversky’s second revolution was about choice. In 1979 they attacked expected utility theory, the elegant framework that had dominated economics since von Neumann and Morgenstern. Under that framework, people should evaluate final states of wealth and weight outcomes by their actual probabilities. In the lab, they did not.
Prospect theory changed the coordinate system. People evaluate outcomes relative to a reference point, not as final wealth. The value curve is concave for gains, convex for losses, and steeper on the loss side. Small probabilities are often overweighted and moderate-to-high probabilities are often underweighted. This yields the famous fourfold pattern: risk aversion for likely gains, risk seeking for likely losses, risk seeking for tiny-probability gains such as lotteries, and risk aversion for tiny-probability losses such as insurance.
Interactive · prospect theory's shape
The loss side is steeper
Drag the loss-aversion coefficient, lambda. Gains and losses are measured from a reference point; both sides flatten, but the loss side falls harder.
A $50 loss feels about 2.25 times as large as a matching gain.
Figure · Prospect Value
The figure uses the classic illustrative value lambda = 2.25 from cumulative prospect theory.
| Feature | What it says | Why it matters |
|---|---|---|
| Reference dependence | A result is coded as gain or loss relative to a baseline. | The same outcome can feel good or bad depending on what was expected. |
| Diminishing sensitivity | Moving from $0 to $100 matters more than moving from $10,000 to $10,100. | The curve flattens as outcomes move away from the reference point. |
| Loss aversion | Losses are often steeper than gains of equal size. | Giving something up can loom larger than acquiring it. |
The cautious version matters. Ruggeri and colleagues ran a large, coordinated direct replication across nineteen countries and 4,098 participants. They reproduced about 94% of the original prospect-theory choices and found similar patterns across samples. But the slogan “losses hurt twice as much as gains feel good” has been overexported. Gal and Rucker argued that evidence for a universal two-times coefficient is weaker than popular summaries suggest; Mrkva and colleagues replied that loss aversion is real but moderated by attention and context. The right landing is not “loss aversion everywhere.” It is “reference-dependent choice is sturdy, while the exact loss-aversion multiplier is not a law of nature.”
Prospect theory also gave the heuristics-and-biases program a bridge into markets. The endowment effect makes a mug worth more after it becomes my mug. status quo bias makes switching feel like giving something up. The disposition effect makes investors reluctant to realize losses. The value function moved the theory from puzzle cards into institutions.
Room IIconfidence
The three faces of overconfidence
Overconfidence sounds like one thing. Moore and Healy’s 2008 review split it into three, and the split prevents confusion.
| Face | Question it answers | Classic pattern |
|---|---|---|
| Overestimation | How good am I in absolute terms? | I think I will finish the paper by Friday. |
| Overplacement | How good am I compared with others? | I think I am above-average at driving. |
| Overprecision | How narrow is my uncertainty? | I give a tight interval and reality lands outside it. |
These faces do not always point in the same direction. Moore and Healy’s twist is that people often overplace themselves on easy tasks and underplace themselves on hard or rare ones: “better than average” is not a universal direction. That is why one clean definition of “overconfidence” keeps breaking in your hands.
The funniest demonstration is Ola Svenson’s 1981 driving study. About 93% of American drivers rated themselves above the median for driving skill. The impossible arithmetic is the point: people are not merely reporting “I am safe enough.” They are placing themselves above the comparison group.
The most expensive face is the planning fallacy. Buehler, Griffin, and Ross showed that people routinely predict best-case timelines while ignoring their own past delays. The same pattern scales up. The Sydney Opera House was expected to cost about $7 million and open in 1963; it opened in 1973 at more than $100 million. Bent Flyvbjerg’s answer is reference-class forecasting: do not begin with your inside story. Begin with how similar projects actually went.
Why calibration is so hard
Calibration is not confidence in general. It asks whether statements made with 70% confidence come true about 70% of the time, 90% statements about 90% of the time, and so on. That requires a feedback loop with many comparable judgments. Everyday life gives us blurry, delayed, excuse-friendly feedback instead. Without records, confidence can feel like a personality trait rather than a measurable error rate.
Room IIIconfirmation
Confirmation bias and Popper in the lab
Karl Popper told scientists to seek refutations. Peter Wason asked whether ordinary reasoners do that naturally. His 2-4-6 task looked simple. Participants saw the sequence 2, 4, 6 and had to discover the rule by proposing other triples. Most guessed something like “ascending by twos” and tested confirming cases: 8-10-12, 20-22-24. The actual rule was broader: any ascending numbers. A test they expected to fail, such as 1-2-3, would have helped far more, but people tended to hunt for yeses.
The Wason selection task repeated the blow in a different form. Given cards showing letters and numbers, people were asked which cards must be turned over to test a rule like “If a card has a vowel on one side, it has an even number on the other.” The required cards are the vowel, which could confirm or violate the rule, and the odd number, which could expose the violation. Many people choose the vowel and the even number instead. But the plot thickened when Leda Cosmides recast the same logic as a social contract, such as checking whether anyone drinking beer is underage. Performance improved sharply. Humans may be bad at abstract falsification and much better at detecting cheaters in a social rule.
The unsettling part is that intelligence does not rescue you. Keith Stanovich and Richard West distinguished general intelligence from myside bias: smart people can be just as likely to protect their own side, and sometimes better equipped to do it. Kahan and colleagues pushed the point with a numeracy task. The same 2-by-2 table was framed as evidence about a skin cream or gun control. Highly numerate participants did well when the data were neutral; when the data crossed their politics, numeracy sometimes helped them polarize.
That last finding needs a warning label. Later work has questioned how general and replicable motivated numeracy is. Pennycook and Rand argue that much misinformation susceptibility looks less like people reasoning too hard for their tribe and more like people failing to reason enough. The frontier is no longer “are people biased?” It is when the failure mode is motivated reasoning, when it is laziness, and when the format of the problem is doing the damage.
Room IVreflection
The bat and the ball
Shane Frederick’s Cognitive Reflection Test fits on a napkin, but the question cuts deep.
A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?
The intuitive answer is 10 cents. It is also wrong. If the ball cost 10 cents, the bat would cost $1.10, and the total would be $1.20. The ball costs 5 cents, the bat $1.05, total $1.10.
The CRT does not merely test arithmetic. It tests whether a fluent answer can be interrupted. Frederick found that CRT performance predicted patience, risk preferences, and susceptibility to several judgment biases. More recent work connects reflection to misinformation: people who stop to check the first answer are less likely to accept fake headlines. The little bat-and-ball problem is System 2 as an audit button, though we should keep the “System 1/System 2” model loose, as the main lesson warned.
Room Vdual process revised
Dual-process 2.0: intuition is smarter than we thought
The popular story says intuition blurts the wrong answer and deliberation corrects it. Wim De Neys and colleagues complicated that story. In two-response experiments, participants first answer under time pressure, then answer again after deliberation. They also report confidence or show signs of conflict. Many people who give the biased answer still show conflict detection: slower responses, lower confidence, or physiological signs that something is off.
Bago and De Neys went further. In some two-response designs, many correct answers appear already in the first, intuitive response. The implication is not that intuition is always wise. It is that intuition may contain more than stereotype and impulse. There may be “logical intuitions” too: fast sensitivity to conflict, number, and structure.
That keeps the field self-correcting. Dual-process theory is still useful as a teaching map, but the map cannot be “dumb System 1 versus smart System 2.” Sometimes the first response is wrong and the second repairs it. Sometimes the first response already contains the repair. Sometimes a single-process model may explain the data just as well. The neat drama survives only as shorthand.
Room VIerror geometry
Bias versus noise
In 2021, Kahneman returned with Olivier Sibony and Cass Sunstein to argue that organizations obsess over bias while undermeasuring noise. Bias is being wrong in a consistent direction. Noise is scatter: two underwriters, doctors, judges, or managers facing similar cases and giving very different answers.
Their insurance example is deliberately uncomfortable. Executives expected underwriters pricing the same risk to differ by perhaps 10-25%. In one audit, the median difference was about 55%. That is not a philosophical bias about categories. It is costly variability inside a system that believed itself professional.
The arithmetic is simple: MSE = bias^2 + noise^2. To reduce error, you can move the average answer toward the target, or you can reduce scatter. Kahneman, Sibony, and Sunstein divide noise into level noise, pattern noise, and occasion noise: some judges are generally harsher, some are harsher only in particular kinds of cases, and some change depending on mood, sequence, or day.
The practical remedy is decision hygiene: independent judgments before discussion, structured criteria, separating evidence from evaluation, using algorithms where they beat humans, and running noise audits. It is less glamorous than debiasing. It is often easier to measure.
Room VIIless with stakes
Less-is-more, with real stakes
The main lesson used the gaze heuristic and recognition heuristic to show why ignoring information can help. Finance makes the point with money attached.
Harry Markowitz invented modern portfolio theory, yet reportedly split his own retirement savings roughly 50/50 between stocks and bonds because he did not want to regret either extreme. That anecdote is not a proof. But DeMiguel, Garlappi, and Uppal gave the deeper result in 2009: compare optimized portfolio strategies with the naive 1/N rule, which divides wealth equally across available assets. Across many datasets, 1/N was extremely hard to beat out of sample. Some optimized strategies would need around 3,000 months of data before their estimation advantage overcame their estimation error.
The reason is Day 10 in portfolio form. Optimization is a model. Models need estimates. Estimates are noisy. When the environment is noisy and data are limited, a beautiful optimizer can fit ghosts. The crude rule has bias, but it has very little variance. In the bias-variance tradeoff, that can win.
Fast-and-frugal trees show the same lesson in classification. A decision tree that asks a few yes/no questions and exits early can match or beat more complex medical or legal models, especially when the target environment has stable cues and limited data. The win is not magic. It is the scissors again: a simple rule plus an environment where the cue structure makes that rule powerful.
So why do the shortcuts exist at all?
One evolutionary answer is error management theory. If two errors have unequal costs, the best system may be biased. A smoke detector is biased toward false alarms because missing a real fire is worse than waking you unnecessarily. In ancestral environments, mistaking a shadow for a predator and mistaking a predator for a shadow did not have equal payoffs.
That does not mean every bias is adaptive. Gigerenzer calls the opposite mistake the “bias bias”: treating every gap from logic as a defect before asking what task and environment produced it. The three-way version is this: some biases are real errors against good norms. Some are artifacts of bad problem formats. Some are adaptive bets under constraints.
Debiasing sits in that same mixed zone. Morewedge and colleagues found that targeted training games reduced several biases months later. Checklists, reference-class forecasting, natural frequencies, base-rate prompts, and structured decision hygiene can help. But the goal is not to erase intuition. It is to know when to trust the shortcut, when to change the representation, and when to force the slower audit.
- Core addition
- Day 11’s main corridor was judgment under limits. The appendix adds choice under risk, confidence errors, confirmation traps, reflection tests, noise, and less-is-more cases with real stakes.
- Best compression
- Bias is not one thing. It can be a logical violation, an adaptive asymmetry, a format artifact, a noisy estimate, or a cheap approximation under computation costs.
- Practical rule
- Before calling a shortcut irrational, ask three questions: what norm is being used, what environment the mind is in, and what the computation would have cost.
Threads today › computation (limits, approximation, and resource-rational choice) · evolution (asymmetric error costs and adaptive heuristics) · information (problem format, natural frequencies, and noisy judgment systems).
↩ Back to the main path
Day 12 will turn from individual shortcuts to social reasoning: how minds coordinate, disagree, signal, defer, and sometimes trap each other in shared mistakes.
Sources & further reading
- Kahneman, D. & Tversky, A. (1979). “Prospect Theory: An Analysis of Decision under Risk.” Econometrica, 47(2), 263-291.
- Tversky, A. & Kahneman, D. (1992). “Advances in Prospect Theory: Cumulative Representation of Uncertainty.” Journal of Risk and Uncertainty, 5, 297-323.
- Kahneman, D., Knetsch, J. L. & Thaler, R. H. (1990). “Experimental Tests of the Endowment Effect and the Coase Theorem.” Journal of Political Economy, 98(6), 1325-1348.
- Ruggeri, K. et al. (2020). “Replicating patterns of prospect theory for decision under risk.” Nature Human Behaviour, 4, 622-633. doi:10.1038/s41562-020-0886-x.
- Gal, D. & Rucker, D. D. (2018). “The Loss of Loss Aversion: Will It Loom Larger Than Its Gain?” Journal of Consumer Psychology, 28(3), 497-516; with the reply by Mrkva et al. (2020), Journal of Consumer Psychology, 30(3), 407-428.
- Moore, D. A. & Healy, P. J. (2008). “The trouble with overconfidence.” Psychological Review, 115(2), 502-517.
- Svenson, O. (1981). “Are we all less risky and more skillful than our fellow drivers?” Acta Psychologica, 47(2), 143-148.
- Buehler, R., Griffin, D. & Ross, M. (1994). “Exploring the planning fallacy.” Journal of Personality and Social Psychology, 67(3), 366-381.
- Flyvbjerg, B. (2008). “Curbing Optimism Bias and Strategic Misrepresentation in Planning.” European Planning Studies, 16(1), 3-21.
- Wason, P. C. (1960). “On the failure to eliminate hypotheses in a conceptual task.” Quarterly Journal of Experimental Psychology, 12(3), 129-140.
- Wason, P. C. (1968). “Reasoning about a rule.” Quarterly Journal of Experimental Psychology, 20(3), 273-281; Cosmides, L. (1989). “The logic of social exchange.” Cognition, 31(3), 187-276.
- Stanovich, K. E. & West, R. F. (2008). “On the Failure of Cognitive Ability to Predict Myside and One-Sided Thinking Biases.” Thinking & Reasoning, 14(2), 129-167; Stanovich, K. E. (2021). The Bias That Divides Us. MIT Press.
- Kahan, D. M., Peters, E., Dawson, E. C. & Slovic, P. (2017). “Motivated Numeracy and Enlightened Self-Government.” Behavioural Public Policy, 1(1), 54-86. doi:10.1017/bpp.2016.2; replication contested, including Ballarini & Sloman (2017) and Kahan & Peters (2017).
- Pennycook, G. & Rand, D. G. (2019). “Lazy, not biased: Susceptibility to partisan fake news is better explained by lack of reasoning than by motivated reasoning.” Cognition, 188, 39-50.
- Frederick, S. (2005). “Cognitive Reflection and Decision Making.” Journal of Economic Perspectives, 19(4), 25-42.
- Bago, B. & De Neys, W. (2019). “The Smart System 1: Evidence for the Intuitive Nature of Correct Responding on the Bat-and-Ball Problem.” Thinking & Reasoning, 25(3), 257-299; see also De Neys & Glumicic (2008), Cognition, 106, 1248-1299; Raoelison, Thompson & De Neys (2020), Cognition, 204, 104381; De Neys (2021), Perspectives on Psychological Science, 16(6), 1412-1427.
- Kahneman, D., Sibony, O. & Sunstein, C. R. (2021). Noise: A Flaw in Human Judgment. Little, Brown Spark.
- DeMiguel, V., Garlappi, L. & Uppal, R. (2009). “Optimal Versus Naive Diversification: How Inefficient is the 1/N Portfolio Strategy?” Review of Financial Studies, 22(5), 1915-1953. doi:10.1093/rfs/hhm075.
- Haselton, M. G. & Nettle, D. (2006). “The paranoid optimist: An integrative evolutionary model of cognitive biases.” Personality and Social Psychology Review, 10(1), 47-66.
- Gigerenzer, G. (2018). “The Bias Bias in Behavioral Economics.” Review of Behavioral Economics, 5(3-4), 303-336.
- Morewedge, C. K. et al. (2015). “Debiasing Decisions: Improved Decision Making With a Single Training Intervention.” Policy Insights from the Behavioral and Brain Sciences, 2(1), 129-140.
Deep dive appendixThe Bleeding EdgeOptional extension.
Appendix I covered prospect theory, overconfidence, confirmation bias, and other findings that have weathered decades and, mostly, replication. This appendix examines five programs from the 2020s, including preprints and contested interpretations. The task is to separate demonstrated results from plausible extensions and overreach.
Where this sits
The main descent ended on the resource-rational synthesis, and Appendix I filled in the older evidence. This appendix follows both threads to the 2020s frontier: trained networks as cognitive models, finite-sampling theories of bias, resource-rational representations, population-scale prebunking, and the strange attempt to use AI systems as stand-ins for human minds. The recurring threads are computation, information, and the AI arc that later dominates Block IX.
Hype-filter labels
Because so much here is new, the labels do extra work. established means peer-reviewed and independently supported. promising means real but early, narrow, contested, or preprint-shaped. hype means the headline outruns the evidence. Keep Day 2’s warning close: theories that recast every bias as secretly rational are exciting because they explain so much, which is also what can make them hard to falsify.
Centaur, and the black box that predicts you
A long-standing goal in psychology is a model that covers many cognitive tasks rather than one bespoke model per task. In 2025, Marcel Binz, Eric Schulz, and collaborators made a broad attempt. They fine-tuned a large language model on Psych-101, a dataset of trial-by-trial choices from 160 psychology experiments, 60,092 participants, and 10,681,650 decisions, all rewritten into plain English. The result, named Centaur, can be used in experiments described in words: gambles, memory tasks, bandit problems, moral dilemmas.
Centaur predicts held-out participants better than many specialized cognitive models built for each task, and it generalizes across cover stories and task variants. The study also reports that fine-tuning on behavior alone made its representations better at predicting human brain activity, even though it was never trained on brain scans.
The pushback is the Day 10 lesson in a sharper form: prediction is not explanation. Jeffrey Bowers and other critics argue that Centaur can behave in non-human ways under probing. The analogy is simple: an analog clock and a digital clock can show precisely the same time while working by different mechanisms. A model can be a powerful predictor and still fail as an explanation of the mind. A later critical preprint argued that a lightly adapted baseline nearly matched Centaur’s predictive edge, which would make the theoretical lesson less decisive than the headline suggests. The restrained reading is not “an LLM is a mind.” It is “a behavioral black box has become a serious instrument, and whether that instrument can become an explanation is the live question.”
Diagram · the crux of the Centaur debate
Same time, different clockwork
Two mechanisms, one output. If a model reproduces human behavior but runs on machinery unlike a mind, has it explained cognition, or only mimicked it?
Binz and Schulz frame Centaur cautiously: a large black box that predicts behavior well, and whose internals may become informative if opened. The overclaim to resist is that a predictive black box is already a theory of mind.
The sampling hypothesis
The main lesson left a paradox unresolved. Research in perception and motor control often models the brain as a probabilistic engine, yet people perform poorly on some explicit probability questions. How can the same system show both patterns?
Adam Sanborn and Nick Chater proposed a reconciliation: the brain is Bayesian, but it does not calculate probabilities. It samples them. Their slogan is Bayesian brains without probabilities. A sampler with infinite draws obeys probability theory exactly. A bounded mind can afford only a few draws, and small samples are noisy, extreme, and sometimes autocorrelated. Run that same machinery on explicit probability questions and it can produce the conjunction fallacy, base-rate neglect, unpacking effects, and anchoring-like drift.
Jian-Qiao Zhu, Sanborn, and Chater turned the slogan into the Bayesian Sampler and then the Autocorrelated Bayesian Sampler. The attraction is unification: probability judgments, estimates, confidence intervals, choices, confidence ratings, and response times can be fit from one family of mechanisms. If it holds, many biases stop looking like separate glitches and start looking like the shared signature of approximate inference with too few samples.
Interactive · biases from finite samples
The sampling engine
A simulated mind judges two things: how likely A is, and how likely A and B are, by drawing a handful of mental samples for each. Logic demands P(A and B) ≤ P(A). With few noisy samples, the order can flip. Drag the sample count and watch the error shrink.
This shows how noise alone can manufacture incoherence. Real conjunction-fallacy rates in classic cases can approach 85% because representativeness also inflates how likely the conjunction feels; the curve is an illustrative model in the spirit of Zhu, Sanborn, and Chater.
Figure · Sampling engine
The figure marks representative sample counts on the same illustrative curve.
Read the curve as three fixed checkpoints: at two samples, noise alone can make a large minority of simulated minds rank the conjunction higher; by eight samples the error has dropped sharply; by twenty-four it is close to gone in the toy model.
How much weight can the sampling account bear? The models are peer reviewed and broad, but many fits are qualitative, rival models such as Costello and Watts’s PT+N account explain overlapping data, and “this can be modeled as sampling” is not the same as “the brain is sampling.” Large language models also produce incoherent probability judgments, and the Bayesian Sampler framework can model some machine errors. If one mechanism accounts for errors in both brains and transformers, it would suggest a cross-system explanation. It remains a hypothesis.
Resource-rationality goes up a level
The main lesson framed resource-rationality as choosing the right heuristic for a limited mind. Work published in 2022 extended the idea to representation: the mind also selects which features of a problem to keep and which to discard before reasoning begins.
Mark Ho, Tom Griffiths, and colleagues made the case in Nature with a paper that opens on the exact Day 10 image: Borges’s empire-sized 1:1 map, perfectly accurate and perfectly useless. Their claim is that the brain builds the opposite: deliberately lossy cognitive maps, or value-guided construals, that keep only what matters for the goal. Across preregistered maze experiments, people built simplified representations close to what the theory predicted, trading representational complexity against usefulness for planning.
A companion line of work reverse-engineers how people plan and asks whether better strategies can be discovered and taught. That is already feeding into cognitive prostheses: tools and training that nudge people toward strategies a resource-rational planner would use. This is Simon’s scissors turned inward. The environment has structure, but the mind also reshapes the problem to fit its limited blade.
Prebunking at platform scale
Everything so far has been about understanding biased thinking. This edge is about fixing it at scale. The engine is William McGuire’s inoculation theory: just as a weakened virus trains the immune system, a weakened dose of a manipulation technique can build mental resistance. The modern version, prebunking, teaches the trick before the false content arrives.
Roozenbeek, van der Linden, Lewandowsky, and colleagues brought the idea to platform scale with Google’s Jigsaw. In Science Advances (2022), five 90-second videos taught manipulation techniques such as fearmongering, false dichotomies, scapegoating, incoherence, and ad hominem attacks. Lab studies and a YouTube field experiment reaching roughly 5.4 million users found that viewers became better at spotting manipulation techniques afterward, at a reported cost of about five cents per person.
The critique is sharp. Does inoculation improve discernment, or does it merely make people skeptical of everything? In signal-detection terms, does it improve sensitivity, or shift response bias toward “do not believe it”? Modirrousta-Galian and Higham argued that some inoculation-game effects look more like blanket skepticism than sharpened judgment. Add decay without booster doses and smaller effects in the wild, and the warranted verdict is: prebunking works, but what it changes and how long it lasts remain unsettled.
Using the machines to stand in for us
Researchers have proposed using language-model responses as proxies or supplements for human participants because they can be generated quickly and cheaply. Lisa Argyle and colleagues call one version algorithmic fidelity: condition an LLM on demographic backstories and it reproduces some aggregate patterns of real samples. The broader practice is often called silicon sampling.
The critiques rhyme with everything Day 11 taught. Synthetic samples can be too smooth, collapsing the variance and missing the tails of real opinion. They can shift unpredictably across model versions. They can mirror the educated, wealthy, liberal overrepresentation of the training data. A model trained to produce the average next word can flatten humanity into an averaged respondent.
A 2024 preprint by Liu, Geng, Peterson, Sucholutsky, and Griffiths adds a further limitation: when LLMs simulate or predict human choices, they often model people as more rational than observed human behavior, closer to the expected-utility agent Day 11 examined. The proposed instrument therefore carries a rationality bias about the people it is meant to represent.
The frontier scoreboard
Five bets on the 2020s, graded
| Bet | Status | Cautious reading |
|---|---|---|
| Centaur as a unified model of cognition (Nature 2025) | predicts behavior / explanation contested | Psych-101 prediction and neural-alignment claims make it a strong behavioral instrument of uncertain theoretical status. |
| The sampling hypothesis (Psychological Review 2020-2023) | promising unification | Finite Bayesian sampling may connect many biases and LLM probability errors, but rival accounts and falsifiability questions remain. |
| Resource-rational representation (Nature / NHB 2022) | solid results / grand theory open | People simplify representations in goal-sensitive ways; a unified normative theory of cognition remains open. |
| Prebunking at scale (Science Advances 2022) | works / discernment contested | Cheap videos reached 5.4M YouTube users and can improve manipulation-spotting; durability and blanket skepticism remain live. |
| Silicon sampling (Political Analysis 2023-2024) | aggregate mimicry / replacement hype | Useful as a research aid or stress test, but variance loss, demographics, version drift, and rationality bias block replacement claims. |
Open questions
What would have to happen next
- Is behavioral prediction ever explanation? Centaur forces cognitive science to say when a model that predicts us teaches us something about the mind, rather than merely mirroring it.
- Can one mechanism own all the biases? Sampling and resource-rational theories are beautiful because they explain so much. That is also what makes the Day 2 demarcation blade necessary.
- If the mind is optimal all the way down, is bias still the right word? The more constraints matter, the less useful the pure-logic benchmark becomes.
- Can population-scale debiasing produce discernment, not just doubt? A citizenry that disbelieves everything is not clearly better off than one that is occasionally fooled.
- What happens when AI becomes the mirror? As we use AI to model, predict, and stand in for human minds, we may study a flattering machine caricature of humanity instead of humanity itself.
- Big idea
- The 2020s frontier tests trained networks, sampling mechanisms, and resource-rational principles as broad accounts of cognition, alongside applied work on prebunking and synthetic participants.
- Best analogy
- Two clocks, analog and digital, show the same time by different mechanisms. The question is whether a machine that reproduces behavior has explained the mind or only imitated its outputs.
- Live controversy
- Nearly all of it: prediction versus explanation, unification versus unfalsifiability, discernment versus skepticism, and whether LLMs that assume we are more rational than we are can stand in for us at all.
Threads here › computation (bounded inference as both bias engine and possible cure) · information (sampling, representation, and inoculation all decide which information a mind keeps, draws, or resists) · AI (machines as models of, cures for, and stand-ins for the reasoning mind).
↩ Back to the main path
That closes the Day 11 wing: main descent, settled cabinet, and bleeding edge. Day 12 turns from individual shortcuts to networks: connections between minds, contagion, hubs, power laws, and the structure of social reasoning.
Sources & further reading
Post-2020 work cited here. Peer-reviewed unless marked as preprint.
- Binz, M., Akata, E., Bethge, M., Griffiths, T. L. & Schulz, E. et al. (2025). “A Foundation Model to Predict and Capture Human Cognition.” Nature, 644(8078), 1002-1009. doi.org/10.1038/s41586-025-09215-4
- O’Grady, C. (2025). “Researchers Claim Their AI Model Simulates the Human Mind. Others Are Skeptical.” Science news, July 2, 2025. science.org
- Orr, M. et al. (2025). “Not Even Wrong: On the Limits of Prediction as Explanation in Cognitive Science.” arxiv.org/abs/2510.03311
- Sanborn, A. N. & Chater, N. (2016). “Bayesian Brains without Probabilities.” Trends in Cognitive Sciences, 20(12), 883-893.
- Zhu, J.-Q., Sanborn, A. N. & Chater, N. (2020). “The Bayesian Sampler: Generic Bayesian Inference Causes Incoherence in Human Probability Judgments.” Psychological Review, 127(5), 719-748. doi.org/10.1037/rev0000190
- Zhu, J.-Q., Sundh, J., Spicer, J., Chater, N. & Sanborn, A. N. (2023). “The Autocorrelated Bayesian Sampler.” Psychological Review.
- Ho, M. K., Abel, D., Correa, C. G., Littman, M. L., Cohen, J. D. & Griffiths, T. L. (2022). “People Construct Simplified Mental Representations to Plan.” Nature, 606(7912), 129-136. doi.org/10.1038/s41586-022-04743-9
- Callaway, F., van Opheusden, B., Gul, S., Das, P., Krueger, P. M., Griffiths, T. L. & Lieder, F. (2022). “Rational Use of Cognitive Resources in Human Planning.” Nature Human Behaviour, 6(8), 1112-1125. See also Lieder et al. (2019), “Cognitive Prostheses for Goal Achievement,” Nature Human Behaviour, 3, 1096-1106; and Binz et al. (2024), “Meta-Learned Models of Cognition,” Behavioral and Brain Sciences, 47, e147.
- Roozenbeek, J., van der Linden, S., Goldberg, B., Rathje, S. & Lewandowsky, S. (2022). “Psychological Inoculation Improves Resilience against Misinformation on Social Media.” Science Advances, 8(34), eabo6254. doi.org/10.1126/sciadv.abo6254
- Modirrousta-Galian, A. & Higham, P. A. (2023). “Gamified Inoculation Interventions Do Not Improve Discrimination between True and Fake News: Reanalyzing Existing Research with Signal Detection Theory.” Journal of Experimental Psychology: General, 152(9), 2411-2437.
- Argyle, L. P., Busby, E. C., Fulda, N., Gubler, J. R., Rytting, C. & Wingate, D. (2023). “Out of One, Many: Using Language Models to Simulate Human Samples.” Political Analysis, 31(3), 337-351.
- Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B. & Larson, J. M. (2024). “Synthetic Replacements for Human Survey Data? The Perils of Large Language Models.” Political Analysis, 32(4), 401-416. See also Santurkar, S. et al. (2023), “Whose Opinions Do Language Models Reflect?” ICML; and Dillion, D. et al. (2023), “Can AI Language Models Replace Human Participants?” Trends in Cognitive Sciences, 27(7), 597-600.
- Liu, R., Geng, J., Peterson, J. C., Sucholutsky, I. & Griffiths, T. L. (2024). “Large Language Models Assume People Are More Rational Than We Really Are.” arxiv.org/abs/2406.17055
End of Day 011 · 169 descents remain