Dead Wrong: The Method Born from one of Intelligence's Worst Failures — Lesson 3
Hello everyone, I am back with the third lesson in my Indications & Warnings series, where I combine my military intelligence experience with that of the investment world. Today, I’m sharing a lesson about one of the most significant intelligence community failures in modern U.S. history, and I’ll also include insights and takeaways from the perspective of an institutional investment allocator.
After September 11, 2001, the U.S. government swore to never let such a grave attack happen again on the U.S. homeland. The internal assessments that followed found an intelligence community that had failed to share and connect what it knew, and reform began. Thirteen months later, the Community handed policymakers the estimate this lesson is about — and revealed the opposite failure. September 11 was a failure to share information. Iraq would be a failure to question it.
In October 2002, the Intelligence Community published a National Intelligence Estimate titled Iraq's Continuing Programs for Weapons of Mass Destruction. This document, which was classified before being released, carried the Community's strongest language, and that was the high confidence that Baghdad possessed chemical and biological weapons and was reconstituting its nuclear program. This one report (among others) helped the Bush administration decide to invade Iraq, five months later. The Iraq Survey Group then spent over a year searching the country and found no stockpiles and no reconstituted program.
Two post-mortems or internal assessments followed. The first one was the Senate Select Committee on Intelligence, which reported in 2004. The second one was the presidential commission chaired by Judge Laurence Silberman and Senator Chuck Robb that reported in 2005, which concluded that the Community had been "dead wrong in almost all of its pre-war judgments" about Iraqi weapons of mass destruction. It’s crucial to acknowledge that the two investigations had distinct scopes and temperaments, yet they ultimately arrived at the same fundamental conclusion. The issue wasn’t a lack of information; rather, alternative explanations for the information they possessed were never seriously considered. The Community formulated a hypothesis and then sought out confirmation.
Saying that hindsight is 20/20 has become a cliché. It’s easy to claim clarity in retrospect, and many people have built careers on making such observations. The more useful question was asked years before the failure, by a CIA officer named Richards J. Heuer: why do trained, intelligent, motivated analysts fail in exactly this way?
The flaw is in the architecture
Heuer's answer, in Psychology of Intelligence Analysis, shows us exactly what that means. The mind does not naturally generate the full set of explanations for the evidence in front of it and weigh them against one another. It seizes the first hypothesis that fits well enough and then goes looking for support. Confirmation bias is not a character flaw that better analysts learn to avoid; it is the default setting of human cognition, and it runs underneath conscious reasoning where willpower cannot reach it. That is why "be more objective" has never fixed anything. If the flaw is in the architecture, the fix has to be in the procedure. (There is an entire future lesson hiding in that sentence. For now, take it as given.)
Heuer then came up with a procedure, and it is known as the “Analysis of Competing Hypotheses (ACH).”
The method
If you were to look up the “official” ACH, it runs in eight steps; however, in practice it compresses to a working sequence. First, write down the full set of hypotheses that could explain what you are seeing; this includes the uncomfortable ones, and the ones nobody in the room favors (you must do this, so you can see every angle). Second, take an inventory of the evidence and arguments, and note what is absent that each hypothesis says should be present. Third, build a matrix: hypotheses across the top, evidence down the side, and a judgment in each cell about whether that item is consistent, inconsistent, or irrelevant to that hypothesis. The discipline lives in the direction of travel. You work across the rows, asking how each piece of evidence bears on every hypothesis, not down your favorite column, collecting support.
Fourth, this is what separates ACH from a brainstorm: you attack by refutation. You proceed by trying to disprove hypotheses, not by trying to prove the one you like. The hypothesis that survives is the one with the least evidence against it, not the one with the most evidence for it, because confirming evidence is cheap and disconfirming evidence is rare. Fifth, check how heavily your surviving conclusion leans on one or two linchpin items, and say so out loud. Finally, report the whole ranked field rather than just the winner, and set indicators that would tell you the ranking is changing.
Here is an example of an ACH Matrix
Diagnosticity, the crown jewel
Buried in that third step is the concept I would keep if I could keep only one thing from the entire method: diagnosticity. Evidence here carries the analytic weight only to the degree that it discriminates between hypotheses. Heuer's illustration is a fever. A 104-degree temperature tells the doctor a great deal about whether you are sick and almost nothing about which illness you have, because it is consistent with dozens of them. The reading is accurate, dramatic, and nearly useless for the decision at hand.
Now let me tell you about the aluminum tubes. In 2001, Iraq was caught procuring high-strength aluminum tubes, and the assessment that reached the NIE read them as rotor casings for uranium-enrichment centrifuges — evidence of a reconstituted nuclear program. If you were to read down that column, the tubes fit. Now read across the row, and the picture begins to change: the same tubes matched, almost exactly, the dimensions of a conventional rocket that Iraq had been producing for years — which is why the Department of Energy's centrifuge experts judged them more likely to be intended for rockets. Evidence consistent with two hypotheses cannot, on its own, elevate either. However, as we now know, the estimate counted it as confirmation anyway.
The mobile-biological-labs reporting failed the same way, differently. A single defector — cryptonym Curveball, handled by German intelligence, never directly interviewed by the analysts relying on him — supplied vivid, detailed descriptions that anchored the biological weapons judgment. The vividness did the work that corroboration should have done. A report can be specific, dramatic, and emotionally compelling while carrying no diagnostic weight at all. If it fits every column, it actually moves nothing.
What the method actually produces
If you were to run this honestly, ACH does not output certainty. It outputs a ranked set of hypotheses; an explicit statement of what evidence would disconfirm the current leader; and collection requirements — the missing evidence that would actually discriminate between the survivors. Those requirements are the real product, as they tell you what to go find next. You already do this instinctively. If you are stuck on a jigsaw puzzle, your brain drafts the collection requirement on its own — straight edge, a bit of red sky — and it sends you hunting. The gap then specifies the search. The matrix does the same job deliberately: its empty rows are the gaps.
I share all of this with you all so that you realize that it took the worst analytic failure in a generation to make this doctrine. The reform legislation of 2004 created the Director of National Intelligence, and the analytic standards that followed — Intelligence Community Directive 203 —made analysis of alternatives a written requirement rather than a personal virtue. Structured techniques are now taught, graded, and inspected. Institutions rarely adopt discipline without catastrophe. Hold that thought, and let’s apply it to the investment world.
The allocator's version
Now, my brothers and sisters in finance, here is the translation, and I promise, I will keep it short, because if you have read this far, you have probably already run it in your head and figured out where I am going with this.
Think of an investment manager's track record as a finished intelligence report. It is the polished product of a process you did not observe, and your job is to explain it. This is where it gets fun and interesting; there are at least five competing hypotheses on the table every single time: repeatable skill; exposure to a factor or regime that happened to be favorable; leverage; vintage or sector luck; and one outsized position doing the work of the whole portfolio (there are definitely more; this is just to name a few).
Take the standard due-diligence file and run it across that matrix. Pedigree from a brand-name firm. Articulate quarterly letters. Growing AUM. References the manager selected for you. The track record itself. Every one of those items is consistent with all five hypotheses. A lucky manager and a skilled one hold the same credentials, write with the same polish, and hand you the same reference list. The file is impressive and almost entirely non-diagnostic. As I discussed earlier, this is the fever chart.
What discriminates? Return decomposition against the relevant factors. Attribution with the best position removed. Behavior in the regime the strategy is supposed to survive, not the one it was born in. References you sourced off list. The economics that keep the second generation of the team in their seats. And when you find that your matrix has empty rows — evidence that would discriminate, but that you do not hold — you have not failed. You have produced exactly what the method produces at Langley: collection requirements. The missing discriminating evidence is the diligence agenda. It also flips the meeting on its head: you stop grading the pitch and start hunting for what would falsify it.
Here is an example of an ACH matrix for those in the Allocator seat
The close
The Intelligence Community bought this discipline at the price of its worst public failure and two commissions' worth of hindsight. You can have it for the price of a matrix. At the end of the day, the question in the data room is never whether the manager looks good — everyone in the data room looks good. The question is what evidence would separate a skilled manager from a lucky one, and whether you actually hold any of it.
In the Intelligence schoolhouse, they taught us to try to prove ourselves wrong before events did it for us. Heuer's way costs a spreadsheet. The other way costs a commission report.