Special Report No. 14
In March 1952 the Air Force decided to evaluate every UFO report it had. Three years later it received the answer: a statistical study of about 4,000 reports, reduced to punched cards, sorted by four analysts and charted as 3,201 sightings. It is still the largest thing of its kind, it has been public since October 1955, and almost everyone who cites it cites a single percentage. That percentage is real. It is also one of ten, because the report counted its file three different ways and there are four defensible things to do with the cases nobody could evaluate at all. Underneath the percentage is something stranger. The analysts noticed that knowns and unknowns seemed to trend alike and did the honest thing: they tested it. Five characteristics of six came back saying the two groups were not one population, at better than one per cent. The report called that inconclusive, went looking for a model of a flying saucer in its twelve best-described cases, found that no two of them agreed, and concluded that the probability any unknown was a saucer is extremely small. Then the Air Force released it with a statement that the question was settled, and a number far below anything in the file entered circulation. Three documents, one year, one cover. This bench prints all three and closes none of them.
Open the interactive ▸ What you're looking at
The Sort view is 3,201 cases as one bar in six bins, with the denominator drawn as a bracket underneath it. Change the counting basis (all sightings, unit sightings, object sightings) and the rule (as published, insufficient information out of the denominator, insufficient information into the numerator, or the report's own re-evaluation applied) and the headline moves in front of you. A ruler at the foot carries every answer at once, with the famous 21.5% and the 1955 public statement marked in different colours because they are different kinds of number.
The Chi Square view is the report's Tables II through XIII as a bench you can run. Knowns rise in amber, unknowns fall in cyan, the dashed lines are the gap and the strip along the bottom shows what each bucket costs in χ². Four switches: which characteristic, whether the astronomical cases are deleted from the knowns as the report deleted them, whether the statistic is the report's or the textbook contingency form, and whether the 'not stated' bucket is in the test at all. The gauge prints the exact p value beside the two critical lines 1955 had.
The Quality view draws Figure 8 as a mosaic: column width is group size, column height is the evaluation split, and the white line is the unknown share. It shows the strongest thing in the whole report, that excellent sightings were twice as likely to defeat identification as poor ones, and immediately shows the part that gets left out, that doubtful came in below poor. The grey dashed line is why. A side panel runs the 4 × 2 contingency test the report never applied to its own reliability column.
The File view runs 1947 to 1956 in twelve cards: Arnold, the March 1952 decision, the flap that tripled the pile, the punched-card machine, the sort, the mirror-graph hypothesis, the test, the hinge, the re-evaluation panel, the failed model, the conclusions, and the October 1955 release with the reply it produced. Two verdicts hang below at equal size, and both of them are the report's.
Why it's here
The station already has one statistics bench. INST-37, Project Stargate, asked what a seven-point margin is worth where chance earns exactly 25. This one asks an uglier and far more common question: what is a percentage worth when nobody agrees on the denominator. Special Report No. 14 is the most-cited government document in the UAP literature, and it is nearly always cited the same way: "the Air Force's own study found 21.5% unexplained." That sentence is true. It is also one of ten answers the same file gives under a handful of equally defensible rules, and those ten answers are spread across nearly fourteen percentage points.
This bench is not here to debunk it. What makes Special Report No. 14 worth an instrument is the tension inside it: it ran an honest test, got a clear result, and then wrote a conclusion pointing the other way. The analysts noticed that knowns and unknowns looked alike, and instead of assuming it they tested it. Five characteristics of six rejected "these are the same thing" at better than 1%. The report then called that result inconclusive, turned to building a model of a "flying saucer" from its twelve best cases, failed, and concluded from that failure that the probability any unknown was a saucer is extremely small. Neither step is absurd. But they are not the same argument, and for seventy years each camp has quoted one half. This bench prints both halves and closes neither.
How it works
Every count on this instrument is printed in the 1955 report. Everything else is arithmetic on those counts, and a script that ships with the site re-derives the report's own published chi squares from them before the build is allowed to pass.
χ²_report = Σ (K·r − n)² / (K·r), r = ΣN / ΣK · χ²_contingency = Σ (O − E)² / E · p = Q(df/2, χ²/2)
The denominator machine is division, and that is the point. Three counting bases and four rules make twelve fractions from one file. The report itself publishes all three bases side by side in Figure 2, so nothing here is smuggled in: the study was explicit that a case, a party of observers and an object are three different units, and it left the reader to notice that the answer depends on which you pick.
The chi square is the report's own, reproduced exactly. Scale the knowns by the ratio of the totals so the two distributions have the same size, take the difference bucket by bucket, square it, divide by the adjusted knowns, and sum. Feed the printed counts back in and the printed totals come out, within the rounding the analysts did by hand. That reproduction is the licence for everything else the bench does with those tables.
The p value is the one thing 1955 genuinely could not have. The report looked up two critical values in a textbook and reported which side of them it landed. This bench evaluates the tail directly, which converts 'better than 1%' into about one in thirty million for the duration table, and leaves brightness sitting at p = 0.33. Both critical lines stay on the gauge, because reading a 1955 document with 2026 arithmetic is only honest if you can see the difference.
The 'not stated' switch prices a problem the report skipped. Every characteristic has a bucket for cases where nobody recorded that attribute, and in brightness it is 66% of the file. Take it out of the speed table and χ² falls from 38.3 to 23.4: still significant, but a third of the signal was living in the absence of data. The report never ran that check, and it is one line of arithmetic.
And the reliability test is the one the report never ran at all. Figure 8 is a contingency table and was never treated as one. Run it: χ² = 63.0 on three degrees of freedom, p about 10⁻¹³, an association several orders of magnitude stronger than any of the six characteristics the report did test. That result does not identify anything. It does establish that the sorting was not independent of how good the evidence was.
What this bench will not do is convert a difference into an identification. Five tables saying two distributions differ is a statement about shape, not about origin, and the report says so in as many words before drawing its conclusion from somewhere else entirely. The instrument's discipline is the station's usual one: print the measurement, print the conclusion, mark which is which, and hand the reader the dial that changes the answer.
The dials that decide what happens
Two of them run the arithmetic of the headline, four run the test, and one runs the table the report forgot. None of them can close the case, and that is the finding.
- The counting basis. All sightings (3,201), unit sightings (2,554), object sightings (2,199). The same file, counted by report, by observing party and by described object. The report publishes all three and the literature quotes them interchangeably, which is where most of the confusion about this document begins.
- The denominator rule. Four ways to handle the cases filed as insufficient information: leave them in, take them out of the bottom, move them into the top, or apply the report's own second look at 57 of its unknowns. Every one has been used in print by someone. On the object basis alone the spread is thirteen and a half points.
- The characteristic. Colour, number, shape, duration, speed, brightness. Six tables, six shapes of disagreement, and six quite different stories about what the unknowns did that the knowns did not. Duration is the strongest; brightness is the null result the report used to soften the other five.
- The knowns, as tested or with astronomical deleted. The report's own revision, and the reasoning behind it is worth the click: every large excess on the knowns' side traced back to identifiable astronomical objects, and on the colour table 98 of the 130 green knowns are astronomical, which the report attributes to the green fireballs of the south-west.
- The statistic. The report's formula treats the knowns as a fixed reference distribution; the textbook contingency chi square treats them as a second sample. Switching between them moves the answers very little, which is the fair verdict on 1955's arithmetic: rough, and not wrong.
- The 'not stated' switch and the quality measure. One asks how much of each result is carried by missing data. The other asks whether the famous quality claim survives removing the insufficient-information cases from every group. Both survive partly, and the parts that do not are on screen.
The claims, as they stand
Six claims about this document, with who made each and where it lands against the printed tables. Two are measured. Four are contested, mostly because they are half-quotations of things the report really does say.
| The study found 21.5% of its cases unexplained proposed by Battelle Memorial Institute, 1955 | MEASURED | Measured, and correct on one of three bases. 689 of 3,201 all sightings is 21.5%. On object sightings the unknowns are 434 = 19.7%, and the number 474 = 21.5% in that same figure belongs to Aircraft. Both percentages are printed in Figure 2. Almost every citation of this report picks one and names neither. |
| The better the report, the more likely it was unexplained proposed by the standard reading of Figure 8 | CONTESTED | Half measured, half misquoted. Excellent reports ran 33.3% unknown against 16.6% for poor ones: that is real, striking and correctly cited. But doubtful reports came in at 13.0%, below poor, so the monotone staircase is not what the figure shows. The mechanism is visible on the instrument: insufficient information climbs from 4.2% to 21.4% as quality falls, draining cases out of every evaluated category. |
| The chi square proved the unknowns were something else proposed by the strong version, in the UFO literature | CONTESTED | Overstated, and the overstatement matters. Five characteristics of six do reject the one-population hypothesis at better than 1%, which is a real result on the report's own data. What it licenses is that the two piles have different shapes. It does not license any statement about what the unknowns were, and the report's own discussion of why each table came out that way is worth reading before anyone builds on it. |
| The chi square results were inconclusive proposed by Special Report No. 14, p. 69 | CONTESTED | The report's own verdict on its own test, and the hinge the whole document turns on. Its reasoning is that unknowns could still be knowns in different proportions, which is fair as far as it goes. It then abandons the statistical approach entirely rather than pursuing it, on the grounds that the inaccuracies would give a distorted and meaningless result. |
| The Air Force misrepresented the report in public proposed by Leon Davidson, 1956 | CONTESTED | Contested, and this bench takes no side on intent. What is checkable: the October 1955 release presented the study as settling the question, and the unknown figure that reached the public was far below every defensible reading of the file, because it described a later case load. Davidson reprinted the whole report at his own expense to make that argument, which is why a private citizen is the reason many people have read it at all. |
| Table X is arithmetically wrong, and it changes a verdict proposed by computed on this bench | MEASURED | Measured, and checkable in ten seconds. Table X's OTHER row prints 1.76; its own two numbers, adjusted knowns 51 and unknowns 54, give (51−54)²/51 = 0.18. Corrected, the table totals 9.88 instead of 11.46, which puts it below the 5% critical value printed on the same page. Twelve tables of hand arithmetic, one slip, and it lands on the one table where it flips the answer. |
Try this
- Start on the sort, all sightings, as published. There is the 21.5%. Now click the dashed marker labelled with it and read what else in Figure 2 carries that same percentage.
- Switch the basis to object sightings. The headline drops to 19.7% and Aircraft, at 474, is now the 21.5% bin. Nothing was recomputed. The unit of counting changed.
- Now change the rule to drop insufficient information. 22.2%. Then to fold it into the numerator: 30.7%. Neither is cheating. Both have been printed by serious people, and the report never says which it prefers.
- Apply the re-review. 17.1%, using the report's own panel outcome. Then notice that this number appears nowhere in the report, because Figure 3 was never redrawn after the panel reported.
- Click the 1952 block in the year strip. 1,501 of 2,199 object sightings. Whatever this study measures, it mostly measures one summer.
- Go to the quality table. Excellent 33.3%, poor 16.6%. This is the strongest single fact in the document and it is quoted correctly. Now follow the white line all the way and find the doubtful group.
- Turn on the insufficient-information line and watch it climb the other way. 4.2% in the excellent group, 21.4% in the poor one. Then switch the measure to 'share of evaluated' and see how much of the anomaly survives.
- Read the side panel. χ² = 63 on the reliability table, an association the report never tested, far stronger than any it did.
- Go to the chi square, duration, as tested. 49.5 against a 1% critical value of 18.5, and a computed p near one in thirty million. This is the report rejecting its own most plausible hypothesis.
- Switch to speed and turn on 'drop not stated'. χ² falls from 38.3 to 23.4. A third of that result lived in the cases where nobody wrote a speed down.
- Switch to shape, knowns revised, and read the gauge. The report printed 11.46 and its own row gives 9.88. One misplaced decimal point in twelve tables, on the one table where the correction crosses the 5% line.
- End on the file and read both verdicts. If you leave certain the report was a cover-up, you have overshot the evidence. If you leave certain it settled the question, you have overshot it in the other direction, and its own Table V is the reason.
Accuracy
The honest line between what the report printed, what is computed here from what it printed, and what is a reading:
| Feature | Tier | What that means |
|---|---|---|
| Figures 2, 3 and 4, complete | T1 Measured | The sort on all three of the report's counting bases: 3,201 all sightings, 2,554 unit sightings, 2,199 object sightings, each split into the same six evaluation bins. Every column sums exactly to its own printed total, which is the first thing the verification script checks. The per-year distribution is Figure 4's, including the 1,501 object sightings from 1952 alone. |
| Figure 8, the reliability table | T1 Measured | All twenty-four cells: four sighting-reliability groups (excellent 213, good 757, doubtful 794, poor 435) by six evaluations. The six cells of each group sum to its total and the four unknown cells sum to the 434 the chi square tables use. Nothing is rounded, collapsed or dropped, including the doubtful group that breaks the trend everyone quotes. |
| All twelve chi square tables | T1 Measured | Tables II-VII and VIII-XIII: six characteristics, the knowns at 1,765 and the revised knowns at 1,286, against the 434 unknowns, with the report's printed chi square, degrees of freedom and both critical values. Every table reconciles in three columns, and the difference between the two knowns columns sums to exactly the 479 astronomical object sightings the report says it removed. |
| The re-evaluation ledger | T1 Measured | Table I and the panel that followed it: 186 unknowns in the daylight sun-elevation groups, 164 folders re-opened, 18 possible aircraft, 20 possible balloons, 19 other possible knowns, 100 still unknown, 7 good unknowns, plus 5 more from the night scan for a total of 12. The 186 figure is not asserted here: it falls out of Table I when you add the right five columns. |
| The exact p values | T2 Modelled | The report could only look up 5% and 1% in a printed table. This bench evaluates the chi square tail with the regularised upper incomplete gamma function, so the duration result reads as roughly one in thirty million rather than "better than 1%". The two critical lines are still drawn, because they are what 1955 actually had in front of it. |
| The contingency form of the test | T2 Modelled | The report's statistic scales the knowns to the unknowns' total and sums (K−n)²/K, which treats the knowns as a fixed reference distribution rather than a second sample. The textbook comparison is a 2 × m contingency chi square with expectations from the margins. Both are on a switch. They barely disagree, which is worth knowing before anyone calls the 1955 arithmetic unsound. |
| The test of Figure 8 | T2 Modelled | A 4 × 2 contingency chi square on report reliability against unknown-or-not: 63.0 on three degrees of freedom, p about 10⁻¹³. It is by a wide margin the strongest association in the document, and the report never ran it. Nothing sinister follows; it is simply the one that would have been hardest to explain away. |
| The ten-rule spread and the 17.1% | T2 Modelled | Three counting bases against four denominator rules, all arithmetic on the published counts, running from 17.1% to 30.8%. Ten of the twelve combinations exist, because the re-evaluation rule reaches only the object basis. The 17.1% is what the report's own re-evaluation implies for the object basis once its 57 reclassified cases are applied. The report published that panel's outcome in prose and left Figure 3 as it was. |
| The "under 3%" marker | T3 Reading | Drawn on the same ruler for contrast and tagged separately, because it is a different kind of number: it belongs to the material released with the report in October 1955 and describes the Air Force's then-current case handling, not this 3,201-case file. The distance between it and every tick on the ruler is why a 1956 pamphlet existed. |
| What the unknowns were | T3 Reading | Open, and the instrument declines to close it. A measured difference between two distributions is not an identification of either. The report understood this, wrote that misidentification in different proportions was the more probable case, and produced the conclusion it is remembered for. Both are on the file view, at equal size. |
In one line: every count on this bench is printed in the 1955 report and carried complete, with Figures 2, 3, 4 and 8, Table I and all twelve chi square tables reconciling in every column; the exact p values, the contingency form of the test, the chi square on Figure 8, the ten-rule spread and the 17.1% are MODELLED arithmetic on those counts, checked by scripts/sr14-tune.mjs before the build passes; the 'under 3%' marker is REPORTED and belongs to the October 1955 release rather than to this file; and what the unknowns actually were is a reading the bench declines to make, hanging the report's conclusion opposite the report's own arithmetic at exactly equal size. Pick a denominator, run the test, and find the sentence the other one contradicts.
Sources
- Primary document: Battelle Memorial Institute, "Analysis of Reports of Unidentified Aerial Objects", Project Blue Book Special Report No. 14, Study No. 102-EL-55/2-79, Project No. 10073, Air Technical Intelligence Center, Wright-Patterson AFB, 5 May 1955. Every count on this instrument is read off a printed figure or table of that document: Figures 1-4 (pp. 17-20), Figure 8 (p. 24), Table I (p. 60), Tables II-VII (pp. 62-67), Tables VIII-XIII (pp. 70-75), the re-evaluation and the "flying saucer" model attempt (pp. 76-93), and the Conclusions (p. 94).
- The report's definition of its terms, Introduction pp. 1-2: J. Allen Hynek's definition of a "flying saucer" as "any aerial phenomenon or sighting that remains unexplained to the viewer at least long enough for him to write a report about it", which the report quotes and then replaces with its own narrower one, "a novel, airborne phenomenon, a manifestation that is not a part of or readily explainable by the fund of scientific knowledge known to be possessed by the Free World".
- The report on the quality of its own data, pp. 3-4: "any conclusions contained in this report are based NOT on facts, but on what many observers thought and estimated the true facts to be"; the note that a few reports originated from people confined to mental institutions; and the Summary's statement that the subjectivity of the data "presented a major limitation to the drawing of significant conclusions, but did not invalidate the application of scientific methods of study".
- The chi square method as the report states it, p. 61: "a statistical test of the likelihood that two distributions come from the same population", with the four-step procedure, the adjustment of knowns to the unknowns' total, and the printed 5% and 1% critical values. The discussion of what drove each result, pp. 68-69, including the finding that 98 of the 130 green knowns were astronomical and attributable to the green fireballs of the south-western United States.
- The re-evaluation and the model attempt, pp. 76-93: 164 folders of 186, the five outcomes, the twelve good unknowns sorted into propeller, aircraft, cigar and disc shapes, the four conditions a model would have had to satisfy, and the conclusion that "out of about 4,000 people who said they saw a 'flying saucer', sufficiently detailed descriptions were given in only 12 cases. Having culled the cream of the crop, it is still impossible to develop a picture of what a 'flying saucer' is."
- The Conclusions, p. 94, quoted on the instrument's file view: that it can never be absolutely proven that "flying saucers" do not exist; that the failures to identify are attributed to reported manoeuvres and to the unavailability of supplemental data such as aircraft flight plans and balloon-launching records; that "there was a complete lack of any valid evidence consisting of physical matter"; and the two probability statements, "extremely small" and "highly improbable".
- The document as released: the U.S. Air Force made the study public in October 1955. The "under 3%" figure carried on the instrument's ruler belongs to that release and to the Air Force's then-current investigation programme, not to the 1947-1952 file the study analyses. This site's own Release 03 briefing on the document covers the released version and its framing.
- Leon Davidson, "Flying Saucers: An Analysis of the Air Force Project Blue Book Special Report No. 14", first published 1956 and reprinted in later editions, which reproduced the report at the author's own expense in order to argue that the public account and the document disagreed. Cited here as the origin of the dispute, not as an endorsement of its conclusions.
- Scanned copies of the full report, including all 240 charts, tables and figures and the 255-page Appendix A of frequency tabulations, are on the Internet Archive and in the Black Vault document archive. Every number used here can be checked against those scans, and scripts/sr14-tune.mjs re-derives the report's own published chi squares from the counts it prints.
Pick a denominator. Run the test the report ran. Then read the conclusion it wrote instead.
Open the interactiveCompiled August 2026