# Do Beale ciphers 1 and 3 behave like book ciphers? Statistics against the solved cipher 2

Statistics line of inquiry. Data: `data/primary/cipher{1,2,3}.txt` (Wikisource transcription of the Library of Congress copy of the 1885 pamphlet; cipher 2 as printed has 762 numbers once the spurious leading "2" of the Wikisource text is dropped, see `results/transcription_check.json`), `data/primary/doi_pamphlet_as_printed.txt` (the Declaration as printed and numbered in the pamphlet), the printed decipherment of cipher 2 and the three "Beale" letters. Code: `code/stats_*.py` (Python 3.11, numpy/scipy/matplotlib; `code/stats_run_all.sh` reruns everything, 60-90 min, the scan-forward grid being the slow part; every Monte-Carlo test uses a fixed seed, given in each script). Machine-readable results: `results/stats_*.json`; the tables below are generated by `code/stats_summary_tables.py` into `results/stats_summary_tables.md`. Figures: `results/stats_hist_values.svg`, `stats_last_digits.svg`, `stats_sequence_plots.svg`, `stats_growth_curves.svg`, `stats_growth_curves_calibrated.svg`, `stats_acf.svg`.

## Conclusion in plain language

Three hypotheses were tested for each unsolved paper: (i) a genuine book cipher keyed to the Declaration of Independence (DOI) as cipher 2 is; (ii) a genuine book cipher keyed to some other text; (iii) numbers written down to look like a cipher.

**Cipher 1 (520 numbers, "locality of the vault").**
- (i) is excluded on these data. Ten of its numbers exceed the 1322 words of the DOI (eighteen exceed 1005, the largest number used in cipher 2). Decoded with the DOI key its letter distribution is exactly what the DOI's initial letters give for numbers unrelated to the key (chi-square against that null 28.7 on 21 df, Monte-Carlo p = 0.19; its distance from English, KL = 0.253, is no smaller than the null's 0.275 +- 0.03, p = 0.27), whereas cipher 2 decoded the same way is unmistakably English (KL = 0.039, p < 0.0003). At the same time Gillogly's alphabetic run is real and sharper than he reported: with the pamphlet's own numbering (the numbering that decodes cipher 2) numbers 188-204 of cipher 1 decode to `abcdefghiijklmmno`, a 17-letter non-decreasing run followed by `hpp`; no shuffle of the same letters in 5000 produced a run of 17 (p = 0.0002; expected longest run 6.3). Nothing like it exists in cipher 2 or 3. So at least those twenty numbers were taken from the numbered DOI in alphabetical order, not by enciphering a text.
- (ii) cannot be excluded by the reuse structure. Cipher 1 re-uses numbers far less than cipher 2 (298 distinct values in 520 numbers, 59 % of them used once; cipher 2: 181 in 762, 24 % once). The encoder habits fitted to cipher 2 (re-use probability 0.78 overall, rising to 0.94 once a letter has 13 homophones in play) would give 164 +- 6 distinct numbers for a 520-letter English text with the DOI and 169 +- 7 with a 2906-word key; a memoryless homophone chooser would give 377 +- 8 and 450 +- 7. Cipher 1's 298 lies between, and a homophonic encoder whose re-use probability is lowered to 0.41-0.45 reproduces its whole frequency profile at once (share of values used once, count of the commonest value, share of the ten commonest, Zipf slope: all between the 8th and 90th percentile of 1000 simulations). Its serial structure (lag-1 rank autocorrelation 0.195, 36 rising runs of length >= 4 against 17 expected) is of the same kind and size as the genuine cipher 2 shows (0.206 and 39 against 24), which no memoryless model reproduces but which a light scan-forward habit does - the encoder taking the next word in the key with the right initial one time in five to ten (section 5.4). Its last digits are only mildly non-uniform once re-use is allowed for (p = 0.035; distinct values p = 0.03; even values 56 % of distinct values, p = 0.03; repeated final digits under-represented, p = 0.03): suggestive of a hand choosing digits, but each of these is a single 5 %-level result among a dozen tests.
- Verdict: cipher 1 is not a DOI cipher; its number statistics are compatible with a homophonic book cipher made with a lighter re-use habit than cipher 2's, and equally with numbers generated to imitate one; the one decisive feature is the alphabetic run, which shows that its author drew at least part of it from the numbered DOI without enciphering anything. The evidence for (iii) rests on that run, not on the digit or reuse statistics.

**Cipher 3 (618 numbers, "names and residences").**
- (i): the DOI decode is indistinguishable from the unrelated-key null (chi-square 22.9, p = 0.70; KL to English 0.246 against 0.273 +- 0.03 under the null, p = 0.20). No letter is in real excess (the flagged `q` is a single hit on word 684, "quartering", against 0.1 expected).
- (ii)/(iii): the frequency profile is again reproducible by a homophonic encoder with re-use probability 0.59 (all profile statistics between the 4th and 87th percentile), but the sequence itself is unlike anything a letter-by-letter encoder produces: lag-1 autocorrelation 0.62 (rank 0.49), lag-2 0.45, only 332 runs up and down where shuffles give 410 +- 11 (z = -7.6), 17 strictly rising runs of six or more numbers against 0.7 +- 0.8 expected, the second half of the cipher averaging 81 higher than the first (p < 0.0001), the largest decile of numbers clustered in position (KS p < 0.001). Cipher 2, whose encoder demonstrably had a "local" habit (51 adjacent pairs within 5 of each other against 33 expected once letter identity is held fixed), shows none of this beyond the mild lag-1 effect. Rising runs and drift are what a hand writing "cipher-like" numbers produces; they are not produced by looking up letters, unless the encoder scanned his key forward from the last word used. That habit was tested (section 5.4): the genuine cipher 2 shows it in a light form (one step in ten), which also covers cipher 1; but no strength of it fits cipher 3. The strength its autocorrelation needs (six steps in ten) produces twice the rising runs it has, and the strength that fits its runs (three in ten) gives half its autocorrelation; cipher 3's correlation persists over many lags, as slow waves of level, not the fast-decaying saw-tooth that scanning makes.
- Capacity: the list the pamphlet says paper 3 contains (thirty men, each with at least one relative and a residence) needs more letters than 618 under any realistic assumption. With real 1810-1850 Virginia name lengths (first name 6.1, surname 6.4 letters on average) and the 78 Virginia county/city names of 1822 (8.5 letters), the bare minimum - each man's first name and surname, one relative's first name, the county name and nothing else - averages 792 letters (5th percentile 760; probability of fitting in 618 numbers: 0 in 20 000 draws). Only lists reduced to surnames and county names (631 +- 17) or initials plus surnames (640 +- 18) come within reach, and even those fit only 12-24 % of the time. Cipher 2 wrote everything out ("in the county of bedford", "nov", "dec", "st" being the only abbreviations).
- Verdict: cipher 3's number sequence has strong internal structure that a book cipher of a written list would not have, and the paper cannot hold the list it is said to hold. On these tests (iii) is the best-supported reading for cipher 3. The residual possibility is an encoder habit of a kind not modelled here - for instance working through a long key a page at a time, which would also make the numbers wander in level; that would be a book cipher of a different sort from cipher 2, and the capacity result stands against it independently.

**Cipher 2 as reference.** Everything above is calibrated on cipher 2, which behaves as a genuine but sloppy homophonic book cipher: 181 distinct numbers for 762 letters, strong preference for early words of the key (a new number is chosen with weight exp(-n/119); 53 % of numbers are <= 100 although only 8 % of the key's words are), a marked preference for numbers ending in 0 or 5 (26 % of distinct values, p = 0.03; the pamphlet's key is numbered every ten words), heavy re-use, and a serial habit (lag-1 rank autocorrelation 0.21, 51 close pairs) that survives permuting numbers within letter classes (p = 0.0002) and is therefore the encoder's, not the language's. Any argument that cipher 1 or 3 is "non-random, therefore real" must reckon with the fact that the genuine cipher is non-random in the same directions, only less so than cipher 3.

## What is new

Confirmations of published work: Gillogly's (1980) alphabetic run in B1 is reproduced and, with the pamphlet's own numbering instead of a straight count, extends from his 14 monotone letters to 17 (`abcdefghiijklmmno`; the `f` and `h` he had to explain away as encoder errors disappear); Wase's (2020) non-uniform last digits are reproduced at the token level for all three ciphers (p < 0.0001); Campanelli's (2022) first-digit difference between B2 and B1/B3 is reproduced at the token level (permutation p = 0.00015 and < 0.0001); the Crossen (1927)/Kruh (1982) objection that 618 signs cannot hold thirty names with relatives and residences is confirmed. Hammer's 1971 claim that the ciphers' "signatures" show them to be real could not be checked against his statistics (his paper is not among our sources); what we can say is that B1 and B3 are indeed not i.i.d. random numbers, but that the genuine B2 is non-random in the same ways, so non-randomness does not separate a cipher from an imitation of one.

Beyond that literature: (1) the last-digit and first-digit anomalies of B2 are shown to be artefacts of homophone re-use - at the level of distinct values B2's digits are uniform (last digit p = 0.40, Benford first digit p = 0.70) and B3's too (p = 0.57, Benford p = 0.006 because of range), so the "B2 is non-uniform in every base but B1/B3 only in base 10" contrast has a mundane component that needs to be removed before the base-10 argument can be evaluated; a type-randomised null that keeps the observed re-use is introduced for this. (2) A quantitative encoder model is fitted to B2 (re-use probability as a function of homophones already in play, preferential re-use with weight count^0.6, preference for early key words with scale 119 words, slip rate 1.4 %) and used to show that B1 and B3 could not have been made with B2's habits with any key, but are compatible with a lighter re-use habit; the best-case calibrated encoders (r = 0.41-0.45 for B1, 0.59 for B3) are the first explicit "book cipher with another key" models the unsolved ciphers have been checked against on their whole frequency profile. (3) The serial structure is measured against shuffles, against letter-preserving shuffles of B2, and against simulated encoders: B2's own lag-1 autocorrelation is shown to be an encoder habit, B1's serial structure to be B2-like, and B3's to be an order of magnitude stronger, with rising runs, drift and clustering of large numbers that no simulated letter-by-letter encoder reproduces. (3a) A scan-forward habit (taking the next word in the key with the right initial) is shown to exist in B2 at a low rate (123 nearest-forward moves against 97 +- 7 expected, p = 0.0015) and, simulated as a mixed encoder, to account for B2's and B1's serial structure at one step in five to ten, while no rate accounts for B3's: its autocorrelation and its rising-run statistics call for incompatible rates, and its correlation persists over several lags where scanning gives a fast-decaying saw-tooth. (4) The capacity argument is put on a numerical footing with real Virginia name and county lengths and explicit scenarios.

## 1. Data, key and the cipher-2 reference (`stats_00_key_check.py`)

- A straight count of the pamphlet's Declaration gives 1321 words (1322 if "self-evident" is counted as two words, which is how Gillogly's Table I and the pamphlet itself count). The pamphlet's printed numbering deviates from a straight count in five places: its "(250)" follows an 11-word group (numbering 1 behind the split count from there), "(480)" is printed twice (11 behind), "(510)" follows a 9-word group (10 behind), "(640)" and "(680)" follow 11-word groups (11, then 12 behind up to the last printed marker, 816).
- Positional agreement of the decode of cipher 2 with the intended message over the first 570 numbers (before the printed "108" that stands for "10 8"): straight count 60.2 %, split count 79.3 %, pamphlet numbering 97.4 %, pamphlet numbering with the encoder's three conventions 811 = y, 1005 = x, 95 = u ("unalienable") 99.5 %. A Needleman-Wunsch alignment of the last ("resolved") decode with the intended message (763 letters, digits spelled out as the decode shows: "ten hundred", "thirty eight hundred", "nov", "dec", "in exchange to save") leaves 11 substitution errors, 7 of them numbers written one too high or too low for a correct homophone (`error_kinds` in `results/stats_key_check.json`), and one skipped letter. The intended letter for each of the 762 numbers, from that alignment, is the basis of the encoder model in section 5. Where a test depends on the exact key it is said so; the key reconstruction of `notes/cryptanalysis_b2_gillogly.md` (`results/b2_decode_summary.json`) agrees on the numbering and its offsets.

## 2. Descriptive statistics (`stats_01_descriptive.py`; figure `results/stats_hist_values.svg`)

| statistic | B1 (locality) | B2 (contents, solved) | B3 (names) |
|---|---|---|---|
| numbers (tokens) | 520 | 762 | 618 |
| minimum | 1 | 1 | 1 |
| maximum | 2906 | 1005 | 975 |
| mean | 273.4 | 161.7 | 154.3 |
| median | 123 | 85 | 94 |
| s.d. | 356.6 | 199.7 | 179.2 |
| distinct values (types) | 298 | 181 | 263 |
| distinct / n | 0.573 | 0.238 | 0.426 |
| tokens per type | 1.74 | 4.21 | 2.35 |
| types used exactly once | 177 | 43 | 124 |
| share of types used once | 0.594 | 0.238 | 0.471 |
| share of tokens that are singletons | 0.340 | 0.056 | 0.201 |
| token share of the 10 commonest values | 0.117 | 0.186 | 0.142 |
| numbers > 1322 (DOI length) | 10 | 0 | 0 |
| share > 1322 | 0.019 | 0.000 | 0.000 |
| numbers > 1005 (largest in B2) | 18 | 0 | 0 |
| share <= 100 | 0.437 | 0.531 | 0.539 |
| share <= 200 | 0.590 | 0.752 | 0.739 |
| share <= 500 | 0.829 | 0.904 | 0.953 |
| share <= 1000 | 0.965 | 0.995 | 1.000 |
| share even (tokens) | 0.590 | 0.528 | 0.531 |
| share even (types) | 0.564 | 0.525 | 0.502 |

Quantiles (10/25/50/75/90 %): B1 18/56/123/346/814; B2 14/37/85/199/483; B3 19/48/94/204/319.

Most re-used values: B1 18 x8, 216 x7, 19 x7, 16 x6, 84 x6, 81 x6, 88 x6, 71 x5, 38 x5, 11 x5, 63 x5, 64 x5, 10 x5, 36 x5, 34 x5. B2 807 x18 (the only v-word, "valuable"), 7 x15, 140 x15, 106 x15, 138 x15, 37 x13, 47 x13, 44 x13, 33 x13, 10 x12, 85 x11, 8 x11, 63 x11, 110 x11, 52 x10. B3 96 x13, 18 x11, 89 x10, 19 x9, 66 x9, 81 x8, 28 x7, 77 x7, 44 x7, 11 x7, 82 x7, 84 x7, 65 x6, 48 x6, 218 x6.

The ten values above 1322 in B1: 1431, 1496, 1629, 1701, 1706, 1780, 1817, 2018, 2160, 2906 - a sparse tail (1.9 % of the numbers) above a body that is even more concentrated on small numbers than B2's (median 123 for a key that would have to be at least 2906 words long).

| digit table | B1 tokens | B2 tokens | B3 tokens | B1 types | B2 types | B3 types |
|---|---|---|---|---|---|---|
| first digit 1 | 142 | 219 | 172 | 74 | 48 | 69 |
| first digit 2 | 88 | 111 | 109 | 52 | 33 | 58 |
| first digit 3 | 59 | 106 | 76 | 31 | 25 | 41 |
| first digit 4 | 50 | 84 | 47 | 30 | 20 | 20 |
| first digit 5 | 23 | 73 | 24 | 18 | 18 | 9 |
| first digit 6 | 44 | 49 | 43 | 27 | 16 | 17 |
| first digit 7 | 24 | 39 | 38 | 16 | 7 | 11 |
| first digit 8 | 57 | 63 | 65 | 28 | 8 | 20 |
| first digit 9 | 33 | 18 | 44 | 22 | 6 | 18 |
| last digit 0 | 68 | 129 | 32 | 41 | 26 | 22 |
| last digit 1 | 72 | 76 | 69 | 41 | 20 | 30 |
| last digit 2 | 46 | 70 | 57 | 29 | 17 | 22 |
| last digit 3 | 39 | 76 | 49 | 25 | 14 | 22 |
| last digit 4 | 67 | 56 | 75 | 35 | 17 | 33 |
| last digit 5 | 36 | 80 | 55 | 22 | 22 | 28 |
| last digit 6 | 71 | 75 | 91 | 34 | 17 | 32 |
| last digit 7 | 29 | 99 | 60 | 23 | 20 | 30 |
| last digit 8 | 55 | 72 | 73 | 29 | 18 | 23 |
| last digit 9 | 37 | 29 | 57 | 19 | 10 | 21 |
| 1-digit numbers | 16 | 51 | 16 | | | |
| 2-digit numbers | 211 | 354 | 317 | | | |
| 3-digit numbers | 274 | 353 | 285 | | | |
| 4-digit numbers | 19 | 4 | 0 | | | |

## 3. Digit preferences (`stats_02_digits.py`; figure `results/stats_last_digits.svg`)

Monte-Carlo p-values from 100 000 multinomial draws (seed 20260914); pairwise comparisons by chi-square homogeneity with 20 000 permutations.

| test | B1 | B2 | B3 |
|---|---|---|---|
| last digit vs uniform, all numbers: chi2 (9 df), p | 47.4, p < 0.0001 | 78.9, p < 0.0001 | 38.1, p < 0.0001 |
| same, type-randomised null (re-use structure kept), p | 0.035 | 0.279 | 0.360 |
| last digit vs uniform, distinct values: chi2, p | 18.2, p = 0.032 | 9.4, p = 0.40 | 7.7, p = 0.57 |
| share even, all numbers (binomial p) | 0.590 (p = 0.00004) | 0.528 (p = 0.14) | 0.531 (p = 0.14) |
| share even, distinct values (binomial p) | 0.564 (p = 0.032) | 0.525 (p = 0.55) | 0.502 (p = 1) |
| share ending in 0 or 5, distinct values (p vs 0.20) | 0.211 (p = 0.61) | 0.265 (p = 0.032) | 0.190 (p = 0.76) |
| share ending in 0, distinct values (p vs 0.10) | 0.138 (p = 0.042) | 0.144 (p = 0.062) | 0.084 (p = 0.47) |
| repeated final digits (dd), distinct values (p vs 0.10) | 0.062 (p = 0.031) | 0.116 (p = 0.45) | 0.082 (p = 0.40) |
| last two digits uniform, distinct values (MC p) | 0.41 | 0.98 | 0.89 |
| first digit vs Benford, all numbers: chi2, p | 52.1, p < 0.0001 | 33.4, p < 0.0001 | 60.8, p < 0.0001 |
| first digit vs Benford, distinct values: chi2, p | 23.6, p = 0.003 | 5.6, p = 0.70 | 21.5, p = 0.006 |
| first digit vs uniform on [1, 1322], distinct values | p < 0.0001 | p < 0.0001 | p < 0.0001 |

| comparison | last digit, all numbers | last digit, distinct values | first digit, all numbers | first digit, distinct values | share even (Fisher p) |
|---|---|---|---|---|---|
| B1 vs B2 | chi2 = 51.3, p < 0.0001 | chi2 = 6.1, p = 0.74 | chi2 = 31.1, p = 0.00015 | chi2 = 11.1, p = 0.20 | 0.590 vs 0.528 (p = 0.03) |
| B3 vs B2 | chi2 = 79.4, p < 0.0001 | chi2 = 6.8, p = 0.67 | chi2 = 42.6, p < 0.0001 | chi2 = 14.9, p = 0.06 | 0.531 vs 0.528 (p = 0.91) |
| B1 vs B3 | chi2 = 31.6, p = 0.0005 | chi2 = 9.0, p = 0.44 | chi2 = 4.2, p = 0.84 | chi2 = 9.7, p = 0.29 | 0.590 vs 0.531 (p = 0.05) |

What this does and does not show.
- Counting every number, all three ciphers have non-uniform last digits (p < 0.0001 each, agreeing with Wase 2020: 0.4 %, 0.02 %, 0.4 % by his KS test). But a homophonic cipher's numbers are not independent draws: B2's excess of 0s and 7s and deficit of 9s is made by 807 (18 times), 140, 10, 110, 540, 7, 37, 47 being repeated. Under a null that keeps each distinct value's multiplicity and gives it a random last digit, B2's token-level statistic is unremarkable (p = 0.28) and so is B3's (p = 0.36); B1's stays mildly significant (p = 0.035). At the level of distinct values only B1 deviates (p = 0.03), through an excess of even digits (56 % of its 298 values; 59 % of all its numbers) and few 5s, 7s and 9s. B1 also has fewer repeated-final-digit values (11, 22, 33 ...) than chance (6.2 % against 10 %, p = 0.03), a known trait of people producing "random" numbers (Wagenaar 1972), while B2 and B3 have the expected share. These are three results at the 3-5 % level out of about a dozen B1 digit tests; taken together they are suggestive, not conclusive, of a hand choosing B1's digits.
- The range of a book cipher's numbers is set by the key and the encoder's preference for early words, so Benford's law is not an appropriate expectation: B2 itself fails it at the token level (p < 0.0001) and is only Benford-like at the type level (p = 0.70) because its distinct values happen to spread over 1-1005 with a decreasing density. B1's and B3's first digits differ from B2's when every number is counted (Campanelli 2022's finding), but at the level of distinct values the difference is within chance for B1 (p = 0.20) and marginal for B3 (p = 0.06); most of the token-level difference is again the re-use structure.
- Numbers ending in 0: B2 has 14.4 % of distinct values ending in 0 (p = 0.06) and 26.5 % ending in 0 or 5 (p = 0.03). The pamphlet's key is numbered every ten words, so the marked words were the easiest to find; this is a process signature of the B2 encoder. B1 shows the same excess of 0-endings (13.8 %, p = 0.04), B3 a deficit (8.4 %). Either a key numbered by tens was used for B1 as well, or its author liked round numbers; the two cannot be told apart here.
- For B1's ten numbers above 1322 the last digits are compatible with uniform (p = 0.43) and 6 of 10 are even; nothing can be said from ten numbers.

## 4. Sequential structure (`stats_03_sequential.py`, `stats_08_serial_extras.py`; figure `results/stats_sequence_plots.svg`)

Each statistic is compared with 20 000 shuffles of the same cipher (two-sided unless stated; the Monte-Carlo standard error of a p-value near 0.05 is 0.0015).

| statistic | B1 obs (null mean +- sd), p | B2 obs (null), p | B3 obs (null), p |
|---|---|---|---|
| lag-1 autocorrelation (Pearson) | 0.252 (0.00 +- 0.044), p < 0.0001 | 0.046 (0.00 +- 0.036), p = 0.19 | 0.622 (0.00 +- 0.040), p < 0.0001 |
| lag-2 autocorrelation (Pearson) | 0.027 (0.00 +- 0.043), p = 0.51 | -0.002 (0.00 +- 0.036), p = 0.97 | 0.454 (0.00 +- 0.040), p < 0.0001 |
| lag-1 rank autocorrelation | 0.195 (0.00 +- 0.044), p = 0.0001 | 0.206 (0.00 +- 0.036), p < 0.0001 | 0.489 (0.00 +- 0.040), p < 0.0001 |
| lag-2 rank autocorrelation | 0.001 (0.00 +- 0.044), p = 0.96 | 0.068 (0.00 +- 0.036), p = 0.058 | 0.332 (0.00 +- 0.040), p < 0.0001 |
| runs up and down (count) | 317 (346 +- 9.5), p = 0.003 | 486 (506 +- 12), p = 0.10 | 332 (410 +- 11), p < 0.0001 |
| runs test, asymptotic z (p) | -3.06 (0.002) | -1.86 (0.062) | -7.61 (3e-14) |
| adjacent pairs with |a-b| <= 5 (one-sided, more) | 10 (15.5 +- 3.8), p = 0.95 | 51 (36.4 +- 5.8), p = 0.009 | 26 (26.2 +- 4.9), p = 0.55 |
| adjacent pairs with |a-b| <= 10 (one-sided, more) | 32 (28.2 +- 5.0), p = 0.25 | 77 (63.8 +- 7.3), p = 0.044 | 76 (49.4 +- 6.5), p < 0.0001 |
| adjacent pairs with |a-b| <= 1 | 0 (4.5 +- 2.1) | 16 (11.8 +- 3.4), p = 0.13 | 1 (7.9 +- 2.8) |
| adjacent pairs with the same last digit | 33 (55.8 +- 7.0), p = 0.001 | 58 (83 +- 8.5), p = 0.004 | 28 (64.7 +- 7.6), p < 0.0001 |
| adjacent pairs with the same first digit | 53 (78.3 +- 7.7), p = 0.001 | 142 (120 +- 9.5), p = 0.021 | 114 (95.2 +- 8.4), p = 0.03 |
| trend: Spearman rho(position, value) | 0.029, p = 0.50 | 0.029, p = 0.42 | 0.169, p = 0.0004 |
| mean of first half minus second half | -33.7 (0 +- 31), p = 0.29 | -13.5 (0 +- 14), p = 0.35 | -80.9 (0 +- 14), p < 0.0001 |
| positions of large numbers (B1: > 1322, n = 10; else top decile): KS vs uniform p | 0.48 | 0.76 | < 0.001 |
| variance of gaps between large numbers | 4990 (1790 +- 1100), p = 0.02 | 60 (84 +- 20), p = 0.18 | 312 (91 +- 24), p = 0.0002 |
| longest strictly rising run of numbers | 7 (5.5 +- 0.7), p = 0.07 | 6 (5.7 +- 0.7), p = 0.55 | 10 (5.6 +- 0.7), p = 0.0002 |
| rising runs of length >= 4 | 36 (17 +- 3.6), p = 0.0002 | 39 (24 +- 4.5), p = 0.001 | 49 (20 +- 4.1), p = 0.0002 |
| rising runs of length >= 6 | 3 (0.6 +- 0.7), p = 0.017 | 2 (0.8 +- 0.9), p = 0.19 | 17 (0.7 +- 0.8), p = 0.0002 |
| mean length of rising runs | 2.27 (1.99 +- 0.05), p = 0.0002 | 2.12 (1.98 +- 0.04), p = 0.001 | 2.58 (1.99 +- 0.05), p = 0.0002 |

Reading these results.
- The premise that a genuine book cipher has no serial correlation is false for this encoder: B2 has a lag-1 rank autocorrelation of 0.21 (z = 5.7), more close pairs than chance and more rising runs of four or more. Permuting B2's numbers only within letter classes (same plaintext letter, `stats_08_serial_extras.py`, 5000 permutations) removes it entirely (null -0.02 +- 0.035, p = 0.0002; close pairs 33 +- 5 expected against 51 observed, p = 0.001), so it is not an effect of English digraphs acting through letter-specific number ranges but a habit of the encoder's hand - after writing a number he tended to take the next one from the same neighbourhood of the key. None of the memoryless simulated encoders of section 5 reproduces it (simulated lag-1 rank autocorrelation -0.01 +- 0.04).
- B1's serial structure is of B2's kind and size: lag-1 rank 0.195 against 0.206, lag-2 nil in both, 36 rising runs of >= 4 (B2: 39), mean rising-run length 2.27 (B2: 2.12), no trend, no clustering of its large numbers (KS p = 0.48; the gap-variance p = 0.02 reflects ten numbers only). It lacks B2's close-pair excess (10 pairs within 5 against 15.5 expected). It is fair to say B1 is serially structured "like the genuine cipher", which neither supports nor undermines its being one.
- B3 is different in degree and in kind: lag-1 autocorrelation 0.62, lag-2 0.45 (both z > 11), 78 fewer runs up and down than chance (z = -7.6), 17 strictly rising runs of six or more numbers (chance: 0.7), a rise of the mean between the first and second half of 81 (p < 0.0001), and its largest numbers clustered together in position (KS p < 0.001, gap variance 3.4 times chance). The sequence plot (`results/stats_sequence_plots.svg`) shows the shape: sawtooth stretches such as `112 136 149 176 180 194 143 171 205 296` and, at the end, `10 22 18 46 137 181 101 39 86 103 116 138 164 212 218 296 815 380 412 460 495 675 820 952`. B2 never does this; its encoder's local habit produces rises of two or three. The excess of rising over falling steps (61 % of B3's steps rise; B1 56 %; B2 53 %; chance 50 %) is what a person writing numbers "at random" produces - counting upward is the path of least resistance - and what a letter-by-letter encoder does not, since plaintext letters arrive in an order unrelated to their positions in the key. Section 5.4 tests the one encoder habit that could produce rises (scanning the key forward from the last word used).
- Too few adjacent pairs share a last digit in all three ciphers (B1 33 against 56, B2 58 against 83, B3 28 against 65). For B2 and B3 this follows from the close-pair excess (two numbers within ten of each other rarely share a last digit); for B1, which has no close-pair excess, it is a second small "digit avoidance" signal in line with section 3.
- Alphabetic runs in the DOI decode (letters obtained by decoding each cipher with the DOI, longest non-decreasing run, against 5000 shuffles of the same letters):

| key | B1 | B2 | B3 |
|---|---|---|---|
| pamphlet numbering (the numbering that decodes cipher 2) | 17 (`abcdefghiijklmmno` at 188), null 6.3, p = 0.0002 | 6, null 6.4, p = 0.88 | 7, null 6.5, p = 0.44 |
| straight count with "self-evident" split (Gillogly's Table I) | 8 (`efghiijp` at 192), null 6.3, p = 0.10 | 7, p = 0.39 | 7, p = 0.51 |
| straight count | 6, p = 0.86 | 7, p = 0.62 | 5, p = 1 |

Gillogly saw `ABFDE FGHII JKLMM NOHPP` because his Table I is one word ahead of the pamphlet's numbering between words 160 and 240 (his text lacks the "a" of "institute a new government"); with the pamphlet's numbering 195 is "changed" (c) and the run reads `abcdefghiijklmmno` + `hpp` (Wikipedia already quotes the corrected string). Under the pamphlet numbering B1 also has 105 adjacent letter pairs that are equal or alphabetic successors against 62 expected (p = 0.0002), but all of that excess is the same run; B3 shows nothing (77 against 79). The run's letters are the initials of DOI words 147, 436, 195, 320, 37, 122, 113, 6, 140, 8, 120, 305, 42, 58, 461, 44, 106: someone picked, for each successive letter of the alphabet, one DOI word beginning with it. That is a use of the numbered DOI as a source of numbers, not an encipherment.

## 5. Homophone-reuse structure (`stats_04_homophone_sim.py`, `stats_09_calibrated_reuse.py`, `stats_simlib.py`; figures `results/stats_growth_curves.svg`, `results/stats_growth_curves_calibrated.svg`)

### 5.1 The encoder of cipher 2

From the 762 (number, intended letter) pairs of section 1 (key = pamphlet numbering with 811 = y, 1005 = x, 95 = u; the DOI has no word beginning with x, y or z, and the encoder never needed z):
- Re-use. When a letter already had d distinct numbers in play and an unused homophone was still available, the encoder re-used one of the d with probability 0.46 (d = 1, n = 37), 0.15 (d = 2, n = 22), 0.43 (d = 3, n = 33), 0.59 (d = 4-5, n = 79), 0.83 (d = 6-8, n = 226), 0.86 (d = 9-12, n = 178), 0.94 (d >= 13, n = 152); overall 0.78. A logistic fit gives logit p = -1.28 + 1.37 ln d. Among the numbers already in play he preferred the ones he had used most (weight count^gamma, gamma = 0.6 by maximum likelihood on a 0-3 grid).
- New numbers. A fresh homophone was chosen with weight exp(-n/tau), tau = 119 words by maximum likelihood (log-likelihood -457 against -702 for uniform choice among unused homophones; a rank-based alternative, weight exp(-rank/6.1) over the letter's homophones in key order, is slightly worse at -479). Hence 53 % of B2's numbers are <= 100 although only 8 % of the key's words are.
- Slips: 11 of 762 numbers (1.4 %) are off by one or otherwise wrong for the intended letter.
- Model check on B2's own text (1000 simulations): distinct numbers 196 +- 7 against 181 observed (2nd percentile), share of types used once 0.32 +- 0.03 against 0.24, commonest number 21.5 +- 3.9 against 18 (23rd percentile), share of tokens in the ten commonest 0.193 +- 0.014 against 0.186 (34th), Zipf slope -0.32 +- 0.05 against -0.29. The model re-uses slightly less than the real encoder did but reproduces the profile; it does not reproduce his serial habit (section 4), his liking for numbers ending in 0 (simulated 10.6 % of types against 14.4 %), or the full strength of his preference for small numbers (simulated median 135 +- 14 against 85; share of numbers <= 100 0.41 against 0.53): he re-used his early, small homophones more than the count^0.6 rule gives. The model is therefore a conservative description of the small-number habit, and the median is not used below to separate the ciphers.

### 5.2 B1 and B3 against simulated encoders

1000 simulations per model; plaintexts are random 520- or 618-letter windows of the three "Beale" letters (12 100 letters of prose in the same hand), and for B3 also a synthetic 30-entry list of real 1810/1850 Virginia names with counties (section 7). Keys: the DOI as numbered in the pamphlet (1310 numbers); hypothetical keys of 2906 and 975 words whose initials are those of the pamphlet's own narrative prose. Reference nulls: i.i.d. uniform numbers on [1, max] and i.i.d. log-uniform numbers on [1, max] (a "human-invented numbers" stand-in: scale-free preference for small numbers, Benford first digits, no memory - the simplest formal model of a person writing numbers without a text; a person's actual serial habits are addressed in section 4 rather than modelled). Entries: observed | simulated mean +- sd (percentile of the observed value among the simulations).

**B1** (observed: 298 distinct; 0.594 of types once; commonest 8; top-10 share 0.117; Zipf slope -0.30; lag-1 rank autocorrelation 0.195; runs z -3.06; 10 close pairs; 0.138 of types end in 0)

| statistic | DOI key, fitted encoder | DOI key, random homophones | 2906-word key, fitted | 2906-word key, random | i.i.d. uniform | i.i.d. log-uniform |
|---|---|---|---|---|---|---|
| distinct numbers | 164 +- 6 (1.000) | 377 +- 8 (0.000) | 169 +- 7 (1.000) | 450 +- 7 (0.000) | 476 +- 6 (0.000) | 274 +- 10 (0.993) |
| share of types used once | 0.36 +- 0.03 (1.000) | 0.75 +- 0.02 (0.000) | 0.37 +- 0.03 (1.000) | 0.87 +- 0.02 (0.000) | 0.91 +- 0.01 (0.000) | 0.77 +- 0.02 (0.000) |
| commonest number | 17 +- 3 (0.000) | 12 +- 4 (0.30) | 16 +- 3 (0.000) | 4.0 +- 0.7 (1.000) | 3.0 +- 0.4 (1.000) | 45 +- 6 (0.000) |
| top-10 share | 0.22 +- 0.02 (0.000) | 0.097 +- 0.010 (0.985) | 0.21 +- 0.02 (0.000) | 0.057 +- 0.005 (1.000) | 0.043 +- 0.003 (1.000) | 0.31 +- 0.02 (0.000) |
| Zipf slope | -0.36 +- 0.05 (0.88) | -0.35 +- 0.07 (0.75) | -0.35 +- 0.05 (0.83) | -0.22 +- 0.04 (0.03) | -0.09 +- 0.05 (0.000) | -0.85 +- 0.05 (1.000) |
| lag-1 rank autocorrelation | -0.01 +- 0.05 (1.000) | 0.00 +- 0.04 (1.000) | 0.02 +- 0.05 (1.000) | 0.00 +- 0.04 (1.000) | -0.01 +- 0.05 (1.000) | 0.00 +- 0.04 (1.000) |
| close pairs |a-b| <= 5 | 18 +- 4 (0.05) | 4.4 +- 2.1 (0.99) | 17 +- 4 (0.05) | 1.9 +- 1.3 (1.000) | 1.9 +- 1.3 (1.000) | 48 +- 7 (0.000) |
| share of types ending in 0 | 0.10 +- 0.02 (0.99) | 0.10 +- 0.01 (1.000) | 0.10 +- 0.02 (0.99) | 0.10 +- 0.01 (1.000) | 0.10 +- 0.01 (1.000) | 0.10 +- 0.02 (0.99) |

**B3** (observed: 263 distinct; 0.471 once; commonest 13; top-10 share 0.142; Zipf -0.33; lag-1 rank 0.489; runs z -7.61; 26 close pairs; 0.084 of types end in 0)

| statistic | DOI key, fitted encoder | DOI key, random homophones | 975-word key, fitted | 975-word key, random | i.i.d. uniform | i.i.d. log-uniform | DOI, fitted, names-list text | 975-word, fitted, names-list text |
|---|---|---|---|---|---|---|---|---|
| distinct numbers | 179 +- 7 (1.000) | 429 +- 9 (0.000) | 183 +- 7 (1.000) | 403 +- 9 (0.000) | 458 +- 8 (0.000) | 256 +- 10 (0.80) | 179 +- 7 (1.000) | 178 +- 6 (1.000) |
| share of types used once | 0.34 +- 0.03 (1.000) | 0.73 +- 0.02 (0.000) | 0.34 +- 0.03 (1.000) | 0.66 +- 0.02 (0.000) | 0.72 +- 0.02 (0.000) | 0.67 +- 0.02 (0.000) | 0.35 +- 0.03 (1.000) | 0.33 +- 0.03 (1.000) |
| commonest number | 19 +- 3 (0.02) | 14 +- 5 (0.42) | 19 +- 3 (0.03) | 6.7 +- 1.0 (1.000) | 4.4 +- 0.6 (1.000) | 63 +- 7 (0.000) | 22 +- 2 (0.000) | 18 +- 3 (0.05) |
| top-10 share | 0.21 +- 0.02 (0.000) | 0.094 +- 0.009 (1.000) | 0.20 +- 0.02 (0.000) | 0.080 +- 0.006 (1.000) | 0.056 +- 0.004 (1.000) | 0.36 +- 0.02 (0.000) | 0.22 +- 0.01 (0.000) | 0.20 +- 0.02 (0.000) |
| Zipf slope | -0.36 +- 0.05 (0.71) | -0.36 +- 0.05 (0.74) | -0.35 +- 0.05 (0.62) | -0.25 +- 0.04 (0.02) | -0.17 +- 0.06 (0.003) | -0.87 +- 0.05 (1.000) | -0.38 +- 0.04 (0.90) | -0.33 +- 0.05 (0.48) |
| lag-1 rank autocorrelation | -0.01 +- 0.04 (1.000) | 0.00 +- 0.04 (1.000) | 0.02 +- 0.04 (1.000) | 0.00 +- 0.04 (1.000) | 0.00 +- 0.04 (1.000) | 0.00 +- 0.04 (1.000) | 0.02 +- 0.04 (1.000) | 0.01 +- 0.04 (1.000) |
| close pairs |a-b| <= 5 | 21 +- 5 (0.89) | 5.2 +- 2.3 (1.000) | 20 +- 5 (0.91) | 7.0 +- 2.7 (1.000) | 7.0 +- 2.7 (1.000) | 76 +- 10 (0.000) | 24 +- 5 (0.70) | 19 +- 5 (0.94) |

The full tables (thirteen statistics per model, growth curves at every twenty numbers, top-ten count profiles) are in `results/stats_homophone_sim.json` and `results/stats_summary_tables.md`.

Reading these results.
- Neither unsolved cipher was made with the B2 encoder's habits, whatever the key: his re-use would leave 164-183 distinct numbers where B1 has 298 and B3 263 (the observed values are beyond the 100th percentile of 1000 simulations in every fitted-encoder model). Nor were they made by a memoryless homophone chooser (377-450 distinct). The growth curves (`stats_growth_curves.svg`) show B1 and B3 running between the two bands from the first fifty numbers on, so this is not a matter of a change of habit part-way through.
- The i.i.d. log-uniform "invented numbers" null matches the distinct counts (B1 274 +- 10, B3 256 +- 10) by coincidence of its small-number weight, but is refuted on the profile: its commonest number would appear 45-63 times (observed 8 and 13), its ten commonest would take 31-36 % of the numbers (observed 12-14 %), and it would have 48-76 adjacent pairs within 5 (observed 10 and 26). Whoever produced B1 and B3 did not simply favour small numbers at random; the frequency profile is flat in the way a homophonic substitution with limited re-use is.
- The names-list plaintext behaves exactly like prose for these statistics (B3 columns 7-8), so the list nature of paper 3 does not change the picture.

### 5.3 The best-case homophonic model (`stats_09_calibrated_reuse.py`)

Holding gamma, tau and the slip rate at their B2 values (tau scaled with key length for the long keys), the re-use probability was replaced by a constant r calibrated so that the simulated number of distinct values matches the observed one on average, and the remaining statistics were compared (1000 simulations).

| statistic | B1 obs | B1, DOI, r = 0.41 | B1, 2906 words, r = 0.45 | B3 obs | B3, DOI, r = 0.59 | B3, 975 words, r = 0.59 |
|---|---|---|---|---|---|---|
| distinct numbers | 298 | 299 +- 11 (0.47) | 300 +- 11 (0.46) | 263 | 263 +- 12 (0.51) | 262 +- 11 (0.56) |
| share of types used once | 0.594 | 0.628 +- 0.027 (0.11) | 0.631 +- 0.025 (0.08) | 0.471 | 0.516 +- 0.029 (0.06) | 0.509 +- 0.028 (0.09) |
| commonest number | 8 | 12.7 +- 3.2 (0.10) | 10.3 +- 2.2 (0.20) | 13 | 17.3 +- 3.2 (0.12) | 16.4 +- 3.5 (0.19) |
| top-10 share | 0.117 | 0.136 +- 0.013 (0.09) | 0.131 +- 0.013 (0.16) | 0.142 | 0.177 +- 0.018 (0.025) | 0.171 +- 0.016 (0.04) |
| Zipf slope | -0.30 | -0.38 +- 0.07 (0.88) | -0.34 +- 0.06 (0.74) | -0.33 | -0.40 +- 0.06 (0.87) | -0.38 +- 0.06 (0.81) |
| top-10 counts | 8 7 7 6 6 6 6 5 5 5 | 12.7 9.0 7.7 7.0 6.5 6.1 5.8 5.6 5.3 5.1 | 10.3 8.5 7.6 7.0 6.5 6.2 5.9 5.6 5.4 5.2 | 13 11 10 9 9 8 7 7 7 7 | 17.3 14.2 12.4 11.2 10.4 9.7 9.2 8.7 8.3 8.0 | 16.4 13.5 11.9 10.9 10.1 9.5 9.0 8.6 8.2 7.9 |
| lag-1 rank autocorrelation | 0.195 | 0.067 +- 0.051 (0.99) | 0.035 +- 0.046 (1.000) | 0.489 | 0.039 +- 0.046 (1.000) | 0.078 +- 0.047 (1.000) |
| runs-up-and-down z | -3.06 | 0.23 +- 0.99 (0.001) | -0.08 +- 0.97 (0.002) | -7.61 | 0.32 +- 1.0 (0.000) | -0.28 +- 1.0 (0.000) |
| close pairs |a-b| <= 5 | 10 | 12.1 +- 3.5 (0.34) | 8.6 +- 2.9 (0.75) | 26 | 17.3 +- 4.4 (0.98) | 20 +- 5 (0.91) |
| share of numbers <= 100 | 0.437 | 0.318 +- 0.023 (1.000) | 0.23 +- 0.02 (1.000) | 0.539 | 0.366 +- 0.032 (1.000) | 0.42 +- 0.03 (1.000) |
| median | 123 | 181 +- 12 (0.000) | 267 +- 24 (0.000) | 94.5 | 155 +- 14 (0.000) | 131 +- 13 (0.000) |

So a homophonic encoder who re-used numbers about half as often as the B2 encoder (r = 0.41-0.45) and, for B3, three-quarters as often (0.59), reproduces every frequency-profile statistic of B1 (percentiles 0.08-0.90) and, slightly less comfortably, of B3 (0.025-0.87; its ten commonest numbers are a little too even). Both ciphers also lean to small numbers more than the calibrated models give (medians 123 and 94.5 against 181-267 and 131-155), but B2 does the same against its own fitted model (85 against 135, section 5.1), so this is a shortcoming of the exp(-n/tau) rule rather than a difference between the ciphers and is not counted as evidence either way. A related artefact matters below: because that rule spends the early words first, the simulated sequences drift upward (rank correlation of value with position about 0.28 in the B1 models and 0.23-0.27 in the B3 models, against 0.03 observed in B1 and B2 and 0.17 in B3), so the drift statistics of section 4 cannot be compared with these models and are not. What the calibrated encoders do not reproduce, for either cipher, is the serial structure - but as section 4 showed, they do not reproduce B2's either, and B1's is B2-sized while B3's is not.

### 5.4 Could a scan-forward habit explain B3? (`stats_10_scan_forward.py`, `stats_11_acf.py`; figure `results/stats_acf.svg`)

A book-cipher encoder who finds each next letter by reading on in the key from the word he last used, and takes the first word with the right initial, produces rising runs of numbers, positive autocorrelation and drift without inventing anything. Since those are B3's distinguishing features, this habit was tested directly.

**Does the B2 encoder scan forward?** For each of the 761 steps of B2 the number was compared with the nearest homophone of the intended letter after the previous number (the first word with the required initial that follows the word just used). B2 takes that word 123 times; letter-preserving shuffles (2000; numbers permuted within letter classes, so the homophone sets and the message are unchanged) give 97.3 +- 6.9 (p = 0.0015). The nearest homophone before the previous word is taken 65 times against 58.0 +- 6.5 (p = 0.16). The B2 encoder did read forward from where he was, but only about 25 times more than chance in 761 steps: a real, modest habit.

**A mixed scan-forward encoder.** At each letter the simulated encoder takes, with probability q, the nearest forward homophone of the letter (the first homophone in the key if none is left) and otherwise behaves as the encoders of sections 5.1 and 5.3 (B2: the fitted encoder; B1 and B3: the calibrated encoders, r = 0.41 and 0.59, with the DOI key and with the 2906- and 975-word keys). q runs over 0, 0.1, ..., 0.8 with 300 simulations per value (seed 20260916; plaintexts as in 5.2). The statistics recorded are the ones that distinguish B3: lag-1 and lag-2 autocorrelation (Pearson and rank), the mean length and the number (>= 4, >= 6) of strictly rising runs, the share of rising steps, the runs-up-and-down z, adjacent pairs within 10, distinct count and median. Entries are simulated mean +- sd (percentile of the observed value); the full grid, including the long keys, is in `results/stats_scan_forward.json` and `stats_summary_tables.md`.

B2 (DOI key, fitted re-use):

| statistic | B2 obs | q = 0 | q = 0.1 | q = 0.2 | q = 0.3 |
|---|---|---|---|---|---|
| lag-1 autocorrelation (Pearson) | 0.046 | -0.030 +- 0.034 (0.98) | 0.080 +- 0.047 (0.22) | 0.182 +- 0.052 (0.00) | 0.291 +- 0.049 (0.00) |
| lag-1 autocorrelation (rank) | 0.206 | -0.012 +- 0.035 (1.000) | 0.084 +- 0.040 (1.000) | 0.178 +- 0.039 (0.77) | 0.276 +- 0.038 (0.03) |
| mean rising-run length | 2.12 | 1.99 +- 0.04 (1.000) | 2.18 +- 0.05 (0.15) | 2.41 +- 0.07 (0.00) | 2.70 +- 0.09 (0.00) |
| rising runs >= 4 | 39 | 24.6 +- 4.5 (1.000) | 39.9 +- 4.7 (0.46) | 56.7 +- 5.9 (0.00) | 70.9 +- 5.6 (0.00) |
| rising runs >= 6 | 2 | 0.8 +- 0.9 (0.96) | 3.4 +- 1.7 (0.34) | 7.9 +- 2.6 (0.01) | 16.2 +- 3.4 (0.00) |
| share of rising steps | 0.530 | 0.499 +- 0.011 (1.000) | 0.541 +- 0.011 (0.15) | 0.586 +- 0.012 (0.00) | 0.630 +- 0.012 (0.00) |
| runs-up-and-down z | -1.86 | 0.47 +- 1.0 (0.01) | -1.59 +- 0.95 (0.39) | -4.02 +- 1.1 (0.97) | -6.81 +- 1.1 (1.000) |
| adjacent pairs within 10 | 77 | 43 +- 6.5 (1.000) | 69 +- 8.3 (0.87) | 93 +- 9.5 (0.05) | 119 +- 11 (0.00) |
| distinct numbers | 181 | 194 +- 7 (0.06) | 208 +- 8 (0.00) | 222 +- 9 (0.00) | 236 +- 10 (0.00) |

B1 (DOI key, r = 0.41; the 2906-word key gives the same picture):

| statistic | B1 obs | q = 0 | q = 0.1 | q = 0.2 | q = 0.3 |
|---|---|---|---|---|---|
| lag-1 autocorrelation (Pearson) | 0.252 | 0.062 +- 0.047 (1.000) | 0.144 +- 0.055 (0.98) | 0.232 +- 0.058 (0.61) | 0.322 +- 0.060 (0.10) |
| lag-1 autocorrelation (rank) | 0.195 | 0.069 +- 0.048 (1.000) | 0.147 +- 0.050 (0.83) | 0.232 +- 0.051 (0.25) | 0.312 +- 0.052 (0.01) |
| lag-2 autocorrelation (Pearson) | 0.026 | 0.074 +- 0.050 (0.15) | 0.078 +- 0.050 (0.14) | 0.105 +- 0.052 (0.07) | 0.139 +- 0.059 (0.03) |
| mean rising-run length | 2.27 | 2.00 +- 0.05 (1.000) | 2.19 +- 0.07 (0.89) | 2.42 +- 0.08 (0.02) | 2.73 +- 0.11 (0.00) |
| rising runs >= 4 | 36 | 16.8 +- 3.8 (1.000) | 28.1 +- 4.6 (0.97) | 38.7 +- 4.4 (0.31) | 48.6 +- 4.6 (0.00) |
| rising runs >= 6 | 3 | 0.7 +- 0.9 (0.99) | 2.4 +- 1.4 (0.78) | 5.8 +- 2.3 (0.16) | 11.3 +- 3.0 (0.00) |
| share of rising steps | 0.561 | 0.501 +- 0.013 (1.000) | 0.545 +- 0.014 (0.89) | 0.588 +- 0.014 (0.02) | 0.634 +- 0.015 (0.00) |
| runs-up-and-down z | -3.06 | 0.15 +- 1.05 (0.00) | -1.44 +- 1.1 (0.09) | -3.48 +- 1.0 (0.66) | -5.70 +- 1.1 (1.000) |
| adjacent pairs within 10 | 32 | 21.6 +- 5.1 (0.97) | 40.0 +- 6.2 (0.12) | 59.9 +- 7.2 (0.00) | 78 +- 8 (0.00) |
| distinct numbers | 298 | 298 +- 11 (0.55) | 300 +- 10 (0.45) | 300 +- 10 (0.44) | 299 +- 11 (0.46) |

B3 (DOI key, r = 0.59; the 975-word key gives the same picture):

| statistic | B3 obs | q = 0.2 | q = 0.3 | q = 0.5 | q = 0.6 | q = 0.7 |
|---|---|---|---|---|---|---|
| lag-1 autocorrelation (Pearson) | 0.622 | 0.221 +- 0.051 (1.000) | 0.310 +- 0.055 (1.000) | 0.495 +- 0.049 (1.000) | 0.585 +- 0.046 (0.81) | 0.666 +- 0.037 (0.11) |
| lag-1 autocorrelation (rank) | 0.489 | 0.215 +- 0.044 (1.000) | 0.296 +- 0.043 (1.000) | 0.473 +- 0.041 (0.67) | 0.560 +- 0.039 (0.04) | 0.645 +- 0.033 (0.00) |
| lag-2 autocorrelation (Pearson) | 0.454 | 0.078 +- 0.050 (1.000) | 0.124 +- 0.056 (1.000) | 0.250 +- 0.061 (1.000) | 0.342 +- 0.062 (0.97) | 0.444 +- 0.054 (0.56) |
| lag-2 autocorrelation (rank) | 0.332 | 0.082 +- 0.047 (1.000) | 0.120 +- 0.051 (1.000) | 0.234 +- 0.054 (0.96) | 0.315 +- 0.052 (0.62) | 0.417 +- 0.047 (0.03) |
| mean rising-run length | 2.58 | 2.42 +- 0.07 (0.99) | 2.71 +- 0.09 (0.09) | 3.60 +- 0.18 (0.00) | 4.32 +- 0.23 (0.00) | 5.45 +- 0.34 (0.00) |
| rising runs >= 4 | 49 | 46.1 +- 4.9 (0.74) | 57.4 +- 4.7 (0.04) | 73.7 +- 4.4 (0.00) | 75.5 +- 3.6 (0.00) | 71.3 +- 3.7 (0.00) |
| rising runs >= 6 | 17 | 6.8 +- 2.4 (1.000) | 13.5 +- 3.1 (0.90) | 31.0 +- 3.7 (0.00) | 39.7 +- 3.9 (0.00) | 45.9 +- 3.4 (0.00) |
| longest rising run | 10 | 7.7 +- 1.1 (0.99) | 9.1 +- 1.4 (0.87) | 12.8 +- 2.1 (0.11) | 15.5 +- 2.5 (0.00) | 19.6 +- 3.4 (0.00) |
| share of rising steps | 0.613 | 0.588 +- 0.012 (0.99) | 0.631 +- 0.013 (0.09) | 0.723 +- 0.014 (0.00) | 0.769 +- 0.012 (0.00) | 0.817 +- 0.011 (0.00) |
| runs-up-and-down z | -7.61 | -3.73 +- 1.0 (0.00) | -6.19 +- 1.1 (0.12) | -12.4 +- 1.1 (1.000) | -16.0 +- 1.2 (1.000) | -20.1 +- 1.1 (1.000) |
| adjacent pairs within 10 | 76 | 74 +- 8 (0.61) | 95 +- 10 (0.03) | 138 +- 10 (0.00) | 160 +- 11 (0.00) | 180 +- 11 (0.00) |
| distinct numbers | 263 | 277 +- 11 (0.11) | 282 +- 12 (0.06) | 292 +- 12 (0.01) | 302 +- 13 (0.00) | 310 +- 14 (0.00) |

Reading these results.
- The B2 encoder's serial habit is reproduced by q = 0.1: with one step in ten taken to the nearest forward homophone, every serial statistic of B2 except the rank autocorrelation lies between the 15th and 87th percentile of the simulations; the rank autocorrelation needs q = 0.2, at which the run statistics are already too strong. B2's habit is stronger in rank than in value (0.21 against 0.05), which a scanning rule does not produce (it raises both alike); part of his serial structure is therefore of another kind - stretches in which he re-used his small early homophones and stretches in which he drew fresh, larger ones. The forward-move count above is of the same order (25 excess moves in 761 steps net of the moves the memoryless encoder makes anyway).
- B1 sits between q = 0.1 and 0.2 with either key. At q = 0.1 the run statistics, the runs z and the adjacent pairs fit (percentiles 0.09-0.97) and only the lag-1 Pearson autocorrelation is high (0.252 against 0.144 +- 0.055, 98th percentile); at q = 0.2 the autocorrelations and the run counts fit (0.16-0.66) while the mean run length and the rising share are at the 2nd percentile and the model has too many adjacent pairs within 10 (60 +- 7 against 32). A habit of B2's kind, a little stronger, covers B1's serial structure; nothing in it is beyond what a letter-by-letter encoder produces. (B1's lack of drift cannot be scored against these models, for the reason given in 5.3.)
- B3 has no q. Its autocorrelations need q = 0.6-0.7 (Pearson lag-1 0.585 +- 0.046 at q = 0.6, lag-2 0.444 +- 0.054 at q = 0.7), but at those values the encoder produces rising runs that B3 does not have: mean run length 4.3-5.5 against 2.58, 40-46 runs of six or more against 17, 77-82 % rising steps against 61 %, 160-180 adjacent pairs within 10 against 76 - every one at the 0th percentile of 300 simulations. B3's run statistics are matched at q = 0.3 (run length 0.09, runs >= 6 0.90, longest run 0.87, rising share 0.09, runs z 0.12), where the autocorrelations are half B3's (0.31 / 0.30 against 0.62 / 0.49 at lag 1, 0.12 against 0.45 at lag 2; all at the 100th percentile). The two families of statistics pull in opposite directions because forward scanning generates its correlation through saw-tooth rising runs, whereas B3's correlation persists across runs: its lag-2 autocorrelation is 73 % of its lag-1 value (0.45 / 0.62), against 58 % for the q = 0.6 encoder (0.34 / 0.59) and 39 % for a first-order autoregression with the same lag-1 value. B3's numbers wander in slow waves of level rather than climb in runs.
- The autocorrelation function makes the same point (`stats_11_acf.py`; `results/stats_acf.svg`; 2000 shuffles per cipher, 300 simulations per model; a least-squares line of value on position is removed from observed and simulated sequences alike, so that the calibrated encoder's artefactual drift does not enter). B3's detrended Pearson autocorrelation at lags 1-7 is 0.59, 0.41, 0.27, 0.18, 0.14, 0.10, 0.06 (shuffles: central 95 % within +- 0.08 at every lag; the observed value is above all 2000 shuffles up to lag 5 and above 99.4 % at lag 6); its rank autocorrelation 0.63, 0.47, 0.36, 0.30, 0.24, 0.19, 0.11. The q = 0.6 encoder matches lag 1 (0.56 +- 0.04) but decays twice as fast - 0.31, 0.16, 0.09, 0.04, 0.02 at lags 2-6 (Pearson; the observed values are at the 91st-98th percentile at each of these lags) and 0.29, 0.15, 0.08, 0.04, 0.02 (rank; the observed values exceed all 300 simulations at lags 2-6). The q = 0.3 encoder is at 0.27 at lag 1 and near zero from lag 3. Block-level variance does not separate the two: the variance of the means of consecutive blocks of 10 / 20 / 50 numbers is 3.3 / 3.5 / 2.2 times its exchangeable value in B3 after detrending (upper-tail p = 0.0005, 0.0005, 0.007 against shuffles; B1 1.0 / 1.1 / 1.2 and B2 1.0 / 1.0 / 0.8, both unremarkable) and the q = 0.6 encoder gives 3.0 +- 0.5 / 3.2 +- 0.8 / 3.2 +- 1.4, because long rising runs move block means too. What separates them is the combination: to get B3's lag-1 correlation and block variance the scanner needs runs twice as long as B3's, and even then its correlation is down to 0.09 by lag 4, where B3's is still 0.18 (Pearson) to 0.30 (rank). B2 shows no autocorrelation beyond lag 1 (rank) or lag 2 (0.07), and its q = 0.1 model fits every lag (percentiles 0.06-0.95). B1 has, besides its lag-1 value (0.25), a negative autocorrelation at lags 3 and 4 (-0.10 and -0.12, at the 0.8th and 0.0th percentile of shuffles; rank lag 7: -0.14, 0.0th percentile) that neither the shuffles nor the scanning models produce (models: -0.01 +- 0.05 at these lags; observed at the 0th-2nd percentile). With sixty lag statistics examined these three values are suggestive rather than established (the lag-4 value alone would be p = 0.005 two-sided), but they point to a sequence that oscillates about a local level - a few high numbers followed by low ones - which is a documented habit of people producing numbers they intend to look random (Wagenaar 1972), not of a key look-up. It is noted as a lead, not a finding.
- Two caveats. Distinct count and median are not used as discriminators here: r could be re-calibrated at each q, and the median shortfall is the model's (section 5.3). And the mixed model is one specific scanning habit (nearest forward homophone from the last number used); a scanner who skipped homophones, or worked forward from a fixed place rather than from the last number, would produce runs of a different length - but any forward-scanning rule makes its correlation out of rising runs, and rising runs are what B3 has too few of for its correlation.

## 6. Letter frequencies under the DOI key (`stats_05_letterfreq.py`)

Each cipher was decoded with (a) the straight count of the pamphlet's Declaration and (b) the pamphlet numbering that decodes cipher 2 ("resolved"), and the resulting letter counts were compared with English (Norvig's letter counts from the Google Books corpus), with the initial-letter distribution of the key, and with the expectation for numbers unrelated to the key. That expectation was made concrete by re-decoding with 4000 locally shuffled keys (initials permuted within blocks of 20 consecutive words, which keeps each number's neighbourhood but destroys the identity of the word): under H0 the decode should look like the key's initial distribution re-weighted by where the cipher's numbers fall, and the chi-square of the observed counts against that expectation, its Monte-Carlo p, the KL divergence to English, and per-letter standardised residuals were computed.

| key | cipher | letters decoded (out of range) | KL to English | KL to key initials | chi2 vs re-weighted key (MC p) | KL to English under H0 | p(closer to English than H0) | letters with z > 3 |
|---|---|---|---|---|---|---|---|---|
| resolved | B1 | 508 (12) | 0.253 | 0.042 | 28.7 (p = 0.19) | 0.275 | 0.27 | none |
| resolved | B3 | 618 (0) | 0.246 | 0.063 | 22.9 (p = 0.70) | 0.273 | 0.20 | q (1 hit, 0.1 expected) |
| resolved | B2 | 762 (0) | 0.039 | 0.408 | 560.7 (p < 0.0003) | 0.289 | < 0.0003 | e, n, u, v, x (t in deficit) |
| straight | B1 | 510 (10) | 0.254 | 0.045 | 14.9 (p = 0.83) | 0.283 | 0.20 | none |
| straight | B3 | 618 (0) | 0.315 | 0.094 | 41.2 (p = 0.058) | 0.283 | 0.84 | none (b: 39 vs 21, z = 2.8; j: 7 vs 2.7, z = 2.6) |
| straight | B2 | 762 (0) | 0.175 | 0.208 | 186.1 (p < 0.0003) | 0.299 | 0.001 | n, u |

Commonest letters of the decodes (pamphlet numbering): B1 t .185, a .142, o .089, s .063, c .059, i .055, b .051, w .043; B3 t .215, a .126, o .086, w .060, e .057, s .053, c .050, r .050; B2 e .135, t .098, n .089, o .083, i .068, d .064, s .059, a .056. The key's initials: t .190, a .131, o .112, h .062, i .054, f .054, s .053, p .045; English: e .125, t .093, a .080, o .076, i .076, n .072, s .065.

B1 and B3 decode to the key's own initial distribution, as they must if their numbers have nothing to do with which word stands at that position; B2 decodes to English. B3 has no letter in real excess under the pamphlet numbering (its `q` is one number, 684 = "quartering", where 0.1 was expected); under the straight count it shows a mild surplus of b and j (p = 0.058 overall, not significant after allowing for the six tests), which does not survive the corrected numbering. The letter test therefore excludes (i) for both ciphers - subject to the caveat that it assumes the same simple scheme as cipher 2 (one number per letter, initial letters); a second layer of transformation cannot be excluded by any letter count.

## 7. Capacity of cipher 3 (`stats_06_capacity.py`)

The letter of 5 January 1822 says paper 3 contains "the names of all my associates" and "opposite to the names of each one ... the names and residences of the relatives and others, to whom they devise their respective portions"; the party was "not less than thirty individuals" and the treasure is to be divided into thirty-one parts, one for Morriss. Cipher 2 shows the encoder's conventions: one number per letter, everything written out, only "nov", "dec" and "st" abbreviated; residences written as "in the county of bedford".

Name lengths were taken from USGenWeb transcriptions of real Virginia census records: Amherst County 1810 heads of household (864 usable names; first name 6.08 letters, surname 6.37, full name 12.45 +- 2.7), Bedford County 1850 (2478 persons; men of 30 and over as associates 11.98; all persons as relatives: first name 5.49, surname 6.20), Botetourt County 1850 index (10 866 persons; 30 and over 12.28). Middle initials were dropped. Residences: the 78 Virginia counties and cities formed by 1822 (Wikipedia list with formation years; mean name length 8.46 letters, shortest "Lee"). Totals for 30 entries were drawn 20 000 times.

| scenario | letters: mean +- sd | 1st / 5th percentile | P(total <= 618) |
|---|---|---|---|
| S0 bare minimum: first name + surname; relative's first name only; county name only | 792 +- 19 | 747 / 760 | 0.0000 |
| S1: S0 plus the relative's surname | 978 +- 21 | 929 / 944 | 0.0000 |
| S2: S1 plus the word "county" | 1158 +- 21 | 1110 / 1124 | 0.0000 |
| S3: S2 plus "virginia" | 1398 +- 21 | 1350 / 1364 | 0.0000 |
| S4 extreme abbreviation: initial + surname for both persons, county name | 691 +- 18 | 651 / 662 | 0.0000 |
| S5: initial + surname; relative's first name only; county name | 640 +- 18 | 599 / 611 | 0.115 |
| S6 floor: associate surname + relative surname + county name, nothing else | 631 +- 17 | 590 / 602 | 0.240 |

618 numbers give 20.6 letters per associate for his name, his heir's name and a residence; a full name alone averages 12.5. Using the other name pools for the associates changes S0 by less than 15 letters (778-792). Unless most first names, all designations and most of the residences were left out - a list that would not let Morriss "easily" find the heirs, as the letter promises - paper 3 cannot contain what the pamphlet says it contains. The 30-entry list built from real names for the simulations of section 5 runs to 1395 letters when written the way cipher 2 writes.

## 8. Classic claims checked against these data

- Hammer (1971, 1979): "the ciphers are not random and probably encode intelligible text". Not substantiable here: his statistics are not in our sources. B1 and B3 are indeed non-random (section 4), but the genuine B2 is non-random in the same directions, and B3's non-randomness is of a kind (rising runs, drift) that intelligible text enciphered letter by letter does not produce.
- Gillogly (1980): confirmed and sharpened (section 4); the anomaly is confined to B1 and to one 20-number stretch.
- Wase (2020): the token-level non-uniformity of last digits is reproduced; its interpretation changes once re-use is allowed for (section 3). His cross-base argument was not re-run.
- Campanelli (2022): the token-level first-digit difference between B2 and B1/B3 is reproduced; at the level of distinct values it is largely a re-use effect (section 3).
- Crossen (1927)/Kruh (1982): the capacity objection is confirmed and quantified (section 7).
- The premise that a genuine homophonic book cipher has no serial correlation is refuted by B2 itself (section 4).

## 9. Methods, seeds and scripts

- `stats_common.py` loaders and helpers; `stats_simlib.py` encoders and sequence statistics.
- `stats_00_key_check.py` numbering variants, decode of B2, alignment (seedless). `stats_01_descriptive.py` tables and histogram. `stats_02_digits.py` digit tests (seed 20260914; 100 000 multinomial draws per goodness-of-fit test, 20 000 draws for the type-randomised null, 20 000 permutations per homogeneity test). `stats_03_sequential.py` (seed 20260914; 20 000 permutations, 5000 for the trend and alphabetic-run tests). `stats_08_serial_extras.py` (seed 20260914; 5000 permutations). `stats_04_homophone_sim.py` (seed 20260914; 1000 simulations per model). `stats_09_calibrated_reuse.py` (seed 20260915; 120 simulations per grid point, 1000 at the calibrated value). `stats_10_scan_forward.py` (seed 20260916; 300 simulations per grid point; 2000 letter-preserving shuffles for the forward-move count). `stats_11_acf.py` (seed 20260917; 2000 shuffles per cipher, 300 simulations per model). `stats_05_letterfreq.py` (seed 20260914; 4000 shuffled keys). `stats_06_capacity.py` (seed 20260914; 20 000 lists). `stats_07_figures.py`. `stats_summary_tables.py`.
- Permutation and Monte-Carlo p-values are (k + 1)/(N + 1); their standard error is sqrt(p(1 - p)/N), i.e. about 0.0015 at p = 0.05 for N = 20 000 and 0.007 for N = 1000. "Percentile" of an observed statistic among simulations is the share of simulations at or below it; a value at 0.000 or 1.000 means none of the 1000 simulations reached it (two-sided p < 0.002). Pearson chi-square p-values are given both asymptotically and by Monte Carlo where expected counts are small.
- Caveats: the encoder model is fitted to a single 762-letter sample; its reuse curve has 22-37 events in the low-d bins. The long keys' initials come from the pamphlet's narrative, a text of the same period and hand; another text would move tau and the letter-availability pattern slightly. Plaintext windows come from the "Beale" letters (12 100 letters), which are the only prose of that hand available; the names-list plaintext gives the same answers. All tests assume the cipher-2 scheme (one number per letter, initial letters); a different scheme would change the letter test (section 6) but not the digit, serial or capacity results. Multiple testing: sections 3 and 4 report about forty p-values; results at the 1-5 % level should be read accordingly, and the conclusions above rest only on results at p < 0.001 or on consistent patterns across tests.

## Sources

- The Beale Papers (Lynchburg, 1885), Wikisource transcription of the Library of Congress copy, pages 1-23: https://en.wikisource.org/wiki/The_Beale_Papers (parsed into `data/primary/`; the letter of 5 January 1822 is on the page reproduced in `data/primary/letter_1822-01-05.txt`, the statement that the party had "not less than thirty individuals" and the division "into thirty-one equal parts" in `letter_1822-01-04.txt`).
- Gillogly, J. J., "The Beale Cipher: A Dissenting Opinion", Cryptologia 4(2), 1980, 116-119, https://doi.org/10.1080/0161-118091854979 ; text used: `sources/gillogly_1980_plain.txt` (Peschel's HTML copy via the Wayback Machine, https://web.archive.org/web/20060126085231/http://members.fortunecity.com/jpeschel/gillog3.htm), which contains his Table I (DOI initials), Table II (B1) and Table III (decode).
- Hammer, C., "Signature simulation and certain cryptographic codes", Communications of the ACM 14(1), 1971, 3-14, https://doi.org/10.1145/362452.362461 (abstract only); Hammer, C., "How did TJB encode B2?", Cryptologia 3(1), 1979, 9-15 (cited from Gillogly 1980 and the Cryptologia table of contents, https://ftp.math.utah.edu/pub/tex/bib/toc/cryptologia.html; not consulted directly).
- Wase, V., "The Role of Base 10 in the Beale Papers", Proceedings of HistoCrypt 2020, Linkoping Electronic Conference Proceedings 171, 153-157, https://ep.liu.se/ecp/171/019/ecp2020_171_019.pdf (p-values 0.4 %, 0.02 %, 0.4 % for last digits; 0.043 %, 3.1 %, 0.004 % for Benford; his quotation of Hammer 1970's abstract; the Mateer 2013 list of numbering peculiarities; the Crossen/Kruh capacity objection).
- Campanelli, L., "A statistical cryptanalysis of the Beale ciphers", Cryptologia 47(5), 2023 (online 2022), https://doi.org/10.1080/01611194.2022.2116614 (abstract: B2's first digits follow an epsilon-Benford law with epsilon about 0.15, B1 and B3 deviate significantly).
- Kruh, L., "A Basic Probe of the Beale Cipher as a Bamboozlement", Cryptologia 6(4), 1982, 378-382, and "The Beale Cipher as a Bamboozlement - Part II", Cryptologia 12(4), 1988, 241-246, https://doi.org/10.1080/0161-118891863007 (the Crossen 1927 capacity argument is cited through Wase 2020).
- Mateer, T. D., "Cryptanalysis of Beale Cipher Number Two", Cryptologia 37(3), 2013, 215-232 (cited through Wase 2020 for the numbering peculiarities; our own list in section 1 is independent).
- Wagenaar, W. A., "Generation of random sequences by human subjects: a critical survey of literature", Psychological Bulletin 77(1), 1972, 65-72, https://doi.org/10.1037/h0032060 (people avoid repetitions when trying to be random).
- Norvig, P., "English Letter Frequency Counts: Mayzner Revisited", http://norvig.com/mayzner.html (letter frequencies from the Google Books corpus, used as "English").
- Phipson, B. and Smyth, G. K., "Permutation p-values should never be zero", Statistical Applications in Genetics and Molecular Biology 9(1), 2010, article 39 (the (k + 1)/(N + 1) convention).
- USGenWeb Census Project transcriptions: Amherst County, Virginia, 1810 (transcribed by V. Moore), https://www.us-census.org/pub/usgenweb/census/va/amherst/1810/index.txt ; Bedford County, Virginia, 1850 (transcribed by E. H. May), https://www.us-census.org/pub/usgenweb/census/va/bedford/1850/ (files pg0140a.txt, pg0151a.txt, pg0162b.txt); Botetourt County, Virginia, 1850 name index (transcribed by S. Carneiro), https://www.us-census.org/pub/usgenweb/census/va/botetourt/1850/ (indx-*.txt). Used for name-length statistics only; a 30-name sample is kept in `results/stats_capacity_names_plaintext.txt` for the simulations (the transcribers permit use in individual research).
- List of cities and counties in Virginia with formation years, Wikipedia, https://en.wikipedia.org/wiki/List_of_cities_and_counties_in_Virginia (wikitext fetched 2026-09-14; the 78 entries formed by 1822 were used; counties now in West Virginia are not in this list).
- Wikipedia, "Beale ciphers", https://en.wikipedia.org/wiki/Beale_ciphers (pointers only; quotes the corrected alphabetic string `abcdefghiijklmmnohpp`).
- Key reconstruction in `notes/cryptanalysis_b2_gillogly.md`: `results/b2_decode_summary.json`, `results/b2_encoder_behaviour.md`, `results/transcription_check.json` (agreement on the pamphlet numbering, the three special numbers and the 762-number count).
