An independent check of The Bixby Letter
The Bixby letter is the condolence Abraham Lincoln signed on 21 November 1864 to Lydia Bixby of Boston, who was reported to have lost five sons in the war. No manuscript survives, and since the 1920s some have held that his secretary John Hay wrote it. In 2019 a team of forensic linguists led by Jack Grieve tested the question with a statistical method. The method, n-gram tracing, asks which of two writers' known works contains more of the letter's short strings of letters and words. By that measure the letter matched Hay's known writing more closely than Lincoln's.
The Bixby Letter, a report on this site, reran that study and added a test the study's authors had not run: it tested the method on prose Lincoln is known to have written in the letter's own years and in its register, the elevated prose of condolence and ceremony rather than the business notes that fill most of his papers. Under the study's design, 30 of 63 such pieces, the Gettysburg Address of 1863 among them, were attributed to Hay. A result like that should be checked by an independent implementation before it is relied on.
This page is that check: a new implementation of n-gram tracing, written from the published paper without sight of the code behind The Bixby Letter (called the first implementation below), run on the same texts and then on texts chosen afresh. It asks whether the tests reproduce. It does not ask who wrote the letter.
The tests reproduce, with one difference of degree. The attribution to Hay reproduces at every n-gram length the study used, and so does the method's accuracy on the two men's ordinary writings. The misattribution of Lincoln's war-time prose also reproduces, at 25 to 30 percent of pieces rather than 48. The first implementation has since traced that difference to 83,000 words of Lincoln's 1858 campaign speeches which it kept and this one left out. A run here with those speeches restored gives 27 of 63. None of the results is evidence that Lincoln wrote the letter.
It does. The letter goes to Hay at all 17 n-gram lengths the study used, by margins (Hay's share of the letter's n-grams minus Lincoln's) of 0.002 at character length 3 to 0.145 at word length 2. Word 5 goes to Lincoln on the phrase "may be found in the", as in the study.
2019 study: all 17 to Hay. First implementation: all 17 to Hay.
It does in outline. Attributing each of 1,397 known texts with itself left out of its author's pool, the character vote is right for 1,380 (98.8 percent), where the study reports none wrong. The scores rise and fall with n-gram length as the study's do: useless at 1 and 2 characters, highest at 7 and 8 (F1 0.96 to 0.99 at character lengths 4 to 10, where 1.00 would be perfect), lower at longer lengths.
2019 study: 1,662 of 1,662 right by the character vote.
Partly. The misattribution appears, but in 19 of its 63 Lincoln pieces (30 percent) under the study's sample rule with 50 random sequences, and in 15 plus one tie (25 percent) under the first implementation's own rule, rather than in 30 of 63. Of its 57 Hay pieces, 2 go to Lincoln.
First implementation: 30 of 63 (48 percent); Hay 1 of 57 wrong.
It does. 27 of 97 pieces (28 percent) from 38 documents go to Hay under the study's design. With pools drawn from the same years and each parent document left out, 1 of 97 is wrong, and the letter splits 5 to 2 for Lincoln on the character vote and 9 to 8 for Hay on all 17, which is evidence for neither man.
There is no earlier figure for these pieces; they are new to this report.
The study's design misattributes Lincoln's elevated war-time prose at a rate well above its 2.6 percent error on his ordinary pre-1860 texts. The rate depends on what is in the Lincoln pool: 25, 28, 30 and 43 percent in four runs here (the last with the first implementation's 1858 speeches restored), 48 percent in the first implementation. The higher figures alone might suggest that the 2019 attribution fails validation. The numbers support only this: the attribution reproduces, and its validation does not transfer to texts of the letter's period and register, where the design misattributes between a quarter and a half of Lincoln's known pieces. The 2019 result should therefore be given little weight in either direction. The evidence does not reverse it.
The following were read in full: the study, Jack Grieve, Emily Chiang, Isobelle Clarke, Hannah Gideon, Annina Heini, Andrea Nini and Emily Waibel, "Attributing the Bixby Letter using n-gram tracing", Digital Scholarship in the Humanities 34(3), 2019, pp. 493 to 512, in the accepted manuscript (the version of record at academic.oup.com refused the connection); the authors' 2016 and 2017 slide decks and the 2017 conference abstract; and The Bixby Letter with the notes that accompany its files. From those files only data was taken: the collections of Lincoln's and Hay's writings it assembled (lincoln_basler.jsonl, hay.jsonl, lapsley.jsonl), its lists of test pieces (register.json, specials.json), and two cleaned essay texts, used only to check this report's own extractions of the same essays.
The following were not opened, so that the two implementations stay independent: the first implementation's Python files, its results and its logs. The method here was written from the study's section 3 and sections 5 and 6, and the war-time test from the first implementation's description of its design (the 63 and 57 pieces, taken as given).
An n-gram is a run of n consecutive characters or words. "Solemn pride" contains the character 4-grams "sole", "olem", "lemn", "emn " and so on, and the word 2-gram "solemn pride". N-gram tracing lists every distinct n-gram of one length in the questioned text, draws for each candidate a random sample from a pool of his known writing (the pools are described in the next section), and scores him by the share of the questioned text's n-grams that occur anywhere in his sample. The candidate with the larger share is taken to be the author. The idea is that a writer's habitual strings, down to spelling and punctuation, turn up in his other work more often than another writer's do.
The details follow the study. Each candidate's sample is a random sequence of his texts, taken until its words reach the size of the smaller candidate's whole sample, so that the two are compared at equal length. Scores are averaged over random sequences: 50 for the letter and the war-time pieces, 10 for the validation. Word n-grams ignore case and punctuation; character n-grams are case-insensitive and keep punctuation and spaces; no n-gram crosses a sentence boundary. Two majority rules decide: at least 4 of the 7 analyses at character lengths 4 to 10, or at least 2 of the 3 at word lengths 1 to 3. A tie counts as wrong for both authors. Below, "the character vote" means the first of these rules, and a text "goes to" the candidate it is attributed to.
Choices the study leaves open were decided here and are documented with the code: where sentences and words begin and end, how a text left out of its own author's sample is replaced (by the next text in the random order, so the sample keeps its size), and an option that scales the sample size (1.0 is the study's rule; 0.95, samples of 95 percent of the smaller pool, reproduces the first implementation's rule).
| Pool | Texts | Words | Built from |
|---|---|---|---|
| Lincoln before 18 May 1860 (the study's cut-off, the day of his nomination for President), this implementation | 505 | 192,404 | Volume I of Roy P. Basler's Collected Works of Abraham Lincoln (1953), the standard edition (363 of its 410 items; dropped: documents only signed by Lincoln and items in another hand), plus Arthur Brooks Lapsley's older edition, The Writings of Abraham Lincoln (1905), for 1849 to 17 May 1860 (142 items; the two volumes of speeches from the 1858 Senate campaign against Stephen A. Douglas left out, see Test 3) |
| Hay, everything, this implementation | 892 | 499,867 | Every item of at least 5 words in the Hay collection assembled for The Bixby Letter: Letters of John Hay and Extracts from Diary (1908), the letters printed in William Roscoe Thayer's 1915 biography, Addresses of John Hay (1906), his 1902 memorial address on President McKinley, his book on Spain Castilian Days (1871), his novel The Bread-Winners (1883), and his verse |
| The 2019 study | 1,085 / 577 | 400,747 / 261,126 | Basler's volumes I to IV, from the University of Michigan's online edition (which refused this implementation's connection); Thayer's biography and four Project Gutenberg texts |
| The first implementation | 580 / 891 | 243,000 / 495,000 | Basler's volumes I and IV plus Lapsley, the 1858 campaign speeches included; the same Hay collection |
The Lincoln side here is half the study's and a fifth smaller than the first implementation's, and the Hay side is twice the study's. Tests 1 and 2 therefore reproduce the study's result with a reconstruction of its design; they do not use its data. data/pools_manifest.csv lists every candidate document, whether it was kept or dropped, and the reason.
results/test1_letter_by_type.csv).The first implementation reports the same 17 to Hay, character 3 by 0.00004 and the rest by 0.01 to 0.14. The margins here are 0.04 to 0.10 at character lengths 5 to 12 and 0.145 at word 2, against its 0.04 to 0.10 and 0.14. The same 17 go to Hay under the first implementation's 0.95 sample rule with 10 sequences.
All 1,397 pool texts, each attributed with itself removed from its author's pool (the study's leave-one-out validation), 10 random sequences, 25 n-gram lengths. The score is F1, which combines how many of an author's texts are recovered with how many of the texts assigned to him are his; 1.00 is perfect. The table sets this implementation's scores beside the study's Tables 3 and 4.
| Length | Hay F1 here | Hay F1 study | Lincoln F1 here | Lincoln F1 study | Accuracy here | Accuracy study |
|---|---|---|---|---|---|---|
| word 1 | 0.97 | 0.93 | 0.95 | 0.95 | 0.96 | 0.94 |
| word 2 | 0.98 | 0.94 | 0.97 | 0.97 | 0.98 | 0.96 |
| word 3 | 0.97 | 0.91 | 0.94 | 0.95 | 0.95 | 0.93 |
| word 4 | 0.92 | 0.80 | 0.87 | 0.92 | 0.89 | 0.85 |
| word 5 | 0.85 | 0.55 | 0.77 | 0.85 | 0.76 | 0.68 |
| char 1 | 0.17 | 0.59 | 0.10 | 0.21 | 0.08 | 0.23 |
| char 2 | 0.66 | 0.74 | 0.73 | 0.70 | 0.58 | 0.58 |
| char 3 | 0.95 | 0.89 | 0.94 | 0.88 | 0.93 | 0.85 |
| char 4 | 0.98 | 0.94 | 0.96 | 0.96 | 0.97 | 0.95 |
| char 5 | 0.98 | 0.95 | 0.96 | 0.97 | 0.97 | 0.96 |
| char 6 | 0.98 | 0.96 | 0.97 | 0.97 | 0.98 | 0.97 |
| char 7 | 0.99 | 0.96 | 0.98 | 0.98 | 0.98 | 0.98 |
| char 8 | 0.99 | 0.96 | 0.98 | 0.98 | 0.99 | 0.98 |
| char 9 | 0.99 | 0.96 | 0.97 | 0.98 | 0.98 | 0.97 |
| char 10 | 0.98 | 0.95 | 0.97 | 0.97 | 0.97 | 0.97 |
| char 12 | 0.97 | 0.93 | 0.95 | 0.96 | 0.96 | 0.96 |
| char 16 | 0.94 | 0.86 | 0.90 | 0.94 | 0.92 | 0.91 |
| char 20 | 0.89 | 0.71 | 0.84 | 0.90 | 0.85 | 0.80 |
The character vote at lengths 4 to 10 is wrong for 17 texts of 1,397 (F1 0.99 for Hay, 0.98 for Lincoln); the study reports all 1,662 of its texts right. The word vote at lengths 1 to 3 is wrong for 25. The 13 Lincoln texts the character vote misses are an 1830 appraisal signed with others, a reporter's third-person account of an 1836 speech, a 70-word note, the 1842 eulogy on Benjamin Ferguson, the 1842 Temperance Address, two letters about poetry and the verse "The Bear Hunt", a 617-word letter to Mary Todd Lincoln, a letter to John D. Johnston, the reported Rock Island Bridge argument, a 109-word letter, and the verse "To Linnie". The 4 Hay texts are a 27-word note and three official letters to senators. The study's authors removed doubtful documents by hand, which would take out several of these. Per-text margins are in results/test2_loo_margins.csv.
The first implementation cut 63 pieces of about 150 words from 30 documents Lincoln wrote in 1861 to 1865 in the letter's register (condolence, ceremony, elevated public prose), and 57 pieces from Hay's prose of the same kind: his two magazine essays of 1861 on officers killed early in the war, "Ellsworth" (Atlantic Monthly, July) and "Colonel Baker" (Harper's, December), and the one condolence letter of his own from the war years that survives in print, 55 words from 1864. Those 120 pieces were taken as given and attributed under the study's design. No parent document is in either pool.
| Run | Lincoln pieces to Hay, character vote | Lincoln pieces to Hay, all 17 | Hay pieces to Lincoln, character vote |
|---|---|---|---|
| This implementation, the study's sample rule, 50 sequences | 19 of 63 (30%) | 7 of 63 (11%) | 2 of 57 |
| This implementation, the first implementation's 0.95 rule, 10 sequences | 15 of 63 plus 1 tie (25%) | 7 of 63 | 2 of 57 |
| The first implementation, its own code | 30 of 63 (48%) | not reported | 1 of 57 |
The method fails at the short character lengths. Accuracy on the Lincoln pieces at character lengths 4, 5 and 6 is 0.29, 0.25 and 0.44, rising to 0.67 to 0.89 at lengths 7 to 10 and 0.87 to 0.90 at 11 to 16. The first implementation reports 21 to 38 percent at lengths 4 to 6 and the same recovery at long lengths. Both pieces of the Gettysburg Address and both pieces of the Ellsworth letter (Lincoln's 1861 condolence to the parents of Colonel Elmer Ellsworth) go to Hay in this implementation, as they did in the first. The McCullough letter (his 1862 condolence to Fanny McCullough) goes to Lincoln in this implementation, 4 votes to 3, and to Hay in the first. Every piece with its 25 margins is in results/test3_pieces.csv.
The Bixby Letter, section 4, ran each implementation's code on the other's pool and traces the gap to the Lincoln pool; the code is not the cause. The first implementation keeps 19 items from Lapsley's third and fourth volumes, 83,000 words of Lincoln's speeches in the 1858 Senate campaign and his side of the seven debates with Stephen A. Douglas, which this implementation had left out because Lapsley prints a letter of Douglas's and his questions among them. By that report's figures, its code on this implementation's pool gives 19 of 63, and this implementation's code on its pool gives 26. The mechanism it describes is the sample rule: adding the speeches enlarges both candidates' samples by the same 84,000 words, and the extra words of Hay's mixed prose cover more of a condolence letter's strings than the extra words of Lincoln's stump oratory do. Restoring those items to the Lincoln pool here (521 texts, 275,806 words) and rerunning with this implementation's own code sends 27 of the 63 pieces to Hay under the study's rule with 50 sequences, so the pool difference accounts for the gap (results/summary_debates.txt). The letter itself still goes to Hay under that pool, by the character vote 7 to 0 and at 16 of the 17 lengths with one tie.
Lincoln: 38 documents of 1861 to 1865 chosen for register (condolence, ceremony, personal and elevated public prose), cut at sentence boundaries into 97 pieces of about 145 words, at most five per document. Nineteen of the 38 are also among the first implementation's 30 documents, because Lincoln left few texts in this register; 19 are not (among them the speech at Independence Hall in 1861, the close of the First Inaugural, the letter to the workingmen of Manchester in 1863, the public letters to Erastus Corning and James C. Conkling of 1863, the speech at the Philadelphia Sanitary Fair in 1864, the remarks to the 166th Ohio Regiment, the reply to the Baltimore committee that presented a Bible, and the last public address of April 1865). Hay: the two 1861 essays cut afresh from their sources, "Ellsworth" from Project Gutenberg text 11154 and "Colonel Baker" from the Internet Archive's scan of Harper's volume 24, 76 pieces; his 55-word condolence of 1864; and, as prose of the same years but not the same register, 185 pieces from his 110 letters and diary entries of 1861 to 1865 of at least 100 words. Every document and its source is in data/test4_sources.csv.
| Design | Lincoln pieces wrong, character vote | Hay essay pieces wrong | Hay letters and diary wrong |
|---|---|---|---|
| A: the study's pools (Lincoln before 1860, all of Hay) | 27 of 97 (28%); 11 by all 17 | 4 of 76 (5%) | 1 of 185 (1%) |
| B: pools of 1861 to 1865 built from these documents, parent left out (Lincoln 38 documents, 14,168 words; Hay 113 documents, 38,506 words) | 1 of 97 (1%) | 7 of 76 (9%) | 8 of 185 (4%) |
| B2: as B, Hay side limited to the essays and the condolence (11,166 words) | 2 of 97 (2%) | 7 of 76 (9%) | not tested |
Under design A the Lincoln documents with wrong pieces include the farewell at Springfield in 1861, the Ellsworth letter (both pieces), the Manchester letter (3 of 4), the Philadelphia fair speech (3 of 3), the 166th Ohio remarks (2 of 2), the Baltimore Sanitary Fair speech of 1864, the Second Inaugural of 1865 and the Conkling letter (2 of 5 each), and the Gettysburg Address (1 of 2). Under designs B and B2 the letter splits (B: 5 to 2 for Lincoln at character lengths 4 to 10, 9 to 8 for Hay over all 17; B2: 4 to 3 and 10 to 7), Hay's own condolence goes to Lincoln, and the four letters Hay is known to have drafted for Lincoln's signature (to George H. Boker in 1863, to Charles Butler in 1864, and to John F. Driggs and William Lloyd Garrison in 1865) go to Lincoln 7 to 0, 7 to 0, 6 to 1 and 7 to 0. Under design A those four go Hay, Hay, Lincoln, Lincoln, and the condolence goes to Hay 7 to 0.
Test 1 shows that the study's attribution is a stable property of its design. The same result is obtained with a fresh implementation, with a Lincoln pool half the study's size drawn partly from a different edition, and with a Hay corpus assembled from different books. It does not show that the design gives correct attributions for texts of the letter's kind.
Test 2 shows that the method separates the two men's ordinary writings about as well as the study says. Its residual errors fall on Lincoln's verse, eulogy, temperance oratory and private letters, which is the direction of the register effect seen in the next two tests.
Tests 3 and 4 show that the design's accuracy on Lincoln's elevated prose of 1861 to 1865 is far below its accuracy on his ordinary texts, and that the loss is confined to short character n-grams. They do not fix the size of the effect: 25 to 30 percent here on the first implementation's pieces, 28 percent on pieces chosen independently, 48 percent in the first implementation's run, and the composition of the Lincoln pool moves it. The method stays right for at least half of these pieces and for 89 percent of the votes over all 17 lengths, so "fails validation" would overstate what the numbers show.
Design B shows that pools drawn from the same years separate the two men in this register too, on very small pools. It does not turn the letter's split under those pools into evidence for Lincoln, because the method under those pools also calls Hay's known writing for Lincoln's signature "Lincoln".
The study's authors had noted the register problem themselves. Their 2016 slides say that they had "not looked at detail at the effect of ... register variation" (slide 15) and "also need to test sensitivity to register variation" (slide 54). The first implementation was the first to test it, and this page finds the same effect at a lower rate.
The study's own corpora could not be obtained, so the Lincoln pool here is 192,000 words against the study's 401,000, and 106,000 of those words are Lapsley's printed text rather than Basler's text edited from the manuscripts. The Hay corpus differs from the study's in source and size. Only the accepted manuscript of the study was read; the published version could not be downloaded. The rules for splitting text into words and sentences are this implementation's own, and the letter counts 138 words here against the study's 139.
The first implementation's 63 and 57 pieces were used as given. Their selection, cutting and OCR were not checked beyond the two essays, where this report's own extractions matched to within nine words in 6,975. Test 4's Lincoln texts overlap the first implementation's by 19 documents, and a fully disjoint set does not exist for this register. Design B's pools are small, so its long n-gram overlaps are near zero and its votes rest on character lengths 4 to 10.
Everything here can be rerun. The folder independent-rerun, kept with the files of The Bixby Letter, holds README.md (the same results with every table), code/ngt.py (the implementation) with the scripts that build the pools and run the four tests, data/ (the pools, the manifest of every candidate document, the two lists of test pieces, and Test 4's documents with their sources), results/ (summaries, per-piece CSV tables, raw JSON) and notes/paper_notes.md (reading notes on the study and the two slide decks, with page references). The random-number seeds are fixed, so a rerun repeats these results exactly. Python 3.11 with numpy:
cd code
python3 build_pools.py
python3 run_tests.py letter --nseq 50 --workers 4
python3 run_tests.py loo --workers 3
python3 run_tests.py summary
python3 test4_build.py && python3 test4_run.py all
python3 export_tables.py
This page is an independent check of one part of The Bixby Letter: whether that report's stylometric tests reproduce.