Refresh site to see updates

Stylometry

Sytometry sty·lom·e·try /stīˈlämətrē/
the statistical analysis of variations in literary style between one writer or genre and another.

FAQ

  • What do you use to perform Stylometric analysis? Stylo is a package developed and maintained by Computational Stylistics Group for R Studio.
  • Can I use Stylo with Python? PyStyl and faststylometry are two Python libraries inspired by Stylo.
  • Where are the peer reviewed studies using Stylo? The Publications link on the Computational Stylistics Group website contains links to papers using Stylometry, Note: there are peer reviewed links on that page, but not all of the links are to peer reviewed papers.
  • Does it utilize machine-learning? Several machine-learning methods are available for supervised runs of Stylo with the classify() function. Only unsupervised runs are utilized in these examples.
  • What's the difference between supervised and unsupervised runs? Supervized runs used labeled data. Meaning you know the author beforehand and you're training their known works to identify if that person has written other works not attributed to them. Unsupervised runs compare the works provided without any influence from a specific writer. This allows the works to be compared to each other independently.
  • Can it be used by AI/LLMs? Yes. It was used in one of the linked publications.
  • Did Stylometry really help convict the Unabomber? Yes but they did it by hand.
  • Why don't people do it by hand? Using a computer for the analysis eliminates virtually all human error (with the exception of bad/malformed data), bias, and is much faster than human analysis.
  • Does the file name influence the results? No. Though, to control for a perceived bias, in the first example the file names have been obfuscated.
  • Do you need to know the language being analyzed? You only need to be able to identify what language is in the file, so you can choose it appropriately in Stylo.
  • Do you need to have writings from throughout the person's life to accurately identify them? No; the individual texts are compared to one another in unsupervised runs.
  • What is the recommended amount of most frequent words (MFW)? It's recommended that at least 100-200 MFWs are used to get a strong authorial signal. However, there are instances where authorship can be determined with fewer words, and that can be observed by outputting the results in increments. Increasing engrams can also help identify someone through repeated phrases. This is demonstrated below.
  • Where did you get the source data? I used two primary sources. For the New Testament I used the Society of Biblical Literature Greek New Testament (SBLGNT). The data is available in XML and can easily be split by chapter. Other Greek texts are available at Perseus at Tufts University. The text of Evangelion is from Github by Mark Bilby.
  • How accurate is Stylometry? It has been shown to identify Greek genres with >97% accuracy. A new methodology has been demonstrated to have an accuracy of around 84.2%-89.6% for texts containing 50 words. The paper shows accuracy rates as high as 92% with 200-word passages. Eder (from Computational Stylistics) has found that 2000 words may be sufficient, but he has also said some authors can be determined with a 100 word sample. The paper he referenced can be accessed here. Take note: The samples he ran are in English, which is considered weakly inflective compared to the Greek and Latin demonstrated on this page.
  • Does Stylometry prove Jesus didn't exist? Nobody is making that claim. However, it can be inferred indirectly. As it says in the definition above, it can identify authorship, including if someone was not the author. By identifying authorship and knowing when that person lived it logically follows that it was written during their lifetime, which in turn can also identify it wasn't written during someone else's lifetime if that person wasn't alive during the identified authors lifetime. It also assists in pinpointing interpolated texts and resolving the chronological debate of Evangelion predating Luke. Which is additional evidence for the 2ND century authorship of the New Testament.
  • Are you a data scientist? I'm not a data scientist, but if the science is sound it is repeatable. I am a senior software engineer by trade but I have also had previous work peer reviewed by a data scientist. A process that took 8 months to defend and was determined to be error free on a dataset of several million records that was updated quarterly.
  • Why not use Bayesian statistics? A design feature of Bayesian statistics is the use of non-zero probabilities in order for future updates; if something is impossible or completely rules out the likelihood or expected value, then it can't be used. This feature can be useful in some cases, but informative non-zero priors are subjective and can introduce a bias because something must be possible/probabile. Example: Bayesian statistics would require the possibility of being able to swallow my own head whole to return a non-zero result. Bayesian statics should also disclose their hypothesis/prior, evidence/likelihood, expected probability, and how those values were determined.
  • Didn't Dustin White debunk the how Christ Before Jesus used Stylo? He tried to, but there were many flaws in his attempt.

Stylometry Examples


Influence of time

This shows the passage of time influencing writing style. There is an indisputable separation between all of the authors. Suetonius lived and wrote during Tacitus' lifetime (early 2ND century), and they're closer to Lucan (1ST century) than Tertullian (2ND-3RD).
This also shows us the passage in Annales 15.44 wasn't written by Tacitus.
 
Here is the file key.
A dendrogram with tokenized filenames.

Control Test

An example of a control test using a shuffled copy of a file. The shuffled text is on the left side and the original text is on the right. The preamble is not verbose enough to influence the results.
The difference in character counts is due to the shuffler using whitespace to split words and then joining by a single space. It doesn't influence the results, as you can see the distance is 0. They are grouped together at the bottom of the chart.
A dendrogram demonstrating the match between an original file and a copied file with identical but shuffled text.
The shuffled file next to the original.

Testimonium Flavianum

Here is the Testimonium Flavianum compared to the works of Josephus and Eusebius. Antiquities of the Jews Book 18 used in the analysis contains the entirety of the Testimonium Flavianum. 3 engram/word phrases show an even closer connection to Eusebius. This demonstrates that even when the full text of a quote remains in the original, it's not enough to influence the result. The only exception would be if the quote makes up the majority of the text.
There is criticism that text can be faked if someone uses multiple word engrams to impersonate someone else. That could be true if someone was writing later and claiming to be Eusebius. In this case it turns out to be a smoking gun; Josephus never uses those phrases, but Eusebius clearly does. Obviously, it would be impossible for Josephus to copy phrases from someone in the future.
Eusebius writes in Praeparatio Evangelica Book XII (12) Section XXXI (31) that he believes in lying for the benefit of those who "require" it: "That it will be necessary sometimes to use falsehood as a remedy for the benefit of those who require such a mode of treatment"
A 5000 MFW single engram file showing the relation of the Testimonium Flavianum to Eusebius.
A 2000 MFW 3 engram file showing the relation of the Testimonium Flavianum to Eusebius.
Proof that the Testimonium Flavianum is included in the original file.
The image below is a 3 engram dendrogram at only 100 MFWs.
As you can see, the results are similar to the 3 engram run at 2000 words.
The sample sizes are still more than what's recommended.
Total words: 413971
Average words: 13353.90322580645
Total unique words: 56806
Average unique words: 1832.4516129032259

* Note: Unlike the run above, this run does not include the TF in book 18 or the "so called Christ" in book 20.
A 100 MFW three engram file showing the relation of the Testimonium Flavianum to Eusebius.
Eusebius of Caesarea Historia ecclesiastica:
  • book 1: 9910
  • book 2: 9180
  • book 3: 12001
  • book 4: 10165
  • book 5: 12957
  • book 6: 12930
  • book 7: 11441
  • book 8: 7745
  • book 9: 6036
  • book 10: 9548
Flavius Josephus Antiquitates Judaicae
  • book 1: 15275
  • book 2: 14941
  • book 3: 14061
  • book 4: 14117
  • book 5: 14987
  • book 6: 18506
  • book 7: 18917
  • book 8: 20259
  • book 9: 13538
  • book 10: 12714
  • book 11: 13866
  • book 12: 17159
  • book 13: 17620
  • book 14: 19626
  • book 15: 16486
  • book 16: 15162
  • book 17: 15268
  • book 18: 15968
  • book 19: 13660
  • book 20: 9839
  • Testimonium Flavianum: 89
This is the count of how many times a word is used and how many words are used that many times.
The first row represents 33,974 words used only once, while the last row is one word used 23,713 times.

Gospel of Luke

This is a consensus tree performed on Luke. The consensus tree requires at least three incremental passes on the data. For the consensus tree below I used four passes incremented by 500 words each time, hence the 500-2000 MFW label. The "consensus strength" value has been increased from the default setting of 0.5 to 0.9, a strength below 0.5 will return an error. This shows us that 1-3 and 23-24 are not written by the same author of the majority of the text. This is apparent on a dendrogram with as little as 100 MFWs.
A 500-2000 MFWs consensus tree. Showing two authors were involved in the making of Luke.
A dendrogram showing the same separation at 100 MFWs.

Luke compared to Acts

The scholarly consensus is that Acts and Luke were written by the same person. That's partially correct; only the beginning and end of Luke were written by the author of Acts. The consensus strength of the tree was set to 0.99 in this case.
A 500-2000 MFWs consensus tree. Showing the author of Acts is responsible for the beginning and ending of Luke.
A dendrogram showing the same separation at 100 MFWs.
Word count of Acts and Luke and what chapters were combined together.

Who wrote most of Luke?

Many scholars believed that Evangelion was written after Luke and Marcion removed details from it. However, the data tells us Marcion wrote Evangelion first.
Luke was utilized by the author of Acts who padded it and added extra details to some of the verses.
If Luke and Acts were written by the same author and Evangelion stole from them, the bulk of Luke wouldn't have the clear separation from Acts and Evangelion would match with Acts.
We also see a distinction from the first 12 chapters of Acts and the last 11. Which might be from a second author or due to the shift in narration as it changes from third person to first person.
This highlights Marcion priority and Marcion is from the 2ND century.
Marcion's canon was assembled in the 2ND century and contained the Gospel of Marcion (Evangelion) and Apostolikon which consisted of ten letters of Paul. While the Modern New Testament canon wasn't established until the 4TH century.
A 500-2000 MFWs consensus tree. Revealing the true author of Luke is Marcion.
A dendrogram of 100 MFWs returns the same results. Marcion is the author of the bulk of Luke.
Luke 1:3 states "... since I myself have carefully investigated everything from the beginning/start, I too decided to write an orderly account for you, most excellent Theophilus."
Since it contains almost the entire contents of Evangelion, that verse exposes their plagiarism. That admission in conjunction with Stylometry reveals that Evangelion is the original text in their "orderly" account.
Luke was compiled after Marcion while a person named Theophilus held a significant position and before Irenaeus of Lyon named the Gospels around 180CE-185CE, meaning the Theophilus addressed in Luke 1:3 and Acts 1:1 must not be Theophilus ben Ananus a Jewish High Priest from 37CE-41CE.
Furthermore, Theophilus ben Ananus doesn't align with the timeline in Acts; Paul visits Colossae/Phrygia after 41CE. Luke would have had to write Acts after the events in Colossians, Philemon, and Galatians. He would have been addressing Theophilus several years or decades after his priesthood.
The best candidates are Theophilus of Antioch who was the Pope/Bishop of Antioch from 169CE to 183/185CE, or Theophilus, bishop of Caesarea (d. 196CE). Writing to a Theophilus who held an exalted position in the Christian church is more likely than writing to a Jewish High Priest.
This narrows down the potential authors to late 2ND century Christians:
  • Melito of Sardis (100CE-180CE) was a Christian convert with Jewish and Hellenistic roots. Most of his works have been lost but he was quoted by Eusebius, Clement of Alexandria, and Origen.
  • Hegesippus (110CE-180CE) he arrived in Rome around 157CE-168CE and wrote around 174CE-180CE.
  • Tatian the Syrian (120CE-185CE) was a convert to Christianity and was expelled sometime after 165CE for his ascetic and gnostic views.
  • Polycrates of Ephesus (130CE-196CE) he came from a family in the center of Christianity. Seven of his relatives were Bishops. He states in his letter to Victor around 186CE-195CE that he has served the Lord for 65 years.
  • Irenaeus of Lyon (130CE-202CE) who grew up in the Church. He wrote Against Heresies around 180CE to refute Gnosticism, promote Monotheism in books 1 and 2, counter Marcion's dual-god belief in Book 3, and names the Gospels.
  • Athenagoras of Athens (133CE-190CE) was a convert to Christianity. He wrote Legatio Pro Christianis to Marcus Aurelius around 176CE-177CE. There are only two mentions of him in early Christian Literature.
  • Clement of Alexandria (150CE-215CE) writes between 195CE-203CE, outside the potential range. However, he does use phrases similar to "most excellent."
The statement in Luke 1:3 "as one having a grasp of everything from the start/beginning/for a long time" Could mean they were in Christianity for a long time or they were born into it. It could be someone who is unknown or one of two candidates that fit the description very well: Polycrates and Irenaeus.

"Genuine" Pauline Epistles

The Most Frequent Words (MFWs) parameter was set to 200, as the computational analysis in Stylo relies on word usage frequency. This parameter selection is informed by the observation that the majority of the files (excluding Philemon) average fewer than 200 unique words across their chapters. Targeting this average effectively eliminates rare terms that appear only once or twice with single-use vocabulary accounting for approximately 10% of total unique words, while preserving high-frequency vocabulary. This configuration yields satisfactory analytical outcomes and doubles the default baseline setting of 100 MFWs.

Here is a tab separated value (.tsv) file with all of the word counts per chapter and book. It is broken down by 20 words per row. In all of the examples 200 words (10 rows) retrieves all of the words used multiple times and words that are only used once, and in some cases all of the words in the chapter itself.

A chart for Philemon is not included since it is only one chapter.

Dendrogram of 1 Corinthians.
Dendrogram of 2 Corinthians.
Dendrogram of Galatians.
Dendrogram of Philippians.
Dendrogram of Romans.
Dendrogram of 1 Thessalonians.

Epistle Data

1 Corinthians
Chapter 1500
Chapter 2287
Chapter 3340
Chapter 4345
Chapter 5221
Chapter 6334
Chapter 7688
Chapter 8226
Chapter 9449
Chapter 10463
Chapter 11529
Chapter 12465
Chapter 13197
Chapter 14606
Chapter 15843
Chapter 16323
Total Words6816
Avg Count426
Total Unique2136
Avg Unique133.5
2 Corinthians
Chapter 1490
Chapter 2285
Chapter 3296
Chapter 4320
Chapter 5338
Chapter 6266
Chapter 7328
Chapter 8410
Chapter 9284
Chapter 10311
Chapter 11500
Chapter 12412
Chapter 13236
Total Words4476
Avg Count344.3
Total Unique1526
Avg Unique117.38
Galatians
Chapter 1364
Chapter 2385
Chapter 3455
Chapter 4444
Chapter 5313
Chapter 6267
Total Words2228
Avg Count371.33
Total Unique920
Avg Unique153.33
Philippians
Chapter 1503
Chapter 2431
Chapter 3337
Chapter 4357
Total Words1628
Avg Count407
Total Unique713
Avg Unique178.25
Romans
Chapter 1545
Chapter 2449
Chapter 3428
Chapter 4399
Chapter 5432
Chapter 6367
Chapter 7467
Chapter 8652
Chapter 9525
Chapter 10338
Chapter 11578
Chapter 12304
Chapter 13270
Chapter 14379
Chapter 15541
Chapter 16383
Total Words7057
Avg Count441.06
Total Unique2146
Avg Unique134.12
1 Thessalonians
Chapter 1213
Chapter 2390
Chapter 3247
Chapter 4309
Chapter 5317
Total Words1476
Avg Count295.2
Total Unique589
Avg Unique117.8
Philemon
Chapter 1336
Total Words336
Avg Count336
Total Unique204
Avg Unique204

All Epistle Chapters Compared

The dendrogram encompassing all chapters of the Epistles reveals three main clusterings, with two exhibiting a higher degree of proximity to each other than to the third. Additionally, two to three sub-clusters are identifiable within these primary groups. Cross-collaboration is evident in Romans and Galatians between the top and middle clusters, whereas the initial and concluding sections of Romans appear to have been expanded by the most distant group. The top cluster accounts for the majority of 1 Corinthians, while the middle cluster is predominantly associated with 2 Corinthians. The bottom cluster of writers is exclusively responsible for, or provided the mojority via expansion to 1 Thessalonians and Philemon, as well as the larger portion of Philippians. This is consistent with Dr. BeDuhn's reconstruction containing a shorter version of Philippians that was subsequently expanded by the bottom group. It is demonstrated that they were responsible for the majority of the contents of chapters 1, 2, and addition of a 4TH chapter. The dendrogram reveals there is enough of the original text in chapter 3 that the authorship is placed into the first group. This is consistent with the pattern observed in Romans and the opening chapter of 1 Corinthians. This finding, combined with the substantial distance separating it from the other clusters, indicates that the bottom group introduced its additions and revisions subsequent to the work of the former two groups.

An examination of the consolidated dendrogram reveals that the majority of chapters maintain their structural proximity relative to the independent analyses, though certain deviations occur when the full corpus is integrated. These shifts signify a deeper lexical affinity with the later additions than the initial metrics suggest when evaluating individual epistles in isolation. Such patterns not only distinguish the presence of various contributors but also uncover a deliberate effort to harmonize these later insertions with the stylistic precursors. While these collaborative layers may remain obscured at the specific chapter level, the full comparative synthesis renders these editorial nuances distinct.

Dendrogram of 'Genuine' Pauline Epistles.

Stylo citation

Eder, M., Rybicki, J. and Kestemont, M. (2016). Stylometry with R:
a package for computational text analysis. R Journal 8(1): 107-121.
<https://journal.r-project.org/archive/2016/RJ-2016-007/index.html>