Stylometry
Sytometry sty·lom·e·try /stīˈlämətrē/
the statistical analysis of variations in literary style between one writer or genre and another.
FAQ
- What do you use to perform Stylometric analysis? Stylo is a package developed and maintained by Computational Stylistics Group for R Studio.
- Can I use Stylo with Python? PyStyl and faststylometry are two Python libraries inspired by Stylo.
- Where are the peer reviewed studies using Stylo? The Publications link on the Computational Stylistics Group website contains links to papers using Stylometry, Note: there are peer reviewed links on that page, but not all of the links are to peer reviewed papers.
- Does it utilize machine-learning? Several machine-learning methods are available for supervised runs of Stylo with the classify() function. Only unsupervised runs are utilized in these examples.
- What's the difference between supervised and unsupervised runs? Supervized runs used labeled data. Meaning you know the author beforehand and you're training their known works to identify if that person has written other works not attributed to them. Unsupervised runs compare the works provided without any influence from a specific writer. This allows the works to be compared to each other independently.
- Can it be used by AI/LLMs? Yes. It was used in one of the linked publications.
- Did Stylometry really help convict the Unabomber? Yes but they did it by hand.
- Why don't people do it by hand? Using a computer for the analysis eliminates virtually all human error (with the exception of bad/malformed data), bias, and is much faster than human analysis.
- Does the file name influence the results? No. Though, to control for a perceived bias, in the first example the file names have been obfuscated.
- Do you need to know the language being analyzed? You only need to be able to identify what language is in the file, so you can choose it appropriately in Stylo.
- Do you need to have writings from throughout the person's life to accurately identify them? No; the individual texts are compared to one another in unsupervised runs.
- What is the recommended amount of most frequent words (MFW)? It's recommended that at least 100-200 MFWs are used to get a strong authorial signal. However, there are instances where authorship can be determined with fewer words, and that can be observed by outputting the results in increments. Increasing engrams can also help identify someone through repeated phrases. This is demonstrated below.
- Where did you get the source data? I used two primary sources. For the New Testament I used the Society of Biblical Literature Greek New Testament (SBLGNT). The data is available in XML and can easily be split by chapter. Other Greek texts are available at Perseus at Tufts University. The text of Evangelion is from Github by Mark Bilby.
- How accurate is Stylometry? It has been shown to identify Greek genres with >97% accuracy. A new methodology has been demonstrated to have an accuracy of around 84.2%-89.6% for texts containing 50 words. The paper shows accuracy rates as high as 92% with 200-word passages. Eder (from Computational Stylistics) has found that 2000 words may be sufficient, but he has also said some authors can be determined with a 100 word sample. The paper he referenced can be accessed here. Take note: The samples he ran are in English, which is considered weakly inflective compared to the Greek and Latin demonstrated on this page.
- Does Stylometry prove Jesus didn't exist? Nobody is making that claim. However, it can be inferred indirectly. As it says in the definition above, it can identify authorship, including if someone was not the author. By identifying authorship and knowing when that person lived it logically follows that it was written during their lifetime, which in turn can also identify it wasn't written during someone else's lifetime if that person wasn't alive during the identified authors lifetime. It also assists in pinpointing interpolated texts and resolving the chronological debate of Evangelion predating Luke. Which is additional evidence for the 2ND century authorship of the New Testament.
- Are you a data scientist? I'm not a data scientist, but if the science is sound it is repeatable. I am a senior software engineer by trade but I have also had previous work peer reviewed by a data scientist. A process that took 8 months to defend and was determined to be error free on a dataset of several million records that was updated quarterly.
- Why not use Bayesian statistics? A design feature of Bayesian statistics is the use of non-zero probabilities in order for future updates; if something is impossible or completely rules out the likelihood or expected value, then it can't be used. This feature can be useful in some cases, but informative non-zero priors are subjective and can introduce a bias because something must be possible/probabile. Example: Bayesian statistics would require the possibility of being able to swallow my own head whole to return a non-zero result. Bayesian statics should also disclose their hypothesis/prior, evidence/likelihood, expected probability, and how those values were determined.
- Didn't Dustin White debunk the how Christ Before Jesus used Stylo? He tried to, but there were many flaws in his attempt.
Stylometry Examples
Influence of time

Control Test


Testimonium Flavianum



As you can see, the results are similar to the 3 engram run at 2000 words.
The sample sizes are still more than what's recommended.
Total words: 413971
Average words: 13353.90322580645
Total unique words: 56806
Average unique words: 1832.4516129032259
* Note: Unlike the run above, this run does not include the TF in book 18 or the "so called Christ" in book 20.

- book 1: 9910
- book 2: 9180
- book 3: 12001
- book 4: 10165
- book 5: 12957
- book 6: 12930
- book 7: 11441
- book 8: 7745
- book 9: 6036
- book 10: 9548
- book 1: 15275
- book 2: 14941
- book 3: 14061
- book 4: 14117
- book 5: 14987
- book 6: 18506
- book 7: 18917
- book 8: 20259
- book 9: 13538
- book 10: 12714
- book 11: 13866
- book 12: 17159
- book 13: 17620
- book 14: 19626
- book 15: 16486
- book 16: 15162
- book 17: 15268
- book 18: 15968
- book 19: 13660
- book 20: 9839
- Testimonium Flavianum: 89
The first row represents 33,974 words used only once, while the last row is one word used 23,713 times.
Gospel of Luke


Luke compared to Acts



Who wrote most of Luke?


- Melito of Sardis (100CE-180CE) was a Christian convert with Jewish and Hellenistic roots. Most of his works have been lost but he was quoted by Eusebius, Clement of Alexandria, and Origen.
- Hegesippus (110CE-180CE) he arrived in Rome around 157CE-168CE and wrote around 174CE-180CE.
- Tatian the Syrian (120CE-185CE) was a convert to Christianity and was expelled sometime after 165CE for his ascetic and gnostic views.
- Polycrates of Ephesus (130CE-196CE) he came from a family in the center of Christianity. Seven of his relatives were Bishops. He states in his letter to Victor around 186CE-195CE that he has served the Lord for 65 years.
- Irenaeus of Lyon (130CE-202CE) who grew up in the Church. He wrote Against Heresies around 180CE to refute Gnosticism, promote Monotheism in books 1 and 2, counter Marcion's dual-god belief in Book 3, and names the Gospels.
- Athenagoras of Athens (133CE-190CE) was a convert to Christianity. He wrote Legatio Pro Christianis to Marcus Aurelius around 176CE-177CE. There are only two mentions of him in early Christian Literature.
- Clement of Alexandria (150CE-215CE) writes between 195CE-203CE, outside the potential range. However, he does use phrases similar to "most excellent."
"Genuine" Pauline Epistles
The Most Frequent Words (MFWs) parameter was set to 200, as the computational analysis in Stylo relies on word usage frequency. This parameter selection is informed by the observation that the majority of the files (excluding Philemon) average fewer than 200 unique words across their chapters. Targeting this average effectively eliminates rare terms that appear only once or twice with single-use vocabulary accounting for approximately 10% of total unique words, while preserving high-frequency vocabulary. This configuration yields satisfactory analytical outcomes and doubles the default baseline setting of 100 MFWs.
Here is a tab separated value (.tsv) file with all of the word counts per chapter and book. It is broken down by 20 words per row. In all of the examples 200 words (10 rows) retrieves all of the words used multiple times and words that are only used once, and in some cases all of the words in the chapter itself.
A chart for Philemon is not included since it is only one chapter.






Epistle Data
| 1 Corinthians | |
|---|---|
| Chapter 1 | 500 |
| Chapter 2 | 287 |
| Chapter 3 | 340 |
| Chapter 4 | 345 |
| Chapter 5 | 221 |
| Chapter 6 | 334 |
| Chapter 7 | 688 |
| Chapter 8 | 226 |
| Chapter 9 | 449 |
| Chapter 10 | 463 |
| Chapter 11 | 529 |
| Chapter 12 | 465 |
| Chapter 13 | 197 |
| Chapter 14 | 606 |
| Chapter 15 | 843 |
| Chapter 16 | 323 |
| Total Words | 6816 |
| Avg Count | 426 |
| Total Unique | 2136 |
| Avg Unique | 133.5 |
| 2 Corinthians | |
|---|---|
| Chapter 1 | 490 |
| Chapter 2 | 285 |
| Chapter 3 | 296 |
| Chapter 4 | 320 |
| Chapter 5 | 338 |
| Chapter 6 | 266 |
| Chapter 7 | 328 |
| Chapter 8 | 410 |
| Chapter 9 | 284 |
| Chapter 10 | 311 |
| Chapter 11 | 500 |
| Chapter 12 | 412 |
| Chapter 13 | 236 |
| Total Words | 4476 |
| Avg Count | 344.3 |
| Total Unique | 1526 |
| Avg Unique | 117.38 |
| Galatians | |
|---|---|
| Chapter 1 | 364 |
| Chapter 2 | 385 |
| Chapter 3 | 455 |
| Chapter 4 | 444 |
| Chapter 5 | 313 |
| Chapter 6 | 267 |
| Total Words | 2228 |
| Avg Count | 371.33 |
| Total Unique | 920 |
| Avg Unique | 153.33 |
| Philippians | |
|---|---|
| Chapter 1 | 503 |
| Chapter 2 | 431 |
| Chapter 3 | 337 |
| Chapter 4 | 357 |
| Total Words | 1628 |
| Avg Count | 407 |
| Total Unique | 713 |
| Avg Unique | 178.25 |
| Romans | |
|---|---|
| Chapter 1 | 545 |
| Chapter 2 | 449 |
| Chapter 3 | 428 |
| Chapter 4 | 399 |
| Chapter 5 | 432 |
| Chapter 6 | 367 |
| Chapter 7 | 467 |
| Chapter 8 | 652 |
| Chapter 9 | 525 |
| Chapter 10 | 338 |
| Chapter 11 | 578 |
| Chapter 12 | 304 |
| Chapter 13 | 270 |
| Chapter 14 | 379 |
| Chapter 15 | 541 |
| Chapter 16 | 383 |
| Total Words | 7057 |
| Avg Count | 441.06 |
| Total Unique | 2146 |
| Avg Unique | 134.12 |
| 1 Thessalonians | |
|---|---|
| Chapter 1 | 213 |
| Chapter 2 | 390 |
| Chapter 3 | 247 |
| Chapter 4 | 309 |
| Chapter 5 | 317 |
| Total Words | 1476 |
| Avg Count | 295.2 |
| Total Unique | 589 |
| Avg Unique | 117.8 |
| Philemon | |
|---|---|
| Chapter 1 | 336 |
| Total Words | 336 |
| Avg Count | 336 |
| Total Unique | 204 |
| Avg Unique | 204 |
All Epistle Chapters Compared
The dendrogram encompassing all chapters of the Epistles reveals three main clusterings, with two exhibiting a higher degree of proximity to each other than to the third. Additionally, two to three sub-clusters are identifiable within these primary groups. Cross-collaboration is evident in Romans and Galatians between the top and middle clusters, whereas the initial and concluding sections of Romans appear to have been expanded by the most distant group. The top cluster accounts for the majority of 1 Corinthians, while the middle cluster is predominantly associated with 2 Corinthians. The bottom cluster of writers is exclusively responsible for, or provided the mojority via expansion to 1 Thessalonians and Philemon, as well as the larger portion of Philippians. This is consistent with Dr. BeDuhn's reconstruction containing a shorter version of Philippians that was subsequently expanded by the bottom group. It is demonstrated that they were responsible for the majority of the contents of chapters 1, 2, and addition of a 4TH chapter. The dendrogram reveals there is enough of the original text in chapter 3 that the authorship is placed into the first group. This is consistent with the pattern observed in Romans and the opening chapter of 1 Corinthians. This finding, combined with the substantial distance separating it from the other clusters, indicates that the bottom group introduced its additions and revisions subsequent to the work of the former two groups.
An examination of the consolidated dendrogram reveals that the majority of chapters maintain their structural proximity relative to the independent analyses, though certain deviations occur when the full corpus is integrated. These shifts signify a deeper lexical affinity with the later additions than the initial metrics suggest when evaluating individual epistles in isolation. Such patterns not only distinguish the presence of various contributors but also uncover a deliberate effort to harmonize these later insertions with the stylistic precursors. While these collaborative layers may remain obscured at the specific chapter level, the full comparative synthesis renders these editorial nuances distinct.

Stylo citation
a package for computational text analysis. R Journal 8(1): 107-121.
<https://journal.r-project.org/archive/2016/RJ-2016-007/index.html>