Stylometric Fingerprinting
Last generated: 2026-07-13 22:08:27 AEST. Detector: translation_guidance_scan_v4.
This page tracks exploratory stylometric work for Stephanos, especially whether formula,
gloss, grammar, and source-form features can expose different epitomising layers. The first
implemented feature family is the existing translation-guidance recogniser output: each entry
gets a vector of rule hits and non-hits, then the vectors are embedded and clustered.
3,570lemmas with at least one active guidance scan
2,895lemmas in the broad formula-vector UMAP slice
320Kappa entries in that UMAP slice
13Parisinus/non-epitomised entries in that UMAP slice
formula: 32, gloss: 70, proper_noun: 75active recogniser rules by kind
0.077KMeans silhouette on broad formula vectors
Current interpretation: the broad formula-vector slice shows an apparent Kappa
signal, but the complete-formula slice does not. That means the first visible separation is
probably mixed with scan-history and rule-coverage effects. Treat this as a progress marker,
not as a demonstrated epitomiser fingerprint.
UMAP
The plotted vectors use formula recogniser rows only, with occurrence counts capped at 5 and
transformed by log1p. The broad slice requires at least 21 of the 32
active formula rules to have been checked. Embedding method: UMAP.
Cluster Summary
| Cluster | Rows | Kappa | Parisinus | Median entry | Mean matched rules | Top over-represented formulae |
|---|
| 1 | 1,057 | 103 (9.7%) | 1 | 101 | 2.1 | X... πλησίον + Y. (genitive) (+0.00) |
| 6 | 666 | 69 (10.4%) | 0 | 114 | 4.3 | X (SETTLEMENT) + Y (genitive REGION) (+0.51); X (SETTLEMENT) + Y (genitive PEOPLE) (+0.47); ὁ πολίτης + X (nominative GENTILIC) (+0.13); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.10) |
| 7 | 419 | 18 (4.3%) | 2 | 104 | 6.2 | ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.84); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.76); ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) (+0.34); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.27) |
| 8 | 264 | 12 (4.5%) | 4 | 106 | 6.0 | Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.90); X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) (+0.79); X (AUTHOR NAME) + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.50); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.13) |
| 2 | 220 | 20 (9.1%) | 3 | 122 | 7.4 | X (nominative PROPER NOUN)... + ἀπό + Y (genitive ARTICLE + genitive ETYMON) (+0.94); X (nominative PROPER NOUN) ... + ἀπό + Y (genitive ETYMON) (+0.79); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.46); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.31) |
| 5 | 168 | 19 (11.3%) | 2 | 94 | 4.8 | X... + πρός + Y (dative) (+0.99); ὁ πολίτης + X (nominative GENTILIC) (+0.05); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.01); X... πλησίον + Y. (genitive) (+0.01) |
| 0 | 69 | 56 (81.2%) | 0 | 127 | 12.3 | ὡς + X (AUTHOR NAME) (+0.91); X (AUTHOR NAME) + Y (NUMERAL) (+0.59); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.55); X (nominative AUTHOR NAME)... + ὡς αὐτός (+0.51) |
| 4 | 12 | 11 (91.7%) | 0 | 104 | 15.1 | X (NEUTER ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+1.00); [NO ANTECEDENT] + ἀφ' οὗ + Y (DERIVED NOUN) (+0.92); X (MASCULINE ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.92); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.65) |
| 3 | 11 | 10 (90.9%) | 0 | 63 | 11.4 | διὰ τοῦ + «X» (GREEK LETTER) + [FORM OF γράφειν] (+1.00); ὡς + X (AUTHOR NAME) (+0.33); Y (EPITHET) + X (DEITY) (+0.27); X (nominative AUTHOR NAME)... + ὡς αὐτός (+0.17) |
| 9 | 9 | 2 (22.2%) | 1 | 107 | 11.0 | ἐκάλειτο X (nominative PROPER NOUN)... + κέκληται Y (nominative PROPER NOUN) (+1.00); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.40); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.34); Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.23) |
Separation Checks
| Check | Rows | Positive class | Result | Interpretation |
| Kappa vs non-Kappa, broad formula slice |
2,895 |
320 Kappa |
0.754 balanced accuracy across folds [0.752, 0.747, 0.765, 0.727, 0.780] |
Useful as a warning flag, but confounded by recogniser coverage and rule history. |
| Kappa vs non-Kappa, complete formula slice |
380 |
320 Kappa |
0.566 balanced accuracy across folds [0.703, 0.547, 0.547, 0.497, 0.534] |
The near-baseline result is evidence against claiming a robust Kappa fingerprint from current formula hits alone. |
| Parisinus/non-epitomised comparison |
380 complete formula rows |
1 Parisinus |
descriptive only |
The current non-epitomised sample is too small for a serious classifier; use it as a qualitative control set. |
Coverage
Corpus Versions
| Version | Rows | Usable | With Greek text |
|---|
| epitome | 3,551 | 3,551 | 3,551 |
| parisinus | 19 | 19 | 19 |
Guidance Scan Counts
| Kind | Scan rows | Lemmas checked | Rules checked | Lemmas matched | Occurrences |
|---|
| formula | 75,798 | 2,897 | 32 | 2,749 | 14,107 |
| gloss | 38,692 | 3,570 | 70 | 906 | 1,825 |
| proper_noun | 57 | 1 | 57 | 1 | 2 |
Formula Coverage Distribution
| Formula rules checked | Lemmas |
|---|
| 0 | 673 |
| 3 | 1 |
| 20 | 1 |
| 21 | 2,505 |
| 29 | 10 |
| 32 | 380 |
Sentence Grammar Feature Tables
| Table | Rows |
|---|
| sentence_grammar_runs | 8 |
| sentence_grammar_evaluations | 111 |
| sentence_grammar_tokens | 495 |
Next Implementation Steps
- Freeze a coverage-balanced formula/gloss feature matrix and rerun the UMAP after every nightly guidance scan.
- Add non-recogniser stylometric baselines: character n-grams, function-word rates, particles, clause connectors, entry length, and normalized type-token measures.
- Populate the sentence-grammar tables over a coverage-balanced sample, then test morphosyntactic vectors separately from formula vectors.
- Expand the non-epitomised control set beyond the current Parisinus rows before making any claim about epitomiser layers.
- Validate clusters by close reading: each cluster needs formula examples and counterexamples before it becomes an argument.