Stylometric Fingerprinting
Last generated: 2026-10-10 21:18:11 AEDT. Detector: translation_guidance_scan_v4.
This page tracks exploratory stylometric work for Stephanos, especially whether formula,
gloss, grammar, and source-form features can expose different epitomising layers. The first
implemented feature family is the existing translation-guidance recogniser output: each entry
gets a vector of rule hits and non-hits, then the vectors are embedded and clustered.
3,571lemmas with at least one active guidance scan
3,571lemmas in the broad formula-vector UMAP slice
321Kappa entries in that UMAP slice
19Parisinus/non-epitomised entries in that UMAP slice
formula: 32, gloss: 70, proper_noun: 75active recogniser rules by kind
0.121KMeans silhouette on broad formula vectors
Current interpretation: the broad formula-vector slice shows an apparent Kappa
signal, but the complete-formula slice does not. That means the first visible separation is
probably mixed with scan-history and rule-coverage effects. Treat this as a progress marker,
not as a demonstrated epitomiser fingerprint.
UMAP
The plotted vectors use formula recogniser rows only, with occurrence counts capped at 5 and
transformed by log1p. The broad slice requires at least 21 of the 32
active formula rules to have been checked. Embedding method: UMAP.
Cluster Summary
| Cluster | Rows | Kappa | Parisinus | Median entry | Mean matched rules | Top over-represented formulae |
|---|
| 2 | 899 | 73 (8.1%) | 2 | 104 | 4.5 | τὸ ἐθνικὸν + X (nominative ETHNONYM) (+0.26); X (SETTLEMENT) + Y (genitive REGION) (+0.14); X... πλησίον + Y. (genitive) (+0.00) |
| 3 | 668 | 66 (9.9%) | 2 | 109 | 5.3 | X (SETTLEMENT) + Y (genitive PEOPLE) (+0.79); X (SETTLEMENT) + Y (genitive REGION) (+0.23); τὸ ἐθνικὸν + X (nominative ETHNONYM) (+0.05); διὰ τοῦ + «X» (GREEK LETTER) + [FORM OF γράφειν] (+0.01) |
| 0 | 521 | 40 (7.7%) | 2 | 100 | 2.9 | X... + πρός + Y (dative) (+0.02); εἰς + «X» (GREEK LETTER) (+0.00); X... πλησίον + Y. (genitive) (+0.00) |
| 7 | 401 | 38 (9.5%) | 1 | 110 | 8.1 | X (nominative PROPER NOUN) ... + ἀπό + Y (genitive ETYMON) (+0.87); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.55); X (nominative PROPER NOUN)... + ἀπό + Y (genitive ARTICLE + genitive ETYMON) (+0.54); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.29) |
| 1 | 358 | 32 (8.9%) | 2 | 100 | 9.2 | ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.78); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.71); ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) (+0.39); ὡς + X (AUTHOR NAME) (+0.30) |
| 6 | 352 | 20 (5.7%) | 8 | 97 | 8.4 | Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.85); X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) (+0.76); X (AUTHOR NAME) + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.45); X (AUTHOR NAME) + Y (NUMERAL) (+0.33) |
| 5 | 177 | 30 (16.9%) | 1 | 53 | 11.4 | X (nominative AUTHOR NAME)... + ὡς αὐτός (+1.00); ὡς + X (AUTHOR NAME) (+0.79); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.63); ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.50) |
| 8 | 133 | 11 (8.3%) | 0 | 79 | 7.5 | X (nominative DERIVED NOUN) + παρά + τό + Y (ETYMON) (+0.98); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.22); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.15); X (nominative DERIVED NOUN) + παρά + τό + Y (INFINITIVE denoting an etymology) (+0.13) |
| 4 | 58 | 11 (19.0%) | 1 | 65 | 12.4 | X (MASCULINE ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.96); X (NEUTER ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.91); [NO ANTECEDENT] + ἀφ' οὗ + Y (DERIVED NOUN) (+0.91); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.31) |
| 9 | 4 | 0 (0.0%) | 0 | 117 | 11.5 | [definite article] + εἰς + «X» (GREEK LETTER) + [FORM OF ληγών AGREEING WITH ART] (+1.00); εἰς + «X» (GREEK LETTER) (+0.49); ὁ πολίτης + X (nominative GENTILIC) (+0.36); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.24) |
Separation Checks
| Check | Rows | Positive class | Result | Interpretation |
| Kappa vs non-Kappa, broad formula slice |
3,571 |
321 Kappa |
0.593 balanced accuracy across folds [0.581, 0.559, 0.627, 0.614, 0.586] |
Useful as a warning flag, but confounded by recogniser coverage and rule history. |
| Kappa vs non-Kappa, complete formula slice |
2,379 |
321 Kappa |
0.540 balanced accuracy across folds [0.552, 0.560, 0.529, 0.562, 0.501] |
The near-baseline result is evidence against claiming a robust Kappa fingerprint from current formula hits alone. |
| Parisinus/non-epitomised comparison |
2,379 complete formula rows |
19 Parisinus |
descriptive only |
The current non-epitomised sample is too small for a serious classifier; use it as a qualitative control set. |
Coverage
Corpus Versions
| Version | Rows | Usable | With Greek text |
|---|
| epitome | 3,664 | 3,664 | 3,551 |
| parisinus | 19 | 19 | 19 |
Guidance Scan Counts
| Kind | Scan rows | Lemmas checked | Rules checked | Lemmas matched | Occurrences |
|---|
| formula | 156,163 | 3,571 | 32 | 3,445 | 26,348 |
| gloss | 174,306 | 3,571 | 70 | 2,452 | 6,060 |
| proper_noun | 57 | 1 | 57 | 1 | 2 |
Formula Coverage Distribution
| Formula rules checked | Lemmas |
|---|
| 21 | 1,190 |
| 31 | 2 |
| 32 | 2,379 |
Sentence Grammar Feature Tables
| Table | Rows |
|---|
| sentence_grammar_runs | 9 |
| sentence_grammar_evaluations | 111 |
| sentence_grammar_tokens | 578 |
Next Implementation Steps
- Freeze a coverage-balanced formula/gloss feature matrix and rerun the UMAP after every nightly guidance scan.
- Add non-recogniser stylometric baselines: character n-grams, function-word rates, particles, clause connectors, entry length, and normalized type-token measures.
- Populate the sentence-grammar tables over a coverage-balanced sample, then test morphosyntactic vectors separately from formula vectors.
- Expand the non-epitomised control set beyond the current Parisinus rows before making any claim about epitomiser layers.
- Validate clusters by close reading: each cluster needs formula examples and counterexamples before it becomes an argument.