Stylometric Fingerprinting
Last generated: 2026-08-24 22:06:24 AEST. Detector: translation_guidance_scan_v4.
This page tracks exploratory stylometric work for Stephanos, especially whether formula,
gloss, grammar, and source-form features can expose different epitomising layers. The first
implemented feature family is the existing translation-guidance recogniser output: each entry
gets a vector of rule hits and non-hits, then the vectors are embedded and clustered.
3,571lemmas with at least one active guidance scan
3,571lemmas in the broad formula-vector UMAP slice
321Kappa entries in that UMAP slice
19Parisinus/non-epitomised entries in that UMAP slice
formula: 32, gloss: 70, proper_noun: 75active recogniser rules by kind
0.091KMeans silhouette on broad formula vectors
Current interpretation: the broad formula-vector slice shows an apparent Kappa
signal, but the complete-formula slice does not. That means the first visible separation is
probably mixed with scan-history and rule-coverage effects. Treat this as a progress marker,
not as a demonstrated epitomiser fingerprint.
UMAP
The plotted vectors use formula recogniser rows only, with occurrence counts capped at 5 and
transformed by log1p. The broad slice requires at least 21 of the 32
active formula rules to have been checked. Embedding method: UMAP.
Cluster Summary
| Cluster | Rows | Kappa | Parisinus | Median entry | Mean matched rules | Top over-represented formulae |
|---|
| 6 | 1,068 | 79 (7.4%) | 3 | 104 | 3.7 | τὸ ἐθνικὸν + X (nominative ETHNONYM) (+0.64); X (SETTLEMENT) + Y (genitive REGION) (+0.04); X (SETTLEMENT) + Y (genitive PEOPLE) (+0.02); X... πλησίον + Y. (genitive) (+0.01) |
| 8 | 969 | 83 (8.6%) | 2 | 98 | 2.6 | ὁ πολίτης + X (nominative GENTILIC) (+0.09); εἰς + «X» (GREEK LETTER) (+0.00); διὰ τοῦ + «X» (GREEK LETTER) + [FORM OF γράφειν] (+0.00) |
| 2 | 432 | 37 (8.6%) | 1 | 108 | 6.7 | X (nominative PROPER NOUN) ... + ἀπό + Y (genitive ETYMON) (+0.85); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.55); X (nominative PROPER NOUN)... + ἀπό + Y (genitive ARTICLE + genitive ETYMON) (+0.51); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.31) |
| 5 | 358 | 66 (18.4%) | 2 | 84 | 9.0 | X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.69); ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.67); ὡς + X (AUTHOR NAME) (+0.49); ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) (+0.43) |
| 1 | 269 | 16 (5.9%) | 0 | 102 | 6.4 | X (nominative PROPER NOUN) + X (genitive PROPER NOUN) (+0.96); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.21); ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.17); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.17) |
| 4 | 219 | 11 (5.0%) | 3 | 100 | 7.2 | X (AUTHOR NAME) + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.92); Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.71); X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) (+0.68); X (AUTHOR NAME) + Y (NUMERAL) (+0.13) |
| 0 | 199 | 9 (4.5%) | 6 | 92 | 6.7 | Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.93); X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) (+0.78); X (AUTHOR NAME) + Y (NUMERAL) (+0.07); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.07) |
| 3 | 27 | 11 (40.7%) | 1 | 97 | 15.4 | X (MASCULINE ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.96); [NO ANTECEDENT] + ἀφ' οὗ + Y (DERIVED NOUN) (+0.96); X (NEUTER ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.92); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.54) |
| 9 | 26 | 9 (34.6%) | 1 | 124 | 12.8 | Y (EPITHET) + X (DEITY) (+1.00); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.36); ὡς + X (AUTHOR NAME) (+0.23); X (AUTHOR NAME) + Y (NUMERAL) (+0.19) |
| 7 | 4 | 0 (0.0%) | 0 | 151 | 6.2 | [definite article] + εἰς + «X» (GREEK LETTER) + [FORM OF ληγών AGREEING WITH ART] (+1.00); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.44); X (SETTLEMENT) + Y (genitive PEOPLE) (+0.36); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.29) |
Separation Checks
| Check | Rows | Positive class | Result | Interpretation |
| Kappa vs non-Kappa, broad formula slice |
3,571 |
321 Kappa |
0.698 balanced accuracy across folds [0.678, 0.655, 0.719, 0.728, 0.711] |
Useful as a warning flag, but confounded by recogniser coverage and rule history. |
| Kappa vs non-Kappa, complete formula slice |
987 |
321 Kappa |
0.481 balanced accuracy across folds [0.502, 0.516, 0.458, 0.489, 0.438] |
The near-baseline result is evidence against claiming a robust Kappa fingerprint from current formula hits alone. |
| Parisinus/non-epitomised comparison |
987 complete formula rows |
17 Parisinus |
descriptive only |
The current non-epitomised sample is too small for a serious classifier; use it as a qualitative control set. |
Coverage
Corpus Versions
| Version | Rows | Usable | With Greek text |
|---|
| epitome | 3,664 | 3,664 | 3,551 |
| parisinus | 19 | 19 | 19 |
Guidance Scan Counts
| Kind | Scan rows | Lemmas checked | Rules checked | Lemmas matched | Occurrences |
|---|
| formula | 111,562 | 3,571 | 32 | 3,410 | 20,377 |
| gloss | 80,930 | 3,571 | 70 | 1,401 | 3,300 |
| proper_noun | 57 | 1 | 57 | 1 | 2 |
Formula Coverage Distribution
| Formula rules checked | Lemmas |
|---|
| 21 | 2,584 |
| 32 | 987 |
Sentence Grammar Feature Tables
| Table | Rows |
|---|
| sentence_grammar_runs | 8 |
| sentence_grammar_evaluations | 111 |
| sentence_grammar_tokens | 495 |
Next Implementation Steps
- Freeze a coverage-balanced formula/gloss feature matrix and rerun the UMAP after every nightly guidance scan.
- Add non-recogniser stylometric baselines: character n-grams, function-word rates, particles, clause connectors, entry length, and normalized type-token measures.
- Populate the sentence-grammar tables over a coverage-balanced sample, then test morphosyntactic vectors separately from formula vectors.
- Expand the non-epitomised control set beyond the current Parisinus rows before making any claim about epitomiser layers.
- Validate clusters by close reading: each cluster needs formula examples and counterexamples before it becomes an argument.