Stylometric Fingerprinting

Last generated: 2026-10-10 21:18:11 AEDT. Detector: translation_guidance_scan_v4.

This page tracks exploratory stylometric work for Stephanos, especially whether formula, gloss, grammar, and source-form features can expose different epitomising layers. The first implemented feature family is the existing translation-guidance recogniser output: each entry gets a vector of rule hits and non-hits, then the vectors are embedded and clustered.

3,571lemmas with at least one active guidance scan
3,571lemmas in the broad formula-vector UMAP slice
321Kappa entries in that UMAP slice
19Parisinus/non-epitomised entries in that UMAP slice
formula: 32, gloss: 70, proper_noun: 75active recogniser rules by kind
0.121KMeans silhouette on broad formula vectors
Current interpretation: the broad formula-vector slice shows an apparent Kappa signal, but the complete-formula slice does not. That means the first visible separation is probably mixed with scan-history and rule-coverage effects. Treat this as a progress marker, not as a demonstrated epitomiser fingerprint.

UMAP

The plotted vectors use formula recogniser rows only, with occurrence counts capped at 5 and transformed by log1p. The broad slice requires at least 21 of the 32 active formula rules to have been checked. Embedding method: UMAP.

Cluster Summary

ClusterRowsKappaParisinusMedian entryMean matched rulesTop over-represented formulae
289973 (8.1%)21044.5τὸ ἐθνικὸν + X (nominative ETHNONYM) (+0.26); X (SETTLEMENT) + Y (genitive REGION) (+0.14); X... πλησίον + Y. (genitive) (+0.00)
366866 (9.9%)21095.3X (SETTLEMENT) + Y (genitive PEOPLE) (+0.79); X (SETTLEMENT) + Y (genitive REGION) (+0.23); τὸ ἐθνικὸν + X (nominative ETHNONYM) (+0.05); διὰ τοῦ + «X» (GREEK LETTER) + [FORM OF γράφειν] (+0.01)
052140 (7.7%)21002.9X... + πρός + Y (dative) (+0.02); εἰς + «X» (GREEK LETTER) (+0.00); X... πλησίον + Y. (genitive) (+0.00)
740138 (9.5%)11108.1X (nominative PROPER NOUN) ... + ἀπό + Y (genitive ETYMON) (+0.87); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.55); X (nominative PROPER NOUN)... + ἀπό + Y (genitive ARTICLE + genitive ETYMON) (+0.54); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.29)
135832 (8.9%)21009.2ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.78); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.71); ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) (+0.39); ὡς + X (AUTHOR NAME) (+0.30)
635220 (5.7%)8978.4Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.85); X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) (+0.76); X (AUTHOR NAME) + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.45); X (AUTHOR NAME) + Y (NUMERAL) (+0.33)
517730 (16.9%)15311.4X (nominative AUTHOR NAME)... + ὡς αὐτός (+1.00); ὡς + X (AUTHOR NAME) (+0.79); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.63); ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.50)
813311 (8.3%)0797.5X (nominative DERIVED NOUN) + παρά + τό + Y (ETYMON) (+0.98); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.22); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.15); X (nominative DERIVED NOUN) + παρά + τό + Y (INFINITIVE denoting an etymology) (+0.13)
45811 (19.0%)16512.4X (MASCULINE ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.96); X (NEUTER ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.91); [NO ANTECEDENT] + ἀφ' οὗ + Y (DERIVED NOUN) (+0.91); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.31)
940 (0.0%)011711.5[definite article] + εἰς + «X» (GREEK LETTER) + [FORM OF ληγών AGREEING WITH ART] (+1.00); εἰς + «X» (GREEK LETTER) (+0.49); ὁ πολίτης + X (nominative GENTILIC) (+0.36); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.24)

Separation Checks

CheckRowsPositive classResultInterpretation
Kappa vs non-Kappa, broad formula slice 3,571 321 Kappa 0.593 balanced accuracy across folds [0.581, 0.559, 0.627, 0.614, 0.586] Useful as a warning flag, but confounded by recogniser coverage and rule history.
Kappa vs non-Kappa, complete formula slice 2,379 321 Kappa 0.540 balanced accuracy across folds [0.552, 0.560, 0.529, 0.562, 0.501] The near-baseline result is evidence against claiming a robust Kappa fingerprint from current formula hits alone.
Parisinus/non-epitomised comparison 2,379 complete formula rows 19 Parisinus descriptive only The current non-epitomised sample is too small for a serious classifier; use it as a qualitative control set.

Coverage

Corpus Versions

VersionRowsUsableWith Greek text
epitome3,6643,6643,551
parisinus191919

Guidance Scan Counts

KindScan rowsLemmas checkedRules checkedLemmas matchedOccurrences
formula156,1633,571323,44526,348
gloss174,3063,571702,4526,060
proper_noun5715712

Formula Coverage Distribution

Formula rules checkedLemmas
211,190
312
322,379

Sentence Grammar Feature Tables

TableRows
sentence_grammar_runs9
sentence_grammar_evaluations111
sentence_grammar_tokens578

Next Implementation Steps

  1. Freeze a coverage-balanced formula/gloss feature matrix and rerun the UMAP after every nightly guidance scan.
  2. Add non-recogniser stylometric baselines: character n-grams, function-word rates, particles, clause connectors, entry length, and normalized type-token measures.
  3. Populate the sentence-grammar tables over a coverage-balanced sample, then test morphosyntactic vectors separately from formula vectors.
  4. Expand the non-epitomised control set beyond the current Parisinus rows before making any claim about epitomiser layers.
  5. Validate clusters by close reading: each cluster needs formula examples and counterexamples before it becomes an argument.