Stylometric Fingerprinting

Last generated: 2026-07-13 22:08:27 AEST. Detector: translation_guidance_scan_v4.

This page tracks exploratory stylometric work for Stephanos, especially whether formula, gloss, grammar, and source-form features can expose different epitomising layers. The first implemented feature family is the existing translation-guidance recogniser output: each entry gets a vector of rule hits and non-hits, then the vectors are embedded and clustered.

3,570lemmas with at least one active guidance scan
2,895lemmas in the broad formula-vector UMAP slice
320Kappa entries in that UMAP slice
13Parisinus/non-epitomised entries in that UMAP slice
formula: 32, gloss: 70, proper_noun: 75active recogniser rules by kind
0.077KMeans silhouette on broad formula vectors
Current interpretation: the broad formula-vector slice shows an apparent Kappa signal, but the complete-formula slice does not. That means the first visible separation is probably mixed with scan-history and rule-coverage effects. Treat this as a progress marker, not as a demonstrated epitomiser fingerprint.

UMAP

The plotted vectors use formula recogniser rows only, with occurrence counts capped at 5 and transformed by log1p. The broad slice requires at least 21 of the 32 active formula rules to have been checked. Embedding method: UMAP.

Cluster Summary

ClusterRowsKappaParisinusMedian entryMean matched rulesTop over-represented formulae
11,057103 (9.7%)11012.1X... πλησίον + Y. (genitive) (+0.00)
666669 (10.4%)01144.3X (SETTLEMENT) + Y (genitive REGION) (+0.51); X (SETTLEMENT) + Y (genitive PEOPLE) (+0.47); ὁ πολίτης + X (nominative GENTILIC) (+0.13); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.10)
741918 (4.3%)21046.2ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.84); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.76); ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) (+0.34); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.27)
826412 (4.5%)41066.0Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.90); X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) (+0.79); X (AUTHOR NAME) + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.50); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.13)
222020 (9.1%)31227.4X (nominative PROPER NOUN)... + ἀπό + Y (genitive ARTICLE + genitive ETYMON) (+0.94); X (nominative PROPER NOUN) ... + ἀπό + Y (genitive ETYMON) (+0.79); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.46); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.31)
516819 (11.3%)2944.8X... + πρός + Y (dative) (+0.99); ὁ πολίτης + X (nominative GENTILIC) (+0.05); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.01); X... πλησίον + Y. (genitive) (+0.01)
06956 (81.2%)012712.3ὡς + X (AUTHOR NAME) (+0.91); X (AUTHOR NAME) + Y (NUMERAL) (+0.59); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.55); X (nominative AUTHOR NAME)... + ὡς αὐτός (+0.51)
41211 (91.7%)010415.1X (NEUTER ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+1.00); [NO ANTECEDENT] + ἀφ' οὗ + Y (DERIVED NOUN) (+0.92); X (MASCULINE ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.92); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.65)
31110 (90.9%)06311.4διὰ τοῦ + «X» (GREEK LETTER) + [FORM OF γράφειν] (+1.00); ὡς + X (AUTHOR NAME) (+0.33); Y (EPITHET) + X (DEITY) (+0.27); X (nominative AUTHOR NAME)... + ὡς αὐτός (+0.17)
992 (22.2%)110711.0ἐκάλειτο X (nominative PROPER NOUN)... + κέκληται Y (nominative PROPER NOUN) (+1.00); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.40); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.34); Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.23)

Separation Checks

CheckRowsPositive classResultInterpretation
Kappa vs non-Kappa, broad formula slice 2,895 320 Kappa 0.754 balanced accuracy across folds [0.752, 0.747, 0.765, 0.727, 0.780] Useful as a warning flag, but confounded by recogniser coverage and rule history.
Kappa vs non-Kappa, complete formula slice 380 320 Kappa 0.566 balanced accuracy across folds [0.703, 0.547, 0.547, 0.497, 0.534] The near-baseline result is evidence against claiming a robust Kappa fingerprint from current formula hits alone.
Parisinus/non-epitomised comparison 380 complete formula rows 1 Parisinus descriptive only The current non-epitomised sample is too small for a serious classifier; use it as a qualitative control set.

Coverage

Corpus Versions

VersionRowsUsableWith Greek text
epitome3,5513,5513,551
parisinus191919

Guidance Scan Counts

KindScan rowsLemmas checkedRules checkedLemmas matchedOccurrences
formula75,7982,897322,74914,107
gloss38,6923,570709061,825
proper_noun5715712

Formula Coverage Distribution

Formula rules checkedLemmas
0673
31
201
212,505
2910
32380

Sentence Grammar Feature Tables

TableRows
sentence_grammar_runs8
sentence_grammar_evaluations111
sentence_grammar_tokens495

Next Implementation Steps

  1. Freeze a coverage-balanced formula/gloss feature matrix and rerun the UMAP after every nightly guidance scan.
  2. Add non-recogniser stylometric baselines: character n-grams, function-word rates, particles, clause connectors, entry length, and normalized type-token measures.
  3. Populate the sentence-grammar tables over a coverage-balanced sample, then test morphosyntactic vectors separately from formula vectors.
  4. Expand the non-epitomised control set beyond the current Parisinus rows before making any claim about epitomiser layers.
  5. Validate clusters by close reading: each cluster needs formula examples and counterexamples before it becomes an argument.