Stylometric Fingerprinting

Last generated: 2026-08-24 22:06:24 AEST. Detector: translation_guidance_scan_v4.

This page tracks exploratory stylometric work for Stephanos, especially whether formula, gloss, grammar, and source-form features can expose different epitomising layers. The first implemented feature family is the existing translation-guidance recogniser output: each entry gets a vector of rule hits and non-hits, then the vectors are embedded and clustered.

3,571lemmas with at least one active guidance scan
3,571lemmas in the broad formula-vector UMAP slice
321Kappa entries in that UMAP slice
19Parisinus/non-epitomised entries in that UMAP slice
formula: 32, gloss: 70, proper_noun: 75active recogniser rules by kind
0.091KMeans silhouette on broad formula vectors
Current interpretation: the broad formula-vector slice shows an apparent Kappa signal, but the complete-formula slice does not. That means the first visible separation is probably mixed with scan-history and rule-coverage effects. Treat this as a progress marker, not as a demonstrated epitomiser fingerprint.

UMAP

The plotted vectors use formula recogniser rows only, with occurrence counts capped at 5 and transformed by log1p. The broad slice requires at least 21 of the 32 active formula rules to have been checked. Embedding method: UMAP.

Cluster Summary

ClusterRowsKappaParisinusMedian entryMean matched rulesTop over-represented formulae
61,06879 (7.4%)31043.7τὸ ἐθνικὸν + X (nominative ETHNONYM) (+0.64); X (SETTLEMENT) + Y (genitive REGION) (+0.04); X (SETTLEMENT) + Y (genitive PEOPLE) (+0.02); X... πλησίον + Y. (genitive) (+0.01)
896983 (8.6%)2982.6ὁ πολίτης + X (nominative GENTILIC) (+0.09); εἰς + «X» (GREEK LETTER) (+0.00); διὰ τοῦ + «X» (GREEK LETTER) + [FORM OF γράφειν] (+0.00)
243237 (8.6%)11086.7X (nominative PROPER NOUN) ... + ἀπό + Y (genitive ETYMON) (+0.85); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.55); X (nominative PROPER NOUN)... + ἀπό + Y (genitive ARTICLE + genitive ETYMON) (+0.51); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.31)
535866 (18.4%)2849.0X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.69); ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.67); ὡς + X (AUTHOR NAME) (+0.49); ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) (+0.43)
126916 (5.9%)01026.4X (nominative PROPER NOUN) + X (genitive PROPER NOUN) (+0.96); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.21); ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) (+0.17); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.17)
421911 (5.0%)31007.2X (AUTHOR NAME) + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.92); Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.71); X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) (+0.68); X (AUTHOR NAME) + Y (NUMERAL) (+0.13)
01999 (4.5%)6926.7Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) (+0.93); X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) (+0.78); X (AUTHOR NAME) + Y (NUMERAL) (+0.07); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.07)
32711 (40.7%)19715.4X (MASCULINE ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.96); [NO ANTECEDENT] + ἀφ' οὗ + Y (DERIVED NOUN) (+0.96); X (NEUTER ETYMON) + ἀφ' οὗ + Y (DERIVED NOUN) (+0.92); X (genitive ETYMON) + Y (nominative DERIVED NOUN) (+0.54)
9269 (34.6%)112412.8Y (EPITHET) + X (DEITY) (+1.00); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.36); ὡς + X (AUTHOR NAME) (+0.23); X (AUTHOR NAME) + Y (NUMERAL) (+0.19)
740 (0.0%)01516.2[definite article] + εἰς + «X» (GREEK LETTER) + [FORM OF ληγών AGREEING WITH ART] (+1.00); X (nominative) + ὡς + Y (nominative HOMOMORPH) (+0.44); X (SETTLEMENT) + Y (genitive PEOPLE) (+0.36); X (nominative DERIVED NOUN) + Y (nominative ETYMON) (+0.29)

Separation Checks

CheckRowsPositive classResultInterpretation
Kappa vs non-Kappa, broad formula slice 3,571 321 Kappa 0.698 balanced accuracy across folds [0.678, 0.655, 0.719, 0.728, 0.711] Useful as a warning flag, but confounded by recogniser coverage and rule history.
Kappa vs non-Kappa, complete formula slice 987 321 Kappa 0.481 balanced accuracy across folds [0.502, 0.516, 0.458, 0.489, 0.438] The near-baseline result is evidence against claiming a robust Kappa fingerprint from current formula hits alone.
Parisinus/non-epitomised comparison 987 complete formula rows 17 Parisinus descriptive only The current non-epitomised sample is too small for a serious classifier; use it as a qualitative control set.

Coverage

Corpus Versions

VersionRowsUsableWith Greek text
epitome3,6643,6643,551
parisinus191919

Guidance Scan Counts

KindScan rowsLemmas checkedRules checkedLemmas matchedOccurrences
formula111,5623,571323,41020,377
gloss80,9303,571701,4013,300
proper_noun5715712

Formula Coverage Distribution

Formula rules checkedLemmas
212,584
32987

Sentence Grammar Feature Tables

TableRows
sentence_grammar_runs8
sentence_grammar_evaluations111
sentence_grammar_tokens495

Next Implementation Steps

  1. Freeze a coverage-balanced formula/gloss feature matrix and rerun the UMAP after every nightly guidance scan.
  2. Add non-recogniser stylometric baselines: character n-grams, function-word rates, particles, clause connectors, entry length, and normalized type-token measures.
  3. Populate the sentence-grammar tables over a coverage-balanced sample, then test morphosyntactic vectors separately from formula vectors.
  4. Expand the non-epitomised control set beyond the current Parisinus rows before making any claim about epitomiser layers.
  5. Validate clusters by close reading: each cluster needs formula examples and counterexamples before it becomes an argument.