Translation Quality Predictor

This page tests whether current source vocabulary, translation-guidance recogniser matches, and mean v3 translation length can predict which passages in 100 Kappa rows from Gabe's final review tracker export are translated badly by ordinary gpt-5.5 v3. It excludes separate reasoning and repeatability experiment lanes.

Best current model: Greek vocabulary + translation length predicting 3-gram F1 badness, CV R^2 0.376, Spearman r 0.615.

Sentence-level alignment and metric rows exist; the worst-sentence review queue below uses the corrected v3 similarity-DP alignment.

Metric Engines

MetricStatus
BLEU-4SacreBLEU sentence BLEU-4
METEORNLTK METEOR with WordNet synonyms
ROUGE-Lrouge-score ROUGE-L with stemming
chrF++SacreBLEU chrF++ with word_order=2
Ground-truth passages100
Scored v3 passages100
Completed v3 runs118
Mean runs per passage1.18

Predictability By Metric

Targets are badness measures: for score metrics, larger means lower translation score; for length, larger means more absolute word-count error. Cross-validation uses fixed five-fold splits where possible. The sample is small, so negative R^2 values should be read as evidence that the feature family is not currently useful for that metric.

Feature family Target Status Passages Features CV R^2 Spearman r CV MAE MAE lift Worst-quartile precision Ridge alpha
Greek vocabulary + translation length 3-gram F1 badness ok 100 186 0.376 0.615 0.1443 0.0378 60.0% 0.3257
Greek vocabulary 3-gram F1 badness ok 100 185 0.367 0.589 0.1452 0.0369 60.0% 0.3257
Greek vocabulary + translation length 2-gram F1 badness ok 100 186 0.365 0.626 0.1104 0.0335 60.0% 0.3257
Greek vocabulary 2-gram F1 badness ok 100 185 0.361 0.585 0.1101 0.0338 56.0% 0.3257
Greek vocabulary + translation length 3-gram Jaccard badness ok 100 186 0.335 0.593 0.1580 0.0329 64.0% 0.4924
Greek vocabulary 3-gram Jaccard badness ok 100 185 0.327 0.563 0.1564 0.0345 60.0% 0.3257
Greek vocabulary + translation length BLEU-4 badness ok 100 186 0.326 0.642 0.1279 0.0355 64.0% 0.4924
Greek vocabulary + translation length Sentence BLEU badness ok 100 186 0.326 0.642 0.1279 0.0355 64.0% 0.4924
Greek vocabulary BLEU-4 badness ok 100 185 0.321 0.600 0.1265 0.0368 68.0% 0.4924
Greek vocabulary Sentence BLEU badness ok 100 185 0.321 0.600 0.1265 0.0368 68.0% 0.4924
Greek vocabulary + translation length ROUGE-L badness ok 100 186 0.319 0.632 0.0722 0.0205 60.0% 0.4924
Greek vocabulary + translation length chrF++ badness ok 100 186 0.315 0.649 0.0714 0.0192 56.0% 0.7444
Greek vocabulary + translation length METEOR badness ok 100 186 0.314 0.630 0.0785 0.0190 64.0% 0.4924
Greek vocabulary METEOR badness ok 100 185 0.313 0.579 0.0775 0.0200 56.0% 0.4924
Greek vocabulary ROUGE-L badness ok 100 185 0.308 0.586 0.0731 0.0195 56.0% 0.4924
Greek vocabulary chrF++ badness ok 100 185 0.302 0.588 0.0729 0.0178 52.0% 0.4924
Translation length chrF++ badness ok 100 1 0.151 0.406 0.0825 0.0081 44.0% 20.3092
Translation length ROUGE-L badness ok 100 1 0.124 0.364 0.0856 0.0070 44.0% 20.3092
Translation length METEOR badness ok 100 1 0.105 0.305 0.0911 0.0064 40.0% 20.3092
Vocabulary + recognisers + translation length ROUGE-L badness ok 100 267 0.094 0.539 0.0837 0.0089 56.0% 2.5719
Translation length BLEU-4 badness ok 100 1 0.091 0.333 0.1531 0.0103 34.6% 20.3092
Translation length Sentence BLEU badness ok 100 1 0.091 0.333 0.1531 0.0103 34.6% 20.3092
Vocabulary + recognisers + translation length chrF++ badness ok 100 267 0.074 0.585 0.0824 0.0082 48.0% 1.1253
Translation length 3-gram Jaccard badness ok 100 1 0.066 0.227 0.1807 0.0103 36.0% 13.4340
Vocabulary + recognisers + translation length 2-gram F1 badness ok 100 267 0.064 0.493 0.1344 0.0095 60.0% 3.8882
Translation length 2-gram F1 badness ok 100 1 0.062 0.260 0.1372 0.0066 36.0% 20.3092
Vocabulary + recognisers + translation length 3-gram Jaccard badness ok 100 267 0.057 0.409 0.1836 0.0073 48.0% 5.8780
Vocabulary + recognisers + translation length 3-gram F1 badness ok 100 267 0.056 0.452 0.1721 0.0100 56.0% 5.8780
Recogniser rules + translation length chrF++ badness ok 100 82 0.055 0.273 0.0877 0.0029 40.0% 4375.4794
Vocabulary + recognisers chrF++ badness ok 100 266 0.053 0.266 0.0878 0.0028 40.0% 4375.4794
Recogniser rules chrF++ badness ok 100 81 0.052 0.265 0.0879 0.0028 40.0% 4375.4794
Recogniser rules + translation length ROUGE-L badness ok 100 82 0.049 0.318 0.0899 0.0027 40.0% 4375.4794
Vocabulary + recognisers ROUGE-L badness ok 100 266 0.047 0.312 0.0900 0.0026 44.0% 4375.4794
Vocabulary + recognisers + translation length METEOR badness ok 100 267 0.047 0.251 0.0944 0.0031 44.0% 4375.4794
Translation length 3-gram F1 badness ok 100 1 0.047 0.232 0.1743 0.0078 36.0% 13.4340
Greek vocabulary Absolute length percent error ok 100 185 0.047 0.071 0.0428 -0.0002 28.0% 2.5719
Recogniser rules ROUGE-L badness ok 100 81 0.046 0.312 0.0900 0.0026 44.0% 4375.4794
Recogniser rules + translation length METEOR badness ok 100 82 0.046 0.249 0.0944 0.0031 44.0% 4375.4794
Vocabulary + recognisers METEOR badness ok 100 266 0.045 0.249 0.0945 0.0030 44.0% 4375.4794
Recogniser rules METEOR badness ok 100 81 0.045 0.251 0.0945 0.0030 44.0% 4375.4794
Greek vocabulary + translation length Absolute length percent error ok 100 186 0.043 0.070 0.0428 -0.0002 28.0% 2.5719
Vocabulary + recognisers + translation length BLEU-4 badness ok 100 267 0.035 0.239 0.1599 0.0034 40.0% 4375.4794
Vocabulary + recognisers + translation length Sentence BLEU badness ok 100 267 0.035 0.239 0.1599 0.0034 40.0% 4375.4794
Recogniser rules + translation length BLEU-4 badness ok 100 82 0.034 0.238 0.1599 0.0034 40.0% 4375.4794
Recogniser rules + translation length Sentence BLEU badness ok 100 82 0.034 0.238 0.1599 0.0034 40.0% 4375.4794
Vocabulary + recognisers BLEU-4 badness ok 100 266 0.033 0.238 0.1601 0.0032 36.0% 4375.4794
Vocabulary + recognisers Sentence BLEU badness ok 100 266 0.033 0.238 0.1601 0.0032 36.0% 4375.4794
Recogniser rules BLEU-4 badness ok 100 81 0.033 0.234 0.1601 0.0032 36.0% 4375.4794
Recogniser rules Sentence BLEU badness ok 100 81 0.033 0.234 0.1601 0.0032 36.0% 4375.4794
Recogniser rules + translation length 2-gram F1 badness ok 100 82 0.027 0.214 0.1424 0.0015 36.0% 4375.4794
Vocabulary + recognisers 2-gram F1 badness ok 100 266 0.025 0.206 0.1425 0.0013 36.0% 4375.4794
Recogniser rules 2-gram F1 badness ok 100 81 0.025 0.202 0.1425 0.0013 36.0% 4375.4794
Recogniser rules + translation length 3-gram F1 badness ok 100 82 0.020 0.234 0.1805 0.0017 40.0% 2894.2661
Vocabulary + recognisers 3-gram F1 badness ok 100 266 0.019 0.223 0.1806 0.0015 40.0% 2894.2661
Recogniser rules 3-gram F1 badness ok 100 81 0.019 0.221 0.1807 0.0014 40.0% 2894.2661
Recogniser rules + translation length 3-gram Jaccard badness ok 100 82 0.016 0.206 0.1896 0.0013 36.0% 2894.2661
Vocabulary + recognisers Absolute length percent error ok 100 266 0.015 0.114 0.0432 -0.0006 44.0% 1266.3802
Recogniser rules Absolute length percent error ok 100 81 0.015 0.111 0.0432 -0.0006 44.0% 1266.3802
Vocabulary + recognisers + translation length Absolute length percent error ok 100 267 0.015 0.111 0.0432 -0.0006 44.0% 1266.3802
Recogniser rules + translation length Absolute length percent error ok 100 82 0.015 0.110 0.0432 -0.0006 44.0% 1266.3802
Recogniser rules 3-gram Jaccard badness ok 100 81 0.015 0.194 0.1898 0.0011 36.0% 2894.2661
Vocabulary + recognisers 3-gram Jaccard badness ok 100 266 0.013 0.398 0.1846 0.0063 48.0% 5.8780
Translation length Absolute length percent error ok 100 1 -0.005 -0.146 0.0428 -0.0002 20.0% 10000.0000

Highest Predicted Risk

This list uses the best cross-validated model in this run and sorts passages by predicted badness for 3-gram F1 badness.

Lemma ID v3 runs Source words Observed badness Predicted badness BLEU-4 chrF++ 3-gram F1 Length error
Καρία 2484 1 181.0 0.5777 0.8025 42.0% 65.8% 42.2% 12.6%
Κάλυτις 2335 2 16.0 0.8531 0.7756 31.9% 60.3% 14.7% 5.8%
Κάρυστος 2603 1 132.0 0.5850 0.7618 46.5% 69.5% 41.5% 0.6%
Καλάσιρις 2085 1 10.0 0.9167 0.6965 32.3% 69.5% 8.3% 15.4%
Καταονία 2628 1 17.0 0.6889 0.6889 36.2% 67.4% 31.1% 4.2%
Κασώριον 2623 1 14.0 0.6471 0.6836 32.1% 66.0% 35.3% 10.0%
Καλαβρία 2080 1 12.0 0.8667 0.6482 23.6% 66.7% 13.3% 11.1%
Καρχηδών 2604 1 88.0 0.6432 0.6351 35.0% 63.9% 35.7% 9.4%
Κεκρυφάλεια 3258 1 17.0 0.2340 0.6252 77.1% 89.4% 76.6% 3.8%
Καππαδοκία 2470 3 57.0 0.3886 0.6138 46.6% 71.2% 61.1% 2.5%
Κριώα 3530 1 16.0 0.6889 0.6086 38.0% 71.1% 31.1% 11.5%
Καικῖνον 2074 1 6.0 1.0000 0.6040 21.4% 66.4% 0.0% 9.1%
Κάναστρον 2455 1 43.0 0.7345 0.6024 32.0% 62.9% 26.5% 5.0%
Κύτα 7254 1 58.0 0.4881 0.5937 54.5% 74.4% 51.2% 4.8%
Καπετώλιον 2468 1 86.0 0.6905 0.5912 30.1% 50.3% 31.0% 14.5%
Καδμεία 2059 1 17.0 0.6316 0.5870 41.9% 68.2% 36.8% 10.0%
Κώμη 7266 1 53.0 0.8630 0.5798 13.8% 47.6% 13.7% 10.1%
Κύρη 7243 1 14.0 0.2821 0.5643 82.2% 91.2% 71.8% 4.5%
Κάλλατις 2119 1 45.0 0.6947 0.5547 31.8% 63.9% 30.5% 7.7%
Κυρτώνιος 7253 1 14.0 0.7222 0.5460 32.9% 64.6% 27.8% 22.2%
Καρπασία 2597 1 80.0 0.4836 0.5383 54.3% 75.2% 51.6% 6.2%
Κυτέριον 7255 1 18.0 0.5652 0.5380 53.0% 76.1% 43.5% 27.3%
Καβασσός 2055 1 69.0 0.6731 0.5334 41.6% 67.5% 32.7% 5.5%
Κάσος 2607 1 44.0 0.5636 0.5330 33.7% 61.5% 43.6% 0.0%
Κύρνος 7247 1 34.0 0.4839 0.5285 47.1% 69.3% 51.6% 6.0%
Κατάνη 2626 1 64.0 0.6333 0.5261 43.8% 64.8% 36.7% 6.3%
Κάλπη 2329 1 36.0 0.2778 0.5235 81.0% 89.1% 72.2% 3.5%
Κάληρος 2116 1 22.0 0.7183 0.5189 29.3% 61.2% 28.2% 12.5%
Κύρτος 7251 1 48.0 0.4677 0.5057 50.8% 68.9% 53.2% 6.1%
Κάθαια 2062 1 16.0 0.8750 0.5047 24.6% 59.2% 12.5% 8.0%

Predictive Features

Translation Length: chrF++ badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
translation_length Mean v3 translation word count z-scored; mean 43.3, SD 36.9 words; present/high = top quartile, absent/low = bottom quartile 0.04045 100 0.2964 0.1529

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent

Vocabulary Terms: 3-gram F1 badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary χωριον 0.35207 2 0.8117 0.4386
vocabulary οικητωρ 0.28127 6 0.7032 0.4297
vocabulary τοις 0.24860 3 0.6520 0.4397
vocabulary τον 0.22986 10 0.6189 0.4269
vocabulary επι 0.22351 2 0.7827 0.4392
vocabulary δευτερω 0.22234 3 0.7329 0.4372
vocabulary οικητωρ θηλυκον 0.22188 2 0.8599 0.4377
vocabulary καλειται 0.22177 2 0.6239 0.4425
vocabulary ωστε 0.20417 3 0.7168 0.4377
vocabulary μοιρα 0.20315 2 0.8028 0.4388
vocabulary τε 0.17563 5 0.5976 0.4381
vocabulary ει 0.16786 4 0.7175 0.4348

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary τεταρτω -0.39870 3 0.1111 0.4565
vocabulary εβδομη -0.31472 2 0.3473 0.4481
vocabulary μεταξυ -0.29994 5 0.3958 0.4487
vocabulary ως εθνικον -0.27046 5 0.2440 0.4567
vocabulary πορρω -0.21394 3 0.1302 0.4559
vocabulary ου πορρω -0.21394 3 0.1302 0.4559
vocabulary εν εθνικον -0.19579 3 0.3009 0.4506
vocabulary ακρα -0.19212 4 0.3969 0.4481
vocabulary εθνικον -0.18558 56 0.4025 0.5015
vocabulary προς -0.18136 15 0.4458 0.4462
vocabulary αι -0.17964 3 0.4240 0.4468
vocabulary ως εν -0.17823 7 0.3972 0.4498

Vocabulary Terms + Translation Length: 3-gram F1 badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary χωριον 0.35739 2 0.8117 0.4386
vocabulary οικητωρ 0.29124 6 0.7032 0.4297
vocabulary τοις 0.23770 3 0.6520 0.4397
vocabulary δευτερω 0.23299 3 0.7329 0.4372
vocabulary οικητωρ θηλυκον 0.23155 2 0.8599 0.4377
vocabulary επι 0.22691 2 0.7827 0.4392
vocabulary καλειται 0.21829 2 0.6239 0.4425
vocabulary μοιρα 0.21622 2 0.8028 0.4388
vocabulary τον 0.21256 10 0.6189 0.4269
vocabulary ωστε 0.18638 3 0.7168 0.4377
vocabulary αυτου 0.14098 2 0.6339 0.4423
vocabulary τε 0.13815 5 0.5976 0.4381

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary τεταρτω -0.38349 3 0.1111 0.4565
vocabulary εβδομη -0.31068 2 0.3473 0.4481
vocabulary μεταξυ -0.29797 5 0.3958 0.4487
vocabulary ως εθνικον -0.26418 5 0.2440 0.4567
vocabulary ως εν -0.20637 7 0.3972 0.4498
vocabulary πορρω -0.20508 3 0.1302 0.4559
vocabulary ου πορρω -0.20508 3 0.1302 0.4559
vocabulary ακρα -0.20006 4 0.3969 0.4481
vocabulary εν -0.19916 37 0.4540 0.4414
vocabulary εστι -0.19439 24 0.4651 0.4401
vocabulary εν εθνικον -0.18880 3 0.3009 0.4506
vocabulary αι -0.18250 3 0.4240 0.4468

Recogniser Rules: chrF++ badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary gloss occurrence count 0.00166 95 0.2299 0.1542
recogniser_summary gloss rule count 0.00128 95 0.2299 0.1542
recogniser_summary matched occurrence count 0.00094 100 0.2261 N/A
recogniser_summary matched rule count 0.00053 100 0.2261 N/A
recogniser_rule gloss: καλεῖται/ἐκαλεῖτο/κέκληται/ἐκλήθη (ἀπο...) is/used to be/is/was called/named after (+ ἀπο) 0.00037 14 0.3102 0.2124
recogniser_rule formula: X (SETTLEMENT) + Y (genitive REGION) Translate as "a X in Y" 0.00028 63 0.2432 0.1970
recogniser_rule formula: X (nominative PROPER NOUN) + X (genitive PROPER NOUN) Translate as "X, X" 0.00026 20 0.2749 0.2139
recogniser_rule gloss: πόλισμα τό * town 0.00015 5 0.3127 0.2216
recogniser_rule formula: X (nominative DERIVED NOUN) + Y (nominative ETYMON) Translate as "'X' is from Y" 0.00013 35 0.2359 0.2209
recogniser_rule gloss: κώμη ἡ village 0.00012 3 0.3208 0.2232
recogniser_rule gloss: χώρα ἡ region; territory (when belonging to a specific people, tribe, or nation); surrounding territory (when juxtaposed against a πόλις); land (remote or poetical places) 0.00012 12 0.3027 0.2157
recogniser_rule gloss: οἰκήτωρ ὁ inhabitant, resident, patron (of a brothel...? - κ123) 0.00012 6 0.3454 0.2185

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary formula rule count -0.00078 99 0.2275 0.0919
recogniser_summary formula occurrence count -0.00075 99 0.2275 0.0919
recogniser_rule formula: ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) Translate as "(just) as 'Y' is from X" -0.00027 34 0.2019 0.2386
recogniser_rule gloss: ἔθνος τό people -0.00026 16 0.1650 0.2378
recogniser_rule formula: X (AUTHOR NAME) + Y (NUMERAL) Translate as "X, book Y" -0.00023 30 0.1892 0.2419
recogniser_rule formula: ὡς + X (AUTHOR NAME) Translate as "as per X" -0.00021 40 0.2121 0.2355
recogniser_rule formula: X (nominative) + ὡς + Y (nominative HOMOMORPH) Translate as "'X' as in 'Y'" -0.00021 42 0.2141 0.2348
recogniser_rule formula: τὸ ἐθνικὸν + X (nominative ETHNONYM) Translate as "the ethnonym is 'X'" -0.00020 61 0.2158 0.2423
recogniser_rule formula: X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) Translate as "Χ, in his *Y*" -0.00020 17 0.1958 0.2323
recogniser_rule formula: Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) Translate as "X, in book Y of his *Z*" -0.00011 11 0.2186 0.2271
recogniser_rule gloss: ἐθνικόν τό ethnonym -0.00008 60 0.2167 0.2403
recogniser_rule formula: ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) Translate as "just as 'Y' is from the name X" -0.00008 19 0.2141 0.2289

Recogniser Rules + Translation Length: chrF++ badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary gloss occurrence count 0.00164 95 0.2299 0.1542
recogniser_summary gloss rule count 0.00127 95 0.2299 0.1542
recogniser_summary matched occurrence count 0.00092 100 0.2261 N/A
translation_length Mean v3 translation word count z-scored; mean 43.3, SD 36.9 words; present/high = top quartile, absent/low = bottom quartile 0.00071 100 0.2964 0.1529
recogniser_summary matched rule count 0.00052 100 0.2261 N/A
recogniser_rule gloss: καλεῖται/ἐκαλεῖτο/κέκληται/ἐκλήθη (ἀπο...) is/used to be/is/was called/named after (+ ἀπο) 0.00036 14 0.3102 0.2124
recogniser_rule formula: X (SETTLEMENT) + Y (genitive REGION) Translate as "a X in Y" 0.00028 63 0.2432 0.1970
recogniser_rule formula: X (nominative PROPER NOUN) + X (genitive PROPER NOUN) Translate as "X, X" 0.00026 20 0.2749 0.2139
recogniser_rule gloss: πόλισμα τό * town 0.00015 5 0.3127 0.2216
recogniser_rule formula: X (nominative DERIVED NOUN) + Y (nominative ETYMON) Translate as "'X' is from Y" 0.00013 35 0.2359 0.2209
recogniser_rule gloss: κώμη ἡ village 0.00012 3 0.3208 0.2232
recogniser_rule gloss: χώρα ἡ region; territory (when belonging to a specific people, tribe, or nation); surrounding territory (when juxtaposed against a πόλις); land (remote or poetical places) 0.00012 12 0.3027 0.2157

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary formula rule count -0.00078 99 0.2275 0.0919
recogniser_summary formula occurrence count -0.00075 99 0.2275 0.0919
recogniser_rule formula: ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) Translate as "(just) as 'Y' is from X" -0.00027 34 0.2019 0.2386
recogniser_rule gloss: ἔθνος τό people -0.00026 16 0.1650 0.2378
recogniser_rule formula: X (AUTHOR NAME) + Y (NUMERAL) Translate as "X, book Y" -0.00023 30 0.1892 0.2419
recogniser_rule formula: ὡς + X (AUTHOR NAME) Translate as "as per X" -0.00021 40 0.2121 0.2355
recogniser_rule formula: X (nominative) + ὡς + Y (nominative HOMOMORPH) Translate as "'X' as in 'Y'" -0.00021 42 0.2141 0.2348
recogniser_rule formula: τὸ ἐθνικὸν + X (nominative ETHNONYM) Translate as "the ethnonym is 'X'" -0.00020 61 0.2158 0.2423
recogniser_rule formula: X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) Translate as "Χ, in his *Y*" -0.00020 17 0.1958 0.2323
recogniser_rule formula: Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) Translate as "X, in book Y of his *Z*" -0.00011 11 0.2186 0.2271
recogniser_rule gloss: ἐθνικόν τό ethnonym -0.00008 60 0.2167 0.2403
recogniser_rule formula: ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) Translate as "just as 'Y' is from the name X" -0.00008 19 0.2141 0.2289

Combined Model Features: chrF++ badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary gloss occurrence count 0.00166 95 0.2299 0.1542
recogniser_summary gloss rule count 0.00128 95 0.2299 0.1542
recogniser_summary matched occurrence count 0.00094 100 0.2261 N/A
recogniser_summary matched rule count 0.00054 100 0.2261 N/A
recogniser_rule gloss: καλεῖται/ἐκαλεῖτο/κέκληται/ἐκλήθη (ἀπο...) is/used to be/is/was called/named after (+ ἀπο) 0.00036 14 0.3102 0.2124
recogniser_rule formula: X (SETTLEMENT) + Y (genitive REGION) Translate as "a X in Y" 0.00028 63 0.2432 0.1970
recogniser_rule formula: X (nominative PROPER NOUN) + X (genitive PROPER NOUN) Translate as "X, X" 0.00026 20 0.2749 0.2139
recogniser_rule gloss: πόλισμα τό * town 0.00015 5 0.3127 0.2216
recogniser_rule formula: X (nominative DERIVED NOUN) + Y (nominative ETYMON) Translate as "'X' is from Y" 0.00013 35 0.2359 0.2209
recogniser_rule gloss: κώμη ἡ village 0.00012 3 0.3208 0.2232
recogniser_rule gloss: χώρα ἡ region; territory (when belonging to a specific people, tribe, or nation); surrounding territory (when juxtaposed against a πόλις); land (remote or poetical places) 0.00012 12 0.3027 0.2157
recogniser_rule gloss: οἰκήτωρ ὁ inhabitant, resident, patron (of a brothel...? - κ123) 0.00012 6 0.3454 0.2185

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary formula rule count -0.00078 99 0.2275 0.0919
recogniser_summary formula occurrence count -0.00075 99 0.2275 0.0919
recogniser_rule formula: ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) Translate as "(just) as 'Y' is from X" -0.00027 34 0.2019 0.2386
recogniser_rule gloss: ἔθνος τό people -0.00026 16 0.1650 0.2378
recogniser_rule formula: X (AUTHOR NAME) + Y (NUMERAL) Translate as "X, book Y" -0.00023 30 0.1892 0.2419
recogniser_rule formula: ὡς + X (AUTHOR NAME) Translate as "as per X" -0.00021 40 0.2121 0.2355
recogniser_rule formula: X (nominative) + ὡς + Y (nominative HOMOMORPH) Translate as "'X' as in 'Y'" -0.00021 42 0.2141 0.2348
recogniser_rule formula: τὸ ἐθνικὸν + X (nominative ETHNONYM) Translate as "the ethnonym is 'X'" -0.00020 61 0.2158 0.2423
recogniser_rule formula: X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) Translate as "Χ, in his *Y*" -0.00020 17 0.1958 0.2323
recogniser_rule formula: Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) Translate as "X, in book Y of his *Z*" -0.00011 11 0.2186 0.2271
vocabulary εθνικον -0.00010 56 0.2116 0.2446
vocabulary ως εθνικον -0.00009 5 0.1420 0.2305

Combined Model Features + Translation Length: ROUGE-L badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_rule formula: X (nominative PROPER NOUN) + X (genitive PROPER NOUN) Translate as "X, X" 0.04917 20 0.2251 0.1770
recogniser_rule gloss: μοῖρα ἡ region or part (in geographic contexts); district (in urban contexts only) 0.04861 4 0.3235 0.1809
translation_length Mean v3 translation word count z-scored; mean 43.3, SD 36.9 words; present/high = top quartile, absent/low = bottom quartile 0.04778 100 0.2545 0.1261
vocabulary τον 0.04740 10 0.2756 0.1767
recogniser_rule formula: καί + X (nominative PROPER NOUN) + Y (nominative PROPER NOUN) Translate as "Y is also 'X'" 0.03953 13 0.2139 0.1825
vocabulary επι 0.03902 2 0.3951 0.1824
recogniser_rule formula: X (AUTHOR NAME) + Y (dative NUMBER) + Z (genitive BOOK NAME) Translate as "X, book Y in his *Z*" 0.03834 28 0.1931 0.1841
recogniser_rule gloss: ἄκρον τό cape (when on the coast, sgl.), headlands (when on the coast, plu.); peak (when inland) 0.03823 8 0.2636 0.1799
recogniser_rule gloss: χωρίον τό locality; point (only in μέσα χωρία: ‘halfway point’) 0.03744 3 0.3365 0.1820
recogniser_rule gloss: κώμη ἡ village 0.03581 3 0.3075 0.1829
vocabulary ωστε 0.03463 3 0.3498 0.1816
vocabulary μοιρα 0.03325 2 0.3520 0.1832

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary τεταρτω -0.05422 3 0.0565 0.1906
recogniser_rule gloss: ἔθνος τό people -0.05140 16 0.1347 0.1965
vocabulary γενος -0.04242 2 0.0880 0.1886
vocabulary μεταξυ -0.03568 5 0.1659 0.1877
vocabulary ως εθνικον -0.02891 5 0.1200 0.1901
recogniser_rule formula: X (AUTHOR NAME) + Y (NUMERAL) Translate as "X, book Y" -0.02754 30 0.1593 0.1983
recogniser_rule formula: X (nominative) + ὡς + Y (nominative HOMOMORPH) Translate as "'X' as in 'Y'" -0.02687 42 0.1726 0.1968
recogniser_rule gloss: μητρόπολις ἡ metropolis -0.02546 5 0.2299 0.1843
recogniser_rule gloss: πολίτης ὁ citizen -0.02443 13 0.1970 0.1851
recogniser_rule gloss: διαίρεσις ἡ diaresis (i.e. the punctuation mark) -0.02405 2 0.2171 0.1860
recogniser_rule formula: X... + πρός + Y (dative) Translate as "X... near Y" -0.02363 12 0.1663 0.1894
recogniser_rule gloss: ὄρος ὁ mountain, mount (X ὄρος = 'Mount X') -0.02282 5 0.2404 0.1838

Worst v3 Translation Sentences

Worst means high average percentile badness across chrF, sentence BLEU, ROUGE-L, 3-gram F1, and absolute word-count delta. This is a review queue, not a human error judgment.

Rank Headword Greek sentence v3 candidate Human-approved translation
1 Καλαμένθη κρεῖττον οὖν ὡς Ἡρόδοτος διὰ τοῦ « ι ». The better form, then, is as per Herodotos, written with ι. It is better to have it with ι, as per Herodotos. A city of the Phoenicians.
2 Κάλλατις ὡς κάλαθος εὑρέθη ἐοικὼς τοῖς θεσμοφοριακοῖς. It is as in 'kalathos', because a basket was found resembling those used at the Thesmophoria. Because a basket similar to that which is ‘Thesmophorian’ was found there.
3 Κατάνη ἀπὸ δὲ τῆς Κατάνης Χαρώνδας, ὁ διάσημος τῶν Ἀθήνησι νομοθετῶν. From Katane came Charondas, renowned among the lawgivers at Athens. From Katane was Charondas, one of the lawmakers made famous by the Athenians.
4 Κοτιάειον δοκεῖ δὲ ἀπὸ τοῦ Κόσα Κόσαιον, ὡς τοῦ Μίδα Μίδαιον, καὶ προσθέσει τοῦ « ι » Κοσιαῖον, καὶ τροπῇ Κοτιαῖον. It seems that 'Kosaion' is from 'Kosa', just as 'Midaion' is from 'Midas'; and by addition of ι, 'Kosiaion', and by alteration, 'Kotiaion'. it does seem that 'Kosaeion' is from the name Kosas (as 'Midaeion' is from the name Midas), then it becomes 'Kosiaeion' by addition of ι and 'Kotiaeion' through sound shift.
5 Κάνωπος καὶ Κανωβίς θηλυκόν. There is also ‘Kanobis’ as a feminine form. The feminine is ‘Kanobis’.
6 Καιρή εἴρηται ἐν τῷ περὶ Ἀγύλλης. It has been stated in the entry on Agylla. It is cited in *On Agylla*.
7 Καρία ἡ χρῆσις δὲ διὰ διφθόγγου κατὰ συνεκδρομὴν τοῦ σώτειρα ὀλέτειρα“. Common usage, however, has it with a diphthong by association with σώτειρα, ὀλέτειρα.' The common usage, however, is with the diphthong, according to the analogy of 'soteira', 'oleteira'.'
8 Καπετώλιον Ἀσκληπιεῖον, Ἀσκληπιός γάρ· Ἀσκληπιεῖον, for there is Ἀσκληπιός; Asklepieion (Ἀσκληπιεῖον) is thus from 'Asklepios' (Ἀσκληπιός);
9 Καπετώλιον Πτολεμαεῖον, Πτολεμαῖος γάρ· Πτολεμαεῖον, for there is Πτολεμαῖος; Ptolemaeion (Πτολεμαεῖον) is thus from 'Ptolemaios' (Πτολεμαῖος);
10 Καπετώλιον Καπετώλιον, ἐν Ῥώμῃ λόφος Ταρπαῖος πάλαι λεγόμενος. Kapetolion: in Rome, a hill formerly called Tarpeios. Kapetolion: a hill in Rome that was long ago called 'Tarpaios'.
11 Κάλυτις ὁ οἰκήτωρ Καλυτίτης, καὶ τὸ θηλυκὸν Καλυτίς, διὰ τὸ προειλῆφθαι τὸν χαρακτῆρα. The inhabitant is 'Kalytites', and the feminine is 'Kalytis', because the characteristic element has already been taken in advance. An inhabitant is a 'Kalytites'; the feminine is also 'Kalytis' due to the form being anticipated.
12 Καλλίπολις δευτέρα κατὰ τὸν Ἀνάπλουν. A second, according to the *Anaplous*. (2) Along the Anaplous.
13 Καρία Ἡρωδιανὸς δὲ ἐν μὲν τῇ Ὀρθογραφίᾳ (2,410,22) ἀμφίβολον αὐτό φησιν. ἐν δὲ τῇ Καθόλου (1,250,14) <τῇ> χρήσει ἑπόμενος διὰ διφθόγγου φησίν, ὑπομνηματίζων δὲ τὸ Περὶ γενῶν Ἀπολλωνίου (2,777,13) διὰ τοῦ ι μακροῦ. „ἔστι γὰρ ὅτε μετὰ τὴν διαίρεσιν ἔκτασις γίγνεται, ὀίομαι ὄιγον ὄιδα παρ’ Αἰολεῦσιν, ἀντὶ τοῦ οἶδα. Herodianos in his *Orthography* is undecided, but, following general usage, says it is with a diphthong; when commenting on Apollonios' *On Genders*, however, he gives it with long ι: 'For there are times when lengthening occurs after separation: ὀίομαι, ὄιγον, ὄιδα among the Aiolians, instead of οἶδα. While Herodian says that this is doubtful in his *Orthography* (and in his *General Prosody* he says that it uses the diphthong following the common usage), he comments on Apollonios’ *On Genders* that it is with long ι: 'for there is occasion when lengthening occurs after diaresis: 'oïomai', 'oïgon', 'oïda' among the Aeolians rather than 'oida'.
14 Κάσιον ὁ πολίτης Κασιώτης ὡς Πηλουσιώτης, καὶ θηλυκὸν Κασιῶτις, καὶ τὸ κτητικὸν Κασιωτικός, ἀφ´ οὗ ἐν τῇ συνηθείᾳ τὰ Κασιωτικὰ ἱμάτια. the feminine is 'Kasiotis', and the possessive is 'Kasiotikos', from which in ordinary usage comes the phrase 'Kasiotika cloaks'. The possessive is 'Kasiotikos', hence the term 'Kasiotic cloaks' in ordinary language.
15 Κάστνιον ἔδει δὲ Καστνιώτης ὡς Πηλιώτης. It ought, however, to be 'Kastniotes', as 'Peliotes' is from Pelion. However, it should be 'Kastniotes' (as in 'Peliotes').
16 Κυρτώνιος τὸ ἐθνικὸν τῷ τῆς χώρας ἔθει Κυρτωνῖνος ὡς Σατορνῖνος. The ethnonym, according to the usage of the region, is 'Kyrtoninos', as in 'Satorninos'. In local usage, the ethnonym is 'Kyrtoninos' (as in 'Saturninos').
17 Κάναι Καναῖος Ζεύς οὐ μόνον ἀπὸ τοῦ Καναίου, ἀλλὰ καὶ ἀπὸ τῆς Κάνης. Kanaios Zeus is named not only after Kanaios, but also after Kane. Zeus Kanaios is not only from the form 'Kanaios', but also from the form 'Kane'.
18 Κάσος ἀπῴκισται δὲ τῆς νήσου καὶ τὸ ἐν Συρίᾳ ὄρος Κάσιον. The mountain Kasion in Syria has also been colonised from the island. Mount Kasios in Syria was also settled from this island.
19 Κορώνεια τετάρτη πόλις Κύπρου. A fourth is a city of Cyprus. (4) a city in Cyprus;
20 Καπετώλιον ὅσα γὰρ ἔχει προϋπάρχοντα εἰς « ος » καθαρόν, παραληγόμενα ἢ μόνῳ τῷ « ι » ἢ προηγουμένου αὐτοῦ τοῦ « α » ὥστε εἶναι πρὸ τέλους τὴν « αι » δίφθογγον, προπερισπᾶται, ἢ καὶ ὅσα κτητικά. For all words which have pre-existing forms ending in pure -ος, and whose penult has either ι alone or this preceded by α, so that the diphthong αι comes before the final syllable, are accented with a circumflex on the penult; so too all possessives. This is because forms whose base already ends in postvocalic -ος—when either a single ι is in the penultimate position or α precedes it so that the diphthong αι stands before the ultima—will be accented with a circumflex on the penult, and the same applies to possessive forms.

Downloadable Tables

Generated: 2026-10-10 10:17:37 UTC. Recogniser detector version: translation_guidance_scan_v4.