Translation Quality Predictor

This page tests whether current source vocabulary, translation-guidance recogniser matches, and mean v3 translation length can predict which passages in 100 Kappa rows from Gabe's final review tracker export are translated badly by ordinary gpt-5.5 v3. It excludes separate reasoning and repeatability experiment lanes.

Best current model: Greek vocabulary + translation length predicting 2-gram F1 badness, CV R^2 0.397, Spearman r 0.645.

Sentence-level alignment and metric rows exist; the worst-sentence review queue below uses the corrected v3 similarity-DP alignment.

Metric Engines

MetricStatus
BLEU-4SacreBLEU sentence BLEU-4
METEORNLTK METEOR with WordNet synonyms
ROUGE-Lrouge-score ROUGE-L with stemming
chrF++SacreBLEU chrF++ with word_order=2
Ground-truth passages100
Scored v3 passages100
Completed v3 runs116
Mean runs per passage1.16

Predictability By Metric

Targets are badness measures: for score metrics, larger means lower translation score; for length, larger means more absolute word-count error. Cross-validation uses fixed five-fold splits where possible. The sample is small, so negative R^2 values should be read as evidence that the feature family is not currently useful for that metric.

Feature family Target Status Passages Features CV R^2 Spearman r CV MAE MAE lift Worst-quartile precision Ridge alpha
Greek vocabulary + translation length 2-gram F1 badness ok 100 207 0.397 0.645 0.1125 0.0339 60.0% 0.3257
Greek vocabulary + translation length 3-gram F1 badness ok 100 207 0.382 0.629 0.1473 0.0388 64.0% 0.3257
Greek vocabulary 2-gram F1 badness ok 100 206 0.371 0.596 0.1145 0.0318 48.0% 0.3257
Greek vocabulary + translation length 3-gram Jaccard badness ok 100 207 0.364 0.625 0.1569 0.0366 60.0% 0.4924
Greek vocabulary + translation length BLEU-4 badness ok 100 207 0.363 0.656 0.1277 0.0386 56.0% 0.4924
Greek vocabulary + translation length Sentence BLEU badness ok 100 207 0.363 0.656 0.1277 0.0386 56.0% 0.4924
Greek vocabulary + translation length chrF++ badness ok 100 207 0.363 0.671 0.0707 0.0220 56.0% 0.4924
Greek vocabulary 3-gram F1 badness ok 100 206 0.355 0.582 0.1495 0.0366 56.0% 0.3257
Greek vocabulary + translation length ROUGE-L badness ok 100 207 0.355 0.668 0.0730 0.0196 60.0% 0.4924
Greek vocabulary 3-gram Jaccard badness ok 100 206 0.340 0.564 0.1578 0.0357 52.0% 0.4924
Greek vocabulary BLEU-4 badness ok 100 206 0.332 0.596 0.1297 0.0366 60.0% 0.4924
Greek vocabulary Sentence BLEU badness ok 100 206 0.332 0.596 0.1297 0.0366 60.0% 0.4924
Greek vocabulary chrF++ badness ok 100 206 0.325 0.594 0.0736 0.0191 56.0% 0.4924
Greek vocabulary ROUGE-L badness ok 100 206 0.321 0.598 0.0741 0.0184 48.0% 0.3257
Greek vocabulary + translation length METEOR badness ok 100 207 0.313 0.631 0.0809 0.0177 64.0% 0.7444
Greek vocabulary METEOR badness ok 100 206 0.291 0.559 0.0812 0.0174 48.0% 0.4924
Translation length chrF++ badness ok 100 1 0.160 0.434 0.0844 0.0083 44.0% 13.4340
Translation length ROUGE-L badness ok 100 1 0.128 0.380 0.0854 0.0072 44.0% 20.3092
Translation length METEOR badness ok 100 1 0.111 0.319 0.0920 0.0065 44.0% 20.3092
Translation length BLEU-4 badness ok 100 1 0.099 0.348 0.1544 0.0118 30.8% 20.3092
Translation length Sentence BLEU badness ok 100 1 0.099 0.348 0.1544 0.0118 30.8% 20.3092
Translation length 3-gram Jaccard badness ok 100 1 0.074 0.246 0.1819 0.0117 32.0% 13.4340
Translation length 2-gram F1 badness ok 100 1 0.069 0.272 0.1379 0.0085 40.0% 20.3092
Vocabulary + recognisers + translation length 3-gram Jaccard badness ok 100 288 0.064 0.440 0.1827 0.0108 52.0% 5.8780
Vocabulary + recognisers + translation length chrF++ badness ok 100 288 0.063 0.377 0.0895 0.0033 44.0% 2894.2661
Recogniser rules + translation length chrF++ badness ok 100 82 0.062 0.377 0.0895 0.0032 44.0% 2894.2661
Vocabulary + recognisers + translation length ROUGE-L badness ok 100 288 0.060 0.519 0.0848 0.0078 52.0% 2.5719
Vocabulary + recognisers chrF++ badness ok 100 287 0.059 0.368 0.0897 0.0031 48.0% 2894.2661
Recogniser rules chrF++ badness ok 100 81 0.059 0.367 0.0897 0.0030 44.0% 2894.2661
Recogniser rules + translation length ROUGE-L badness ok 100 82 0.057 0.366 0.0892 0.0033 44.0% 4375.4794
Vocabulary + recognisers + translation length METEOR badness ok 100 288 0.056 0.319 0.0954 0.0031 52.0% 4375.4794
Recogniser rules + translation length METEOR badness ok 100 82 0.055 0.318 0.0954 0.0031 52.0% 4375.4794
Vocabulary + recognisers ROUGE-L badness ok 100 287 0.055 0.357 0.0894 0.0032 44.0% 4375.4794
Recogniser rules ROUGE-L badness ok 100 81 0.054 0.355 0.0894 0.0032 44.0% 4375.4794
Vocabulary + recognisers METEOR badness ok 100 287 0.054 0.315 0.0955 0.0030 52.0% 4375.4794
Recogniser rules METEOR badness ok 100 81 0.054 0.314 0.0956 0.0030 52.0% 4375.4794
Translation length 3-gram F1 badness ok 100 1 0.052 0.239 0.1762 0.0099 32.0% 13.4340
Greek vocabulary Absolute length percent error ok 100 206 0.048 0.070 0.0419 0.0002 28.0% 2.5719
Greek vocabulary + translation length Absolute length percent error ok 100 207 0.044 0.061 0.0420 0.0001 28.0% 2.5719
Recogniser rules + translation length BLEU-4 badness ok 100 82 0.041 0.276 0.1624 0.0038 40.0% 4375.4794
Recogniser rules + translation length Sentence BLEU badness ok 100 82 0.041 0.276 0.1624 0.0038 40.0% 4375.4794
Vocabulary + recognisers BLEU-4 badness ok 100 287 0.040 0.272 0.1626 0.0037 40.0% 4375.4794
Vocabulary + recognisers Sentence BLEU badness ok 100 287 0.040 0.272 0.1626 0.0037 40.0% 4375.4794
Vocabulary + recognisers + translation length BLEU-4 badness ok 100 288 0.040 0.532 0.1519 0.0143 52.0% 3.8882
Vocabulary + recognisers + translation length Sentence BLEU badness ok 100 288 0.040 0.532 0.1519 0.0143 52.0% 3.8882
Recogniser rules BLEU-4 badness ok 100 81 0.040 0.271 0.1626 0.0036 40.0% 4375.4794
Recogniser rules Sentence BLEU badness ok 100 81 0.040 0.271 0.1626 0.0036 40.0% 4375.4794
Vocabulary + recognisers + translation length 2-gram F1 badness ok 100 288 0.035 0.271 0.1431 0.0033 40.0% 4375.4794
Recogniser rules + translation length 2-gram F1 badness ok 100 82 0.035 0.268 0.1431 0.0033 40.0% 4375.4794
Vocabulary + recognisers 2-gram F1 badness ok 100 287 0.033 0.264 0.1432 0.0031 40.0% 4375.4794
Recogniser rules 2-gram F1 badness ok 100 81 0.033 0.262 0.1432 0.0031 40.0% 4375.4794
Vocabulary + recognisers + translation length 3-gram F1 badness ok 100 288 0.027 0.263 0.1819 0.0042 36.0% 2894.2661
Recogniser rules + translation length 3-gram F1 badness ok 100 82 0.027 0.263 0.1819 0.0042 36.0% 2894.2661
Vocabulary + recognisers 3-gram F1 badness ok 100 287 0.026 0.258 0.1821 0.0040 36.0% 2894.2661
Recogniser rules 3-gram F1 badness ok 100 81 0.025 0.256 0.1821 0.0040 36.0% 2894.2661
Vocabulary + recognisers Absolute length percent error ok 100 287 0.024 0.188 0.0426 -0.0005 48.0% 1266.3802
Recogniser rules Absolute length percent error ok 100 81 0.024 0.188 0.0426 -0.0005 48.0% 1266.3802
Recogniser rules + translation length 3-gram Jaccard badness ok 100 82 0.024 0.258 0.1901 0.0034 36.0% 2894.2661
Vocabulary + recognisers + translation length Absolute length percent error ok 100 288 0.024 0.184 0.0426 -0.0005 48.0% 1266.3802
Recogniser rules + translation length Absolute length percent error ok 100 82 0.023 0.186 0.0426 -0.0005 48.0% 1266.3802
Vocabulary + recognisers 3-gram Jaccard badness ok 100 287 0.022 0.248 0.1903 0.0032 36.0% 2894.2661
Recogniser rules 3-gram Jaccard badness ok 100 81 0.022 0.245 0.1904 0.0032 36.0% 2894.2661
Translation length Absolute length percent error ok 100 1 -0.003 -0.081 0.0422 -0.0001 20.0% 10000.0000

Highest Predicted Risk

This list uses the best cross-validated model in this run and sorts passages by predicted badness for 2-gram F1 badness.

Lemma ID v3 runs Source words Observed badness Predicted badness BLEU-4 chrF++ 3-gram F1 Length error
Καρία 2484 1 181.0 0.4484 0.6578 42.0% 65.8% 42.2% 12.6%
Κασώριον 2623 1 14.0 0.5556 0.6149 32.1% 66.0% 35.3% 10.0%
Κάρυστος 2603 1 132.0 0.4327 0.5948 46.5% 69.5% 41.5% 0.6%
Καρχηδών 2604 1 88.0 0.4897 0.5172 35.0% 63.9% 35.7% 9.4%
Καλάσιρις 2085 1 10.0 0.6923 0.5165 32.3% 69.5% 8.3% 15.4%
Κάλυτις 2335 2 16.0 0.6765 0.5153 31.9% 60.3% 14.7% 5.8%
Καππαδοκία 2470 2 57.0 0.4280 0.5018 20.0% 56.8% 41.7% 3.7%
Καδμεία 2059 1 17.0 0.5000 0.4964 41.9% 68.2% 36.8% 10.0%
Κριώα 3530 1 16.0 0.4894 0.4870 38.0% 71.1% 31.1% 11.5%
Καταονία 2628 1 17.0 0.5319 0.4804 36.2% 67.4% 31.1% 4.2%
Καλαβρία 2080 1 12.0 0.6250 0.4716 23.6% 66.7% 13.3% 11.1%
Κύρνος 7247 1 34.0 0.3895 0.4605 47.1% 69.3% 51.6% 6.0%
Καπετώλιον 2468 1 86.0 0.5827 0.4502 30.1% 50.3% 31.0% 14.5%
Κάναστρον 2455 1 43.0 0.5826 0.4495 32.0% 62.9% 26.5% 5.0%
Κύτα 7254 1 58.0 0.4000 0.4314 51.9% 73.6% 47.6% 4.8%
Καρπασία 2597 1 80.0 0.3767 0.4269 54.3% 75.2% 51.6% 6.2%
Κατάνη 2626 1 64.0 0.4945 0.4203 43.8% 64.8% 36.7% 6.3%
Κυτέριον 7255 1 18.0 0.4167 0.4147 53.0% 76.1% 43.5% 27.3%
Κώμη 7266 1 53.0 0.7432 0.4128 13.8% 47.6% 13.7% 10.1%
Καβασσός 2055 1 69.0 0.5048 0.4127 41.6% 67.5% 32.7% 5.5%
Κάλπη 2329 1 36.0 0.2000 0.4114 81.0% 89.1% 72.2% 3.5%
Καικῖνον 2074 1 6.0 0.6842 0.4061 21.4% 66.4% 0.0% 9.1%
Κάλλατις 2119 1 45.0 0.4887 0.4044 31.8% 63.9% 30.5% 7.7%
Κωνώπη 7267 1 47.0 0.3667 0.4044 57.6% 75.7% 50.8% 9.4%
Κάσος 2607 1 44.0 0.4464 0.3992 33.7% 61.5% 43.6% 0.0%
Κωλιάς 7264 1 42.0 0.3445 0.3914 56.9% 76.7% 53.0% 1.6%
Κυρταία 7250 1 29.0 0.3600 0.3900 50.3% 74.5% 43.8% 7.5%
Κεκρυφάλεια 3258 1 17.0 0.1429 0.3883 77.1% 89.4% 76.6% 3.8%
Κάληρος 2116 1 22.0 0.5890 0.3851 29.3% 61.2% 28.2% 12.5%
Κύρη 7243 1 14.0 0.1707 0.3798 82.2% 91.2% 71.8% 4.5%

Predictive Features

Translation Length: chrF++ badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
translation_length Mean v3 translation word count z-scored; mean 43.3, SD 36.9 words; present/high = top quartile, absent/low = bottom quartile 0.04413 100 0.3080 0.1563

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent

Vocabulary Terms: 2-gram F1 badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary χωριον 0.25555 2 0.5835 0.3338
vocabulary οικητωρ 0.20658 6 0.5544 0.3250
vocabulary επι 0.17847 2 0.6113 0.3332
vocabulary τον 0.17278 10 0.4599 0.3253
vocabulary τοις 0.17030 3 0.4885 0.3342
vocabulary ωστε 0.16318 3 0.5518 0.3322
vocabulary και φασι 0.15981 2 0.5012 0.3355
vocabulary καλειται 0.15368 2 0.4723 0.3361
vocabulary οικητωρ και 0.15340 3 0.6190 0.3301
vocabulary και 0.14608 66 0.3827 0.2536
vocabulary ει 0.14562 4 0.5657 0.3293
vocabulary εκαλειτο 0.13836 7 0.4612 0.3296

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary τεταρτω -0.28380 3 0.0929 0.3464
vocabulary εθνικον -0.22072 56 0.3040 0.3831
vocabulary εβδομη -0.20172 2 0.2444 0.3407
vocabulary ως εθνικον -0.17656 5 0.1882 0.3467
vocabulary μεταξυ και -0.15920 5 0.2941 0.3412
vocabulary μεταξυ -0.15920 5 0.2941 0.3412
vocabulary πορρω -0.15309 3 0.0962 0.3463
vocabulary ου πορρω -0.15309 3 0.0962 0.3463
vocabulary πολιτης -0.14055 18 0.3400 0.3385
vocabulary παιδος -0.13934 4 0.2105 0.3441
vocabulary απο παιδος -0.13934 4 0.2105 0.3441
vocabulary εν -0.13888 37 0.3442 0.3356

Vocabulary Terms + Translation Length: 2-gram F1 badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary χωριον 0.26122 2 0.5835 0.3338
vocabulary οικητωρ 0.20964 6 0.5544 0.3250
vocabulary επι 0.18037 2 0.6113 0.3332
vocabulary και φασι 0.17586 2 0.5012 0.3355
vocabulary οικητωρ και 0.16837 3 0.6190 0.3301
vocabulary τοις 0.15968 3 0.4885 0.3342
vocabulary τον 0.15541 10 0.4599 0.3253
vocabulary μοιρα 0.15287 2 0.6121 0.3332
vocabulary καλειται 0.15058 2 0.4723 0.3361
vocabulary ωστε 0.14669 3 0.5518 0.3322
vocabulary και 0.14182 66 0.3827 0.2536
vocabulary δευτερω 0.13076 3 0.5157 0.3333

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary τεταρτω -0.26899 3 0.0929 0.3464
vocabulary εθνικον -0.20497 56 0.3040 0.3831
vocabulary εβδομη -0.19909 2 0.2444 0.3407
vocabulary ως εθνικον -0.17236 5 0.1882 0.3467
vocabulary εν -0.16318 37 0.3442 0.3356
vocabulary μεταξυ και -0.15863 5 0.2941 0.3412
vocabulary μεταξυ -0.15863 5 0.2941 0.3412
vocabulary ως εν -0.14703 7 0.3095 0.3410
vocabulary πορρω -0.14563 3 0.0962 0.3463
vocabulary ου πορρω -0.14563 3 0.0962 0.3463
vocabulary παιδος -0.13816 4 0.2105 0.3441
vocabulary απο παιδος -0.13816 4 0.2105 0.3441

Recogniser Rules: chrF++ badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary gloss occurrence count 0.00206 95 0.2344 0.1542
recogniser_summary gloss rule count 0.00161 95 0.2344 0.1542
recogniser_summary matched occurrence count 0.00116 100 0.2304 N/A
recogniser_rule gloss: καλεῖται/ἐκαλεῖτο/κέκληται/ἐκλήθη (ἀπο...) is/used to be/is/was called/named after (+ ἀπο) 0.00061 14 0.3218 0.2155
recogniser_summary matched rule count 0.00046 100 0.2304 N/A
recogniser_rule formula: X (SETTLEMENT) + Y (genitive REGION) Translate as "a X in Y" 0.00045 62 0.2472 0.2029
recogniser_rule formula: X (nominative PROPER NOUN) + X (genitive PROPER NOUN) Translate as "X, X" 0.00038 20 0.2758 0.2190
recogniser_rule gloss: πόλισμα τό * town 0.00021 5 0.3127 0.2260
recogniser_rule formula: εἰς + «X» (GREEK LETTER) Translate as "ending in X" 0.00018 3 0.3507 0.2266
recogniser_rule gloss: χώρα ἡ region; territory (when belonging to a specific people, tribe, or nation); surrounding territory (when juxtaposed against a πόλις); land (remote or poetical places) 0.00018 12 0.3147 0.2189
recogniser_rule gloss: κώμη ἡ village 0.00018 3 0.3208 0.2276
recogniser_rule gloss: οἰκήτωρ ὁ inhabitant, resident, patron (of a brothel...? - κ123) 0.00015 6 0.3454 0.2230

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary formula rule count -0.00120 99 0.2309 0.1780
recogniser_summary formula occurrence count -0.00095 99 0.2309 0.1780
recogniser_rule formula: ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) Translate as "(just) as 'Y' is from X" -0.00040 34 0.2030 0.2445
recogniser_rule gloss: ἔθνος τό people -0.00040 16 0.1709 0.2417
recogniser_rule formula: X (AUTHOR NAME) + Y (NUMERAL) Translate as "X, book Y" -0.00035 30 0.1911 0.2472
recogniser_rule formula: X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) Translate as "Χ, in his *Y*" -0.00034 17 0.1958 0.2374
recogniser_rule formula: X (nominative) + ὡς + Y (nominative HOMOMORPH) Translate as "'X' as in 'Y'" -0.00033 42 0.2152 0.2413
recogniser_rule formula: ὡς + X (AUTHOR NAME) Translate as "as per X" -0.00027 40 0.2173 0.2391
recogniser_rule formula: τὸ ἐθνικὸν + X (nominative ETHNONYM) Translate as "the ethnonym is 'X'" -0.00020 61 0.2206 0.2456
recogniser_rule formula: Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) Translate as "X, in book Y of his *Z*" -0.00019 11 0.2186 0.2318
recogniser_rule formula: ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) Translate as "just as 'Y' is from the name X" -0.00013 19 0.2141 0.2342
recogniser_rule gloss: ἐθνικόν τό ethnonym -0.00009 60 0.2229 0.2415

Recogniser Rules + Translation Length: chrF++ badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary gloss occurrence count 0.00203 95 0.2344 0.1542
recogniser_summary gloss rule count 0.00159 95 0.2344 0.1542
recogniser_summary matched occurrence count 0.00112 100 0.2304 N/A
translation_length Mean v3 translation word count z-scored; mean 43.3, SD 36.9 words; present/high = top quartile, absent/low = bottom quartile 0.00104 100 0.3080 0.1563
recogniser_rule gloss: καλεῖται/ἐκαλεῖτο/κέκληται/ἐκλήθη (ἀπο...) is/used to be/is/was called/named after (+ ἀπο) 0.00061 14 0.3218 0.2155
recogniser_rule formula: X (SETTLEMENT) + Y (genitive REGION) Translate as "a X in Y" 0.00045 62 0.2472 0.2029
recogniser_summary matched rule count 0.00045 100 0.2304 N/A
recogniser_rule formula: X (nominative PROPER NOUN) + X (genitive PROPER NOUN) Translate as "X, X" 0.00038 20 0.2758 0.2190
recogniser_rule gloss: πόλισμα τό * town 0.00021 5 0.3127 0.2260
recogniser_rule formula: εἰς + «X» (GREEK LETTER) Translate as "ending in X" 0.00018 3 0.3507 0.2266
recogniser_rule gloss: χώρα ἡ region; territory (when belonging to a specific people, tribe, or nation); surrounding territory (when juxtaposed against a πόλις); land (remote or poetical places) 0.00018 12 0.3147 0.2189
recogniser_rule gloss: κώμη ἡ village 0.00018 3 0.3208 0.2276

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary formula rule count -0.00119 99 0.2309 0.1780
recogniser_summary formula occurrence count -0.00095 99 0.2309 0.1780
recogniser_rule gloss: ἔθνος τό people -0.00040 16 0.1709 0.2417
recogniser_rule formula: ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) Translate as "(just) as 'Y' is from X" -0.00040 34 0.2030 0.2445
recogniser_rule formula: X (AUTHOR NAME) + Y (NUMERAL) Translate as "X, book Y" -0.00035 30 0.1911 0.2472
recogniser_rule formula: X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) Translate as "Χ, in his *Y*" -0.00034 17 0.1958 0.2374
recogniser_rule formula: X (nominative) + ὡς + Y (nominative HOMOMORPH) Translate as "'X' as in 'Y'" -0.00032 42 0.2152 0.2413
recogniser_rule formula: ὡς + X (AUTHOR NAME) Translate as "as per X" -0.00027 40 0.2173 0.2391
recogniser_rule formula: τὸ ἐθνικὸν + X (nominative ETHNONYM) Translate as "the ethnonym is 'X'" -0.00020 61 0.2206 0.2456
recogniser_rule formula: Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) Translate as "X, in book Y of his *Z*" -0.00019 11 0.2186 0.2318
recogniser_rule formula: ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) Translate as "just as 'Y' is from the name X" -0.00013 19 0.2141 0.2342
recogniser_rule gloss: ἐθνικόν τό ethnonym -0.00010 60 0.2229 0.2415

Combined Model Features: chrF++ badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary gloss occurrence count 0.00206 95 0.2344 0.1542
recogniser_summary gloss rule count 0.00161 95 0.2344 0.1542
recogniser_summary matched occurrence count 0.00115 100 0.2304 N/A
recogniser_rule gloss: καλεῖται/ἐκαλεῖτο/κέκληται/ἐκλήθη (ἀπο...) is/used to be/is/was called/named after (+ ἀπο) 0.00061 14 0.3218 0.2155
recogniser_summary matched rule count 0.00046 100 0.2304 N/A
recogniser_rule formula: X (SETTLEMENT) + Y (genitive REGION) Translate as "a X in Y" 0.00045 62 0.2472 0.2029
recogniser_rule formula: X (nominative PROPER NOUN) + X (genitive PROPER NOUN) Translate as "X, X" 0.00038 20 0.2758 0.2190
recogniser_rule gloss: πόλισμα τό * town 0.00021 5 0.3127 0.2260
recogniser_rule formula: εἰς + «X» (GREEK LETTER) Translate as "ending in X" 0.00018 3 0.3507 0.2266
recogniser_rule gloss: χώρα ἡ region; territory (when belonging to a specific people, tribe, or nation); surrounding territory (when juxtaposed against a πόλις); land (remote or poetical places) 0.00018 12 0.3147 0.2189
recogniser_rule gloss: κώμη ἡ village 0.00018 3 0.3208 0.2276
recogniser_rule gloss: οἰκήτωρ ὁ inhabitant, resident, patron (of a brothel...? - κ123) 0.00015 6 0.3454 0.2230

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_summary formula rule count -0.00119 99 0.2309 0.1780
recogniser_summary formula occurrence count -0.00095 99 0.2309 0.1780
recogniser_rule formula: ὡς + X (nominative ETYMON) + Y (nominative DERIVED NOUN) Translate as "(just) as 'Y' is from X" -0.00040 34 0.2030 0.2445
recogniser_rule gloss: ἔθνος τό people -0.00040 16 0.1709 0.2417
recogniser_rule formula: X (AUTHOR NAME) + Y (NUMERAL) Translate as "X, book Y" -0.00035 30 0.1911 0.2472
recogniser_rule formula: X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) Translate as "Χ, in his *Y*" -0.00034 17 0.1958 0.2374
recogniser_rule formula: X (nominative) + ὡς + Y (nominative HOMOMORPH) Translate as "'X' as in 'Y'" -0.00033 42 0.2152 0.2413
recogniser_rule formula: ὡς + X (AUTHOR NAME) Translate as "as per X" -0.00027 40 0.2173 0.2391
recogniser_rule formula: τὸ ἐθνικὸν + X (nominative ETHNONYM) Translate as "the ethnonym is 'X'" -0.00020 61 0.2206 0.2456
recogniser_rule formula: Χ (AUTHOR NAME) + ἐν + Y (dative NUMBER) + Z (genitive BOOK NAME) Translate as "X, in book Y of his *Z*" -0.00019 11 0.2186 0.2318
vocabulary εθνικον -0.00014 56 0.2168 0.2475
recogniser_rule formula: ὡς + X (definite ARTICLE + ETYMON) + Y (nominative DERIVED NOUN) Translate as "just as 'Y' is from the name X" -0.00013 19 0.2141 0.2342

Combined Model Features + Translation Length: 3-gram Jaccard badness

Positive coefficients predict worse translation scores for the selected target. Negative coefficients predict better scores. Translation length is z-scored; vocabulary features exclude detected proper-noun tokens. These are exploratory ridge coefficients, not causal claims.

Features associated with worse scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
recogniser_rule formula: X (AUTHOR NAME) + Y (dative NUMBER) + Z (genitive BOOK NAME) Translate as "X, book Y in his *Z*" 0.09656 28 0.6084 0.5755
recogniser_rule gloss: χωρίον τό locality; point (only in μέσα χωρία: ‘halfway point’) 0.07078 3 0.8982 0.5750
translation_length Mean v3 translation word count z-scored; mean 43.3, SD 36.9 words; present/high = top quartile, absent/low = bottom quartile 0.05909 100 0.7014 0.4515
recogniser_rule gloss: οἰκήτωρ ὁ inhabitant, resident, patron (of a brothel...? - κ123) 0.05560 6 0.8128 0.5702
recogniser_rule formula: καί + X (nominative PROPER NOUN) + Y (nominative PROPER NOUN) Translate as "Y is also 'X'" 0.05482 13 0.6457 0.5756
vocabulary τον 0.05386 10 0.7556 0.5657
recogniser_rule formula: X (nominative PROPER NOUN) + X (genitive PROPER NOUN) Translate as "X, X" 0.05347 20 0.6427 0.5702
vocabulary δευτερω 0.05004 3 0.7988 0.5781
recogniser_rule gloss: μοῖρα ἡ region or part (in geographic contexts); district (in urban contexts only) 0.04777 4 0.8089 0.5754
recogniser_rule formula: X (nominative PROPER NOUN)... + ἀπό + Y (genitive ARTICLE + genitive ETYMON) Translate as "X... from the form 'Y'" 0.04451 13 0.6309 0.5778
vocabulary χωριον 0.04124 2 0.8840 0.5786
recogniser_rule formula: ὡς + X (AUTHOR NAME) + Y (dative NUMBER) + Z (genitive BOOK NAME) Translate as 'as per X, in book Y of his *Z*' 0.03841 5 0.6647 0.5805

Features associated with better scores

Type Feature Detail Coefficient Passages Mean badness present Mean badness absent
vocabulary τεταρτω -0.11219 3 0.1333 0.5987
recogniser_rule gloss: ἔθνος τό people -0.10664 16 0.4585 0.6088
vocabulary ως εθνικον -0.08812 5 0.3441 0.5974
recogniser_rule formula: X (AUTHOR NAME) + Y (NUMERAL) Translate as "X, book Y" -0.06212 30 0.5140 0.6150
recogniser_rule formula: X... + πρός + Y (dative) Translate as "X... near Y" -0.04795 12 0.5323 0.5919
recogniser_rule formula: X (AUTHOR NAME) + ἐν + Y (dative BOOK NAME) Translate as "Χ, in his *Y*" -0.04681 17 0.4694 0.6083
vocabulary εθνικον -0.04595 56 0.5500 0.6289
recogniser_rule formula: X (nominative) + ὡς + Y (nominative HOMOMORPH) Translate as "'X' as in 'Y'" -0.04579 42 0.5464 0.6125
vocabulary εβδομη -0.04482 2 0.4099 0.5883
vocabulary πορρω -0.04359 3 0.2178 0.5961
vocabulary ου πορρω -0.04359 3 0.2178 0.5961
vocabulary εν εθνικον -0.04205 3 0.4126 0.5900

Worst v3 Translation Sentences

Worst means high average percentile badness across chrF, sentence BLEU, ROUGE-L, 3-gram F1, and absolute word-count delta. This is a review queue, not a human error judgment.

Rank Headword Greek sentence v3 candidate Human-approved translation
1 Καλαμένθη κρεῖττον οὖν ὡς Ἡρόδοτος διὰ τοῦ « ι ». The better form, then, is as per Herodotos, written with ι. It is better to have it with ι, as per Herodotos. A city of the Phoenicians.
2 Κάλλατις ὡς κάλαθος εὑρέθη ἐοικὼς τοῖς θεσμοφοριακοῖς. It is as in 'kalathos', because a basket was found resembling those used at the Thesmophoria. Because a basket similar to that which is ‘Thesmophorian’ was found there.
3 Κατάνη ἀπὸ δὲ τῆς Κατάνης Χαρώνδας, ὁ διάσημος τῶν Ἀθήνησι νομοθετῶν. From Katane came Charondas, renowned among the lawgivers at Athens. From Katane was Charondas, one of the lawmakers made famous by the Athenians.
4 Κοτιάειον δοκεῖ δὲ ἀπὸ τοῦ Κόσα Κόσαιον, ὡς τοῦ Μίδα Μίδαιον, καὶ προσθέσει τοῦ « ι » Κοσιαῖον, καὶ τροπῇ Κοτιαῖον. It seems that 'Kosaion' is from 'Kosa', just as 'Midaion' is from 'Midas'; and by addition of ι, 'Kosiaion', and by alteration, 'Kotiaion'. it does seem that 'Kosaeion' is from the name Kosas (as 'Midaeion' is from the name Midas), then it becomes 'Kosiaeion' by addition of ι and 'Kotiaeion' through sound shift.
5 Κάνωπος καὶ Κανωβίς θηλυκόν. There is also ‘Kanobis’ as a feminine form. The feminine is ‘Kanobis’.
6 Καιρή εἴρηται ἐν τῷ περὶ Ἀγύλλης. It has been stated in the entry on Agylla. It is cited in *On Agylla*.
7 Καρία ἡ χρῆσις δὲ διὰ διφθόγγου κατὰ συνεκδρομὴν τοῦ σώτειρα ὀλέτειρα“. Common usage, however, has it with a diphthong by association with σώτειρα, ὀλέτειρα.' The common usage, however, is with the diphthong, according to the analogy of 'soteira', 'oleteira'.'
8 Καπετώλιον Ἀσκληπιεῖον, Ἀσκληπιός γάρ· Ἀσκληπιεῖον, for there is Ἀσκληπιός; Asklepieion (Ἀσκληπιεῖον) is thus from 'Asklepios' (Ἀσκληπιός);
9 Καπετώλιον Πτολεμαεῖον, Πτολεμαῖος γάρ· Πτολεμαεῖον, for there is Πτολεμαῖος; Ptolemaeion (Πτολεμαεῖον) is thus from 'Ptolemaios' (Πτολεμαῖος);
10 Καπετώλιον Καπετώλιον, ἐν Ῥώμῃ λόφος Ταρπαῖος πάλαι λεγόμενος. Kapetolion: in Rome, a hill formerly called Tarpeios. Kapetolion: a hill in Rome that was long ago called 'Tarpaios'.
11 Κάλυτις ὁ οἰκήτωρ Καλυτίτης, καὶ τὸ θηλυκὸν Καλυτίς, διὰ τὸ προειλῆφθαι τὸν χαρακτῆρα. The inhabitant is 'Kalytites', and the feminine is 'Kalytis', because the characteristic element has already been taken in advance. An inhabitant is a 'Kalytites'; the feminine is also 'Kalytis' due to the form being anticipated.
12 Καλλίπολις δευτέρα κατὰ τὸν Ἀνάπλουν. A second, according to the *Anaplous*. (2) Along the Anaplous.
13 Καρία Ἡρωδιανὸς δὲ ἐν μὲν τῇ Ὀρθογραφίᾳ (2,410,22) ἀμφίβολον αὐτό φησιν. ἐν δὲ τῇ Καθόλου (1,250,14) <τῇ> χρήσει ἑπόμενος διὰ διφθόγγου φησίν, ὑπομνηματίζων δὲ τὸ Περὶ γενῶν Ἀπολλωνίου (2,777,13) διὰ τοῦ ι μακροῦ. „ἔστι γὰρ ὅτε μετὰ τὴν διαίρεσιν ἔκτασις γίγνεται, ὀίομαι ὄιγον ὄιδα παρ’ Αἰολεῦσιν, ἀντὶ τοῦ οἶδα. Herodianos in his *Orthography* is undecided, but, following general usage, says it is with a diphthong; when commenting on Apollonios' *On Genders*, however, he gives it with long ι: 'For there are times when lengthening occurs after separation: ὀίομαι, ὄιγον, ὄιδα among the Aiolians, instead of οἶδα. While Herodian says that this is doubtful in his *Orthography* (and in his *General Prosody* he says that it uses the diphthong following the common usage), he comments on Apollonios’ *On Genders* that it is with long ι: 'for there is occasion when lengthening occurs after diaresis: 'oïomai', 'oïgon', 'oïda' among the Aeolians rather than 'oida'.
14 Κάσιον ὁ πολίτης Κασιώτης ὡς Πηλουσιώτης, καὶ θηλυκὸν Κασιῶτις, καὶ τὸ κτητικὸν Κασιωτικός, ἀφ´ οὗ ἐν τῇ συνηθείᾳ τὰ Κασιωτικὰ ἱμάτια. the feminine is 'Kasiotis', and the possessive is 'Kasiotikos', from which in ordinary usage comes the phrase 'Kasiotika cloaks'. The possessive is 'Kasiotikos', hence the term 'Kasiotic cloaks' in ordinary language.
15 Κάστνιον ἔδει δὲ Καστνιώτης ὡς Πηλιώτης. It ought, however, to be 'Kastniotes', as 'Peliotes' is from Pelion. However, it should be 'Kastniotes' (as in 'Peliotes').
16 Κυρτώνιος τὸ ἐθνικὸν τῷ τῆς χώρας ἔθει Κυρτωνῖνος ὡς Σατορνῖνος. The ethnonym, according to the usage of the region, is 'Kyrtoninos', as in 'Satorninos'. In local usage, the ethnonym is 'Kyrtoninos' (as in 'Saturninos').
17 Κάναι Καναῖος Ζεύς οὐ μόνον ἀπὸ τοῦ Καναίου, ἀλλὰ καὶ ἀπὸ τῆς Κάνης. Kanaios Zeus is named not only after Kanaios, but also after Kane. Zeus Kanaios is not only from the form 'Kanaios', but also from the form 'Kane'.
18 Κάσος ἀπῴκισται δὲ τῆς νήσου καὶ τὸ ἐν Συρίᾳ ὄρος Κάσιον. The mountain Kasion in Syria has also been colonised from the island. Mount Kasios in Syria was also settled from this island.
19 Κορώνεια τετάρτη πόλις Κύπρου. A fourth is a city of Cyprus. (4) a city in Cyprus;
20 Καπετώλιον ὅσα γὰρ ἔχει προϋπάρχοντα εἰς « ος » καθαρόν, παραληγόμενα ἢ μόνῳ τῷ « ι » ἢ προηγουμένου αὐτοῦ τοῦ « α » ὥστε εἶναι πρὸ τέλους τὴν « αι » δίφθογγον, προπερισπᾶται, ἢ καὶ ὅσα κτητικά. For all words which have pre-existing forms ending in pure -ος, and whose penult has either ι alone or this preceded by α, so that the diphthong αι comes before the final syllable, are accented with a circumflex on the penult; so too all possessives. This is because forms whose base already ends in postvocalic -ος—when either a single ι is in the penultimate position or α precedes it so that the diphthong αι stands before the ultima—will be accented with a circumflex on the penult, and the same applies to possessive forms.

Downloadable Tables

Generated: 2026-07-13 12:08:00 UTC. Recogniser detector version: translation_guidance_scan_v4.