Translation quality distributions

Entry-level histograms compare prompt v1, v2, and v3 on the frozen 100-entry Kappa paper corpus. The default view is GPT-5.6 Sol; the same fixed bins are available for every dated OpenAI model and the three Claude comparison models.

What “quality” means here: reference similarity to the approved house-style translation, not an independent expert judgment of correctness. The default score is the unweighted mean of BLEU-4, chrF++, METEOR, and ROUGE-L.

Prompt-version score distributions

v1v2v3Vertical rule: mean scoreBin width: 5 percentage points
Three normalized histograms comparing prompt versions for the selected translation model.

All benchmark models

Each compact chart uses the same 0–1 score axis and three directly labelled histogram rows. Select a model to open its larger comparison above.

Model and prompt summary

The table follows the selected metric. Percentiles are calculated across the 100 entry-level scores in each complete cell.

ProviderModelPromptnMeanMedianSDP10P90

Downloads

Method and comparability

Each model/prompt cell contains the same 100 entries from Gabe's final Kappa review-tracker export. Scores compare the latest completed or approved translation run in that cell with the approved human translation. Histograms use 20 fixed-width bins from 0 to 1 and show the percentage of entries in each bin, so every panel has the same score scale. GPT-5.2 v1 retains the actual run-model labels recorded in PostgreSQL; the profile cell is not relabelled from those run records.

Claude Opus 4.8 v2 is displayed as a missing cell because no translations exist for that profile/version. It is not treated as a zero score or removed from the model list.

Generated 2026-07-25 12:54:33 UTC from the paper-corpus prompt-evaluation rows.