Translation quality distributions
Entry-level histograms compare prompt v1, v2, and v3 on the frozen 100-entry Kappa paper corpus. The default view is GPT-5.6 Sol; the same fixed bins are available for every dated OpenAI model and the three Claude comparison models.
What “quality” means here: reference similarity to the approved house-style translation, not an independent expert judgment of correctness. The default score is the unweighted mean of BLEU-4, chrF++, METEOR, and ROUGE-L.
Prompt-version score distributions
All benchmark models
Each compact chart uses the same 0–1 score axis and three directly labelled histogram rows. Select a model to open its larger comparison above.
Model and prompt summary
The table follows the selected metric. Percentiles are calculated across the 100 entry-level scores in each complete cell.
| Provider | Model | Prompt | n | Mean | Median | SD | P10 | P90 |
|---|
Downloads
Method and comparability
Each model/prompt cell contains the same 100 entries from Gabe's final Kappa review-tracker export. Scores compare the latest completed or approved translation run in that cell with the approved human translation. Histograms use 20 fixed-width bins from 0 to 1 and show the percentage of entries in each bin, so every panel has the same score scale. GPT-5.2 v1 retains the actual run-model labels recorded in PostgreSQL; the profile cell is not relabelled from those run records.
Claude Opus 4.8 v2 is displayed as a missing cell because no translations exist for that profile/version. It is not treated as a zero score or removed from the model list.
Generated 2026-07-25 12:54:33 UTC from the paper-corpus prompt-evaluation rows.