Translation quality distributions
Entry-level histograms compare prompt v1, v2, and v3 on the frozen 100-entry Kappa paper corpus. The default view is GPT-5.6 Sol; the same fixed bins are available for every dated OpenAI model and the three Claude comparison models.
What “quality” means here: reference similarity to the approved house-style translation, not an independent expert judgment of correctness. The default score is the unweighted mean of BLEU-4, chrF++, METEOR, and ROUGE-L.
Prompt-version score distributions
All benchmark models
Each compact chart uses the same 0–1 score axis and three directly labelled histogram rows. Select a model to open its larger comparison above.
Model and prompt summary
The table follows the selected metric. Percentiles are calculated across the 100 entry-level scores in each complete cell.
| Provider | Model | Prompt | n | Mean | Median | SD | P10 | P90 |
|---|
Downloads
Method and comparability
Each model/prompt cell contains the same 100 entries from Gabe's final Kappa review-tracker export. Scores compare the latest completed or approved translation run in that cell with the approved human translation. Histograms use 20 fixed-width bins from 0 to 1 and show the percentage of entries in each bin, so every panel has the same score scale. GPT-5.2 v1 retains the actual run-model labels recorded in PostgreSQL; the profile cell is not relabelled from those run records.
All three Claude models have complete v1, v2 and v3 cells. As with the OpenAI comparisons, each cell contains the same 100 entries.
Generated 2026-09-20 11:23:17 UTC from the paper-corpus prompt-evaluation rows.