Compact Modern Chinese travels best.
zh_compact is the most stable choice across model families, often preserving quality while controlling completion length.
Cross-lingual LLM evaluation
A controlled study of whether English, compact Modern Chinese, and Classical Chinese move an LLM to a different quality–cost frontier — with tasks and model backends held fixed.
The result is not “one language wins.” Prompt language interacts with model capacity, task type, and reasoning visibility.
zh_compact is the most stable choice across model families, often preserving quality while controlling completion length.
Classical Chinese is strongest overall for GPT-5.4 in the corrected matrix, but its advantage does not transfer reliably to smaller open models.
Hidden reasoning mainly cuts visible-token cost; compact-visible reasoning preserves the strongest GPT-5.4/Wényán point at much lower cost than explicit process.
Only the system-prompt language changes. The task subset, model backend, evaluator, and output budget stay fixed.
base
The conventional control: clear, terse, and familiar to every model family in the study.
“Follow the user’s instructions carefully… eliminate redundancy and aim for compact, clear expression.”
zh_compact
A compact Chinese control that separates language effects from the special literary style of Wényán.
“请严格遵循用户要求……删繁就简,要求表达精炼、意思清楚。”
wy
An information-dense natural language whose ellipsis and terse grammar make it a plausible prompt-compression regime.
“凡码与引文外,释理叙事宜从简。禁绝白话冗词,务求辞约义明。”
This public-manifest MATH-500 item keeps the mathematical problem and answer fixed while changing the language used for both reasoning and response. The illustrative outputs make the compression contrast directly visible.
Problem / 题目: If a and b are positive integers such that gcd(a,b)=210, lcm(a,b)=210³, and a<b, how many possible values are there for a?
Let a=210x and b=210y. Then gcd(x,y)=1 and xy=210²=2²·3²·5²·7². Each prime-square factor must go wholly to one of x,y, giving 2⁴=16 ordered assignments. Since a<b, exactly half remain. <answer>8</answer>
题目 / Problem: 如果 a 和 b 是正整数,且 gcd(a,b)=210,lcm(a,b)=210³,并且 a<b,那么 a 有多少个可能的取值?
令 a=210x,b=210y。由 gcd(a,b)=210 可知 gcd(x,y)=1;再由 lcm(a,b)=210³ 得 xy=210²=2²·3²·5²·7²。因为 x 与 y 互质,每个素因子的平方必须整体分配给其中一方,共有 2⁴=16 种有序分配。条件 a<b 恰好保留一半,所以 a 有 8 个可能值。<answer>8</answer>
題目: 若 a 與 b 為正整數,且 gcd(a,b)=210,lcm(a,b)=210³,並且 a<b,則 a 之可能值有幾何?
令 a=210x,b=210y,則 (x,y)=1,而 xy=210²=2²·3²·5²·7²。四素之方各歸一數,凡 2⁴ 配;以 a<b,取其半,故八。<answer>8</answer>
Task counts cover the displayed task rendering; output counts include the required answer tag. These responses illustrate the format and are not additional leaderboard runs.
The 2048-token rerun removes the largest truncation artifact from the pilot and reports score together with total inference tokens.
Switch between the aggregate view and each benchmark to inspect score, token cost, and cost-normalized efficiency for every model–language pair.
| Rank | Model | Prompt language | N | Score | Prompt tok | Output tok | Total tok | Score / 1k tok |
|---|---|---|---|---|---|---|---|---|
| 1 | gpt-5.4 | Wényán | 59 | 0.763BEST | 186.93 | 45.15 | 232.08 | 3.286 |
| 2 | gpt-5.4 | Simplified Chinese | 59 | 0.746 | 199.93 | 49.05 | 248.98 | 2.995 |
| 3 | gpt-5.4 | English | 59 | 0.729 | 181.93 | 54.64 | 236.58 | 3.081 |
| 4 | gpt-4o | Simplified Chinese | 59 | 0.627 | 200.93 | 29.59 | 230.53 | 2.720 |
| 5 | gpt-4o | English | 59 | 0.593 | 182.93 | 37.12 | 220.05 | 2.696 |
| 6 | gpt-4o | Wényán | 59 | 0.576 | 187.93 | 32.86 | 220.80 | 2.610 |
| 7 | qwen3-4b | Simplified Chinese | 59 | 0.475 | 203.85 | 69.98 | 273.83 | 1.733 |
| 8 | qwen3-4b | Wényán | 59 | 0.475 | 197.85 | 98.71 | 296.56 | 1.600 |
| 9 | qwen3-4b | English | 59 | 0.441 | 197.85 | 121.95 | 319.80 | 1.378 |
| 10 | qwen3-1.7b | English | 59 | 0.339 | 197.85 | 127.00 | 324.85 | 1.044 |
| 11 | qwen3-1.7b | Simplified Chinese | 59 | 0.271 | 203.85 | 125.20 | 329.05 | 0.824 |
| 12 | qwen3-1.7b | Wényán | 59 | 0.254 | 197.85 | 228.12 | 425.97 | 0.597 |
| Rank | Model | Prompt language | N | Score | Prompt tok | Output tok | Total tok | Score / 1k tok |
|---|---|---|---|---|---|---|---|---|
| 1 | qwen3-4b | Simplified Chinese | 11 | 1.000BEST | 107.00 | 129.64 | 236.64 | 4.226 |
| 2 | qwen3-4b | Wényán | 11 | 1.000BEST | 101.00 | 151.73 | 252.73 | 3.957 |
| 3 | qwen3-4b | English | 11 | 1.000BEST | 101.00 | 177.82 | 278.82 | 3.587 |
| 4 | gpt-5.4 | Wényán | 11 | 1.000BEST | 96.45 | 218.18 | 314.64 | 3.178 |
| 5 | gpt-5.4 | Simplified Chinese | 11 | 1.000BEST | 109.45 | 239.09 | 348.55 | 2.869 |
| 6 | gpt-5.4 | English | 11 | 1.000BEST | 91.45 | 269.36 | 360.82 | 2.771 |
| 7 | qwen3-1.7b | Wényán | 11 | 1.000BEST | 101.00 | 391.45 | 492.45 | 2.031 |
| 8 | gpt-4o | Simplified Chinese | 11 | 0.909 | 110.45 | 141.82 | 252.27 | 3.604 |
| 9 | gpt-4o | English | 11 | 0.909 | 92.45 | 182.36 | 274.82 | 3.308 |
| 10 | qwen3-1.7b | English | 11 | 0.909 | 101.00 | 314.55 | 415.55 | 2.188 |
| 11 | qwen3-1.7b | Simplified Chinese | 11 | 0.909 | 107.00 | 319.91 | 426.91 | 2.129 |
| 12 | gpt-4o | Wényán | 11 | 0.818 | 97.45 | 161.45 | 258.91 | 3.160 |
| Rank | Model | Prompt language | N | Score | Prompt tok | Output tok | Total tok | Score / 1k tok |
|---|---|---|---|---|---|---|---|---|
| 1 | gpt-5.4 | Wényán | 24 | 0.708BEST | 150.79 | 6.00 | 156.79 | 4.518 |
| 2 | gpt-4o | Simplified Chinese | 24 | 0.708BEST | 164.79 | 5.46 | 170.25 | 4.161 |
| 3 | gpt-5.4 | Simplified Chinese | 24 | 0.667 | 163.79 | 6.00 | 169.79 | 3.926 |
| 4 | gpt-5.4 | English | 24 | 0.625 | 145.79 | 5.88 | 151.67 | 4.121 |
| 5 | gpt-4o | Wényán | 24 | 0.583 | 151.79 | 4.71 | 156.50 | 3.727 |
| 6 | gpt-4o | English | 24 | 0.542 | 146.79 | 5.29 | 152.08 | 3.562 |
| 7 | qwen3-4b | Wényán | 24 | 0.292 | 158.21 | 165.50 | 323.71 | 0.901 |
| 8 | qwen3-4b | Simplified Chinese | 24 | 0.250 | 164.21 | 108.46 | 272.67 | 0.917 |
| 9 | qwen3-4b | English | 24 | 0.250 | 158.21 | 155.62 | 313.83 | 0.797 |
| 10 | qwen3-1.7b | English | 24 | 0.208 | 158.21 | 78.12 | 236.33 | 0.882 |
| 11 | qwen3-1.7b | Simplified Chinese | 24 | 0.083 | 164.21 | 23.54 | 187.75 | 0.444 |
| 12 | qwen3-1.7b | Wényán | 24 | 0.000 | 158.21 | 120.92 | 279.12 | 0.000 |
| Rank | Model | Prompt language | N | Score | Prompt tok | Output tok | Total tok | Score / 1k tok |
|---|---|---|---|---|---|---|---|---|
| 1 | gpt-5.4 | English | 24 | 0.708BEST | 259.54 | 5.00 | 264.54 | 2.678 |
| 2 | gpt-5.4 | Wényán | 24 | 0.708BEST | 264.54 | 5.00 | 269.54 | 2.628 |
| 3 | gpt-5.4 | Simplified Chinese | 24 | 0.708BEST | 277.54 | 5.00 | 282.54 | 2.507 |
| 4 | gpt-4o | English | 24 | 0.500 | 260.54 | 2.38 | 262.92 | 1.902 |
| 5 | gpt-4o | Wényán | 24 | 0.458 | 265.54 | 2.08 | 267.62 | 1.713 |
| 6 | qwen3-4b | Simplified Chinese | 24 | 0.458 | 287.88 | 4.17 | 292.04 | 1.569 |
| 7 | gpt-4o | Simplified Chinese | 24 | 0.417 | 278.54 | 2.29 | 280.83 | 1.484 |
| 8 | qwen3-4b | Wényán | 24 | 0.417 | 281.88 | 7.62 | 289.50 | 1.439 |
| 9 | qwen3-4b | English | 24 | 0.375 | 281.88 | 62.67 | 344.54 | 1.088 |
| 10 | qwen3-1.7b | English | 24 | 0.208 | 281.88 | 89.92 | 371.79 | 0.560 |
| 11 | qwen3-1.7b | Simplified Chinese | 24 | 0.167 | 287.88 | 137.62 | 425.50 | 0.392 |
| 12 | qwen3-1.7b | Wényán | 24 | 0.167 | 281.88 | 260.46 | 542.33 | 0.307 |
The corrected matrix uses a uniform 2048-token output budget. All tracks are free of cap hits and empty outputs except qwen3-1.7b, which retains four length truncations. Token counts are most meaningful within a model/backend.
A second three-run study varies whether reasoning is explicit, hidden, or compact-visible while keeping the language conditions fixed.
Every run records the exact task subset, prompt condition, configured token budget, finish reason, token usage, and evaluator result. The pipeline keeps raw execution, aggregation, and plotting separate.
fixed manifest
→ task × language × model runs
→ results.jsonl + usage metadata
→ truncation audit
→ cross-run aggregation
→ score / token Pareto plots
The public repository includes manifests, runners, aggregated reports, plotting scripts, and publish-safe result artifacts.