Surveys◆ AI-generated · Sourced

Paper: "First-Language Bias" in LLM Essay Scoring — Robust Across Prompts, Yet Scores Differ by Test-Taker's L1

Paper: "First-Language Bias" in LLM Essay Scoring — Robust Across Prompts, Yet Scores Differ by Test-Taker's L1
TL;DR

A study runs LoRA-adapted open model Gemma-3-27B on the full TOEFL11 corpus (12,100 essays, 11 first languages, 8 prompts): 77.79% band agreement, QWK 0.702, robust cross-prompt — but reveals first-language (L1) scoring effects.

Stress-testing on 12,100 TOEFL essays

This paper (arXiv:2607.14605) examines cross-prompt generalization and first-language (L1) scoring effects in LLM automated essay scoring (AES). Reusing the exact model and inference config of the "AiAWE" system — a LoRA-adapted Gemma-3-27B-it open model fine-tuned on 480 argumentative essays from two prompts — it evaluates on the full TOEFL11 corpus: 12,100 essays by test-takers from 11 first-language backgrounds across 8 prompts, none seen in training.

Result: robust across prompts

Raw scores (0.5–5.0) are mapped to ETS's three proficiency bands (low/medium/high) for direct comparison. Overall band agreement was 77.79%, quadratic weighted kappa (QWK) 0.702, with adjacent-band agreement of 99.98%. Crucially, accuracy was stable across all eight unseen prompts, with no advantage for prompts thematically related to training — indicating robust cross-prompt generalization.

The concern: first-language (L1) bias

But the study also found scoring differences correlated with the test-taker's first language — an "L1 bias." The implication: even an AES model that generalizes well across prompts still needs separate scrutiny for fairness across native-language groups. For high-stakes uses of LLMs (exams, hiring, education) this is a signal that must be faced: high overall accuracy ≠ fair to every group. The open, reproducible setup (same model, same config, public corpus) also makes such fairness audits easier to verify independently.

Sources (compliance trail)
https://arxiv.org/abs/2607.14605
Umi Intelligence · Enroll / Contact

Turn “understanding the frontier” into “putting it to work”

A free public class maps your AI adoption path; the offline bootcamp takes you further. Reach out anytime.

✉ hello@umi6.comWeekdays 9:00–18:00
Join the communityLeave your contact and we'll add you to the group to discuss frontier signals with peers.