Lexicon Engine Benchmark Suite
Rigorous empirical evaluation across 4 independent random seeds (100 unique sentences per engine, 500 evaluations total) measuring generalization accuracy, latency, and standard deviation (±1σ) across expanded 1,050-sentence authentic corpus (350 JFLEG, 350 BEA-2019, 350 LOCNESS).
How to read these scores
Grammar correction rewards careful edits more than aggressive ones. We report the usual GEC metrics below, plus corpus-specific fluency and grammar hit rates.
- Precision
- Of the edits the engine made, how many were actually correct? High precision means fewer wrong or noisy changes.
- Recall
- Of the real errors in the text, how many did the engine catch? High recall means fewer missed mistakes.
- F₀.₅
- A single combined score that weights precision about twice as heavily as recall, the standard for grammar error correction, because a bad rewrite is usually worse than leaving text alone.
- Clean FPR
- How often the engine “fixes” sentences that were already correct. Lower is better; 0% means it stayed quiet on clean prose.
- Latency
- Average time to process one sentence. Lower is better for live typing; Quality trades speed for accuracy.
- JFLEG & BEA
- Corpus hit rates: JFLEG measures fluency and natural phrasing; BEA measures grammar and usage. Shown as correct / eligible items.
Standard 4B: local edge model delivering 221 ms live keystroke speed.
Quality 27B: top F₀.₅ score (75.7%), highest recall (45.7%), and zero clean FPR.
LanguageTool: instant syntax at keystroke speed.
100 unique sentences across 4 seeds per engine from 1,050 authentic pool.
Cross-Validation Ladder
Pareto Efficiency Frontier
Evaluation Leaderboard
Executive summary of the 4-seed cross-validation benchmark across 5 engines.
Click
[ 4 Runs ▾ ] on any tier to inspect all individual seeds, or toggle to Detailed
Matrix for the full 25-row dataset.