How to read these scores

Grammar correction rewards careful edits more than aggressive ones. We report the usual GEC metrics below, plus corpus-specific fluency and grammar hit rates.

Precision
Of the edits the engine made, how many were actually correct? High precision means fewer wrong or noisy changes.
Recall
Of the real errors in the text, how many did the engine catch? High recall means fewer missed mistakes.
F₀.₅
A single combined score that weights precision about twice as heavily as recall, the standard for grammar error correction, because a bad rewrite is usually worse than leaving text alone.
Clean FPR
How often the engine “fixes” sentences that were already correct. Lower is better; 0% means it stayed quiet on clean prose.
Latency
Average time to process one sentence. Lower is better for live typing; Quality trades speed for accuracy.
JFLEG & BEA
Corpus hit rates: JFLEG measures fluency and natural phrasing; BEA measures grammar and usage. Shown as correct / eligible items.
Spotlight Presets:
Sweet Spot Recommended
71.2 F₀.₅ · 4-run mean (±9.5%)

Standard 4B: local edge model delivering 221 ms live keystroke speed.

84.1% precision 45.1% recall 221 ms
Publication Quality SOTA
75.7 F₀.₅ · 4-run mean (±6.9%)

Quality 27B: top F₀.₅ score (75.7%), highest recall (45.7%), and zero clean FPR.

85.5% peak run 45.7% recall 91.2% precision
Live Speed Rules
35 ms / sentence

LanguageTool: instant syntax at keystroke speed.

64.6% F₀.₅ mean 95.0% precision
Eval Scale 5×4 Matrix
500 sentence evaluations

100 unique sentences across 4 seeds per engine from 1,050 authentic pool.

5 engines 4 seeds ±1σ variance

Cross-Validation Ladder

4-Seed Macro Mean with Peak Run Potential across 100 Sentences
4-Run Macro Mean (±1σ)
Peak Run Score
Tap or hover bars for variance breakdown

Pareto Efficiency Frontier

Macro Mean F₀.₅ Quality vs. Latency Trade-Off
Macro Mean F₀.₅
Normal (LT)
Light (1B)
Standard (4B)
Quality (27B)
QuillBot (Free)

Evaluation Leaderboard

Executive summary of the 4-seed cross-validation benchmark across 5 engines.
Click [ 4 Runs ▾ ] on any tier to inspect all individual seeds, or toggle to Detailed Matrix for the full 25-row dataset.

View
Rank