Tighter on-device Cleanup with Model 4.1

Tighter on-device Cleanup with Model 4.1

Today we are shipping Epilude Model 4.1, a refinement of the on-device model that powers Local Mode. It is built for the longer dictations where Model 4 could occasionally ramble or repeat itself. The result is tighter writing, fewer critical errors, and the same private, offline experience on your Mac.

What Model 4.1 is

Local Mode runs two models directly on your Mac. A speech model turns your voice into text. A cleanup model turns that raw transcript into writing you would actually send: it removes filler, fixes punctuation, resolves false starts and mid-sentence corrections, and follows the tone you asked for.

Model 4.1 is a new version of the cleanup model. It is our fine-tune of an open-weight model in the Qwen family. The speech model is unchanged. Together they keep the same on-device footprint and hardware requirements as Model 4.

SpecModel 4.1
Download size1.5 GB, one time
HardwareApple Silicon Macs
Full AI CleanupMacs with 16 GB of memory or more
DecodingGreedy, fully deterministic
ConnectivityWorks entirely offline

How we evaluate

We maintain an internal benchmark of 90 dictation scenarios spanning punctuation, formatting, self-corrections, tone control, multiple languages, mixed-language speech, and long-input structure. Every candidate runs the full suite repeatedly. Frontier-model judges cast multiple independent votes on every output, and a scenario counts as passed only when it passes in every judged run. We also re-grade identical outputs so grader noise does not look like a model change.

We do not publish this benchmark, and its results are not comparable to anything external. It is how we decide whether a cleanup model is safe enough to ship.

Results

On that internal benchmark, Model 4.1 clears one more scenario than Model 4 and reduces the critical-error set without introducing a new one:

Internal benchmarkModel 4Model 4.1
Scenarios passed, of 907879
Critical errors43

A scenario counts as passed only when it passes in every repeated judged run. The remaining gains concentrate in the work that becomes visible only after you have been talking for a while: keeping a long answer on track, avoiding repetition, and staying faithful to the parts that are easy to over-clean.

What we learned building it

This release reinforced a lesson that has shaped each model generation: evaluation quality has to match the failures you are trying to prevent. A broad score can hide the one sentence a person would never send. We made the quality bar more specific around long-input faithfulness, then held the new model to it across repeated runs.

It also reinforced the value of model averaging, a known technique in the research literature. We selected a stable combined model to avoid rare failure modes and deliver more dependable cleanup.

Limitations

Model 4.1 is not perfect. Long, highly structured dictations remain harder than short messages, especially when a single thought contains several corrections, quoted material, and a change of tone. Cloud mode is still ahead on our internal suite overall. Local Mode remains the choice when keeping your audio and text on your Mac matters most.

What's next

We will keep working on the difficult cases that matter most in real writing: longer inputs, multilingual dictation, and models that preserve a speaker's intent under cleanup. If Model 4.1 changes how dictation feels on your Mac, or if you catch it making a mistake, we want to hear about it through the Help Center.

Epilude Team4 min readproductlaunchlocal-modemodels