Local AI transcription finds every filler word in your footage and cuts them with frame-accurate precision. 14 patterns detected including um, uh, like, you know, basically and literally.
EditLint AI ships with 14 default filler word patterns covering the most common verbal hesitations in English. The list is fully editable — add your own words and remove any you want to keep.
These are detected automatically in every transcription pass. Each match includes a start time, end time and confidence score so you can review before cutting.
Context-aware detection means "like" is only flagged when used as a hesitation filler — not when used grammatically as a comparison.
EditLint AI uses an on-device Whisper model to transcribe your audio track. Processing happens entirely on your machine — no files are sent to any server. Transcription takes roughly 1/6th of real time, so a 60-second clip transcribes in about 10 seconds.
The transcript is parsed against your filler word list. Every match is returned with a precise start time, end time, word matched and confidence score. Context-aware logic ensures words like "like" are only flagged when used as hesitation fillers, not when used grammatically in a sentence.
Every detected filler appears in the EditLint panel. You can click any item to jump the playhead to that exact word, preview the audio around it, and tick or untick individual detections. Low-confidence matches are highlighted so you can give them extra scrutiny before applying.
Click Apply and EditLint AI places frame-accurate cuts on your timeline with a configurable padding of 1–5 frames before and after each removed word. Your sequence is duplicated automatically before any cuts are applied, so the original is always one click away.
Stop scrubbing through transcripts manually. Let EditLint AI find every filler word in seconds and review them all before making a single cut.