Making a student film fully accessible on a real deadline

Cinema, Radio and Television Beginner 180 min Near zero if the campus licenses a captioning vendor; outsourced audio description runs $15–75 per minute

The situation

Student films screen publicly and go on the department site, which puts them under the same accessibility obligations as everything else the college publishes. What actually happens is a student takes YouTube's auto-captions and ships a file where every proper noun in their own film is misspelled. Audio description usually does not happen at all, because nobody has taught it.

Steps

  1. Export a clean dialogue stem and lock picture first

    Premiere, Resolve or Avid

    Generate the transcript from a dialogue-only stem, not the final mix. Music and effects beds are the single largest cause of garbage machine transcription. Lock picture first — every timecode downstream breaks if the edit changes.

    What you only learn by doing it: This is the argument for why accessibility belongs in post rather than after post. If picture is not locked you will do the caption pass twice, and the second pass is not faster than the first. Students learn this exactly once.

  2. Run recognition, then correct against a names list you wrote yourself

    Descript

    Before correcting, the student writes out every proper noun, place name, piece of slang and technical term in their own film, then corrects against that list. Machine captions are rarely accurate enough to meet accessibility requirements on their own.

    What you only learn by doing it: 99% accuracy sounds generous until you do the arithmetic: at 150 words per minute that still permits about 1.5 errors per minute, and the errors cluster on names, negations and numbers. Have students count errors against word count — the ones who “feel like it's basically fine” are usually around 94%.

  3. Fix what recognition structurally cannot do

    Subtitle Edit, Aegisub, or Premiere's caption panel

    Add speaker identification, bracketed non-speech information that carries meaning, and break lines at grammatical boundaries at 160–180 words per minute, maximum two lines on screen. Machine output gives you words; this gives you a track someone can read while watching.

    What you only learn by doing it: The non-speech cues are a directing lesson in disguise. Students discover that a sound they designed carries plot information a deaf viewer only gets if they write it down — and that half their atmospheric sound work carries nothing and should not be captioned.

  4. Write audio description into the gaps, to the Description Key standard

    DCMP Description Key (free), YouDescribe for practice

    Find every dialogue gap long enough to hold description and write only what is needed: present tense, active voice, third person, general to specific, describe what you see without interpretation or comment.

    What you only learn by doing it: The universal mistake is interpretation. Students write “she looks devastated,” which tells a blind viewer what to feel and steals the performance from them. The correct line is “she turns toward the window; her jaw tightens.” Equally universal: trying to fill every silence. Good describers use silence.

  5. Voice and mix description as an alternate track

    Audition or Fairlight; a neural voice or a student VO

    Produce the described version as a separate deliverable rather than a re-edit. Duck the programme audio under description only where necessary and by the minimum amount.

    What you only learn by doing it: Record at the film's pace, not at reading pace. Students read description 20–30% faster than the film moves, which makes a contemplative short feel like an instructional video. And pick one synthetic voice for the whole film — switching reads as a new character to someone who cannot see the screen.

  6. Package, then verify with the monitor physically off

    SRT/WebVTT sidecars; NVDA or VoiceOver, both free

    Upload captions as a sidecar rather than burning them in, confirm the described version is discoverable, and have the student navigate to their own film using only a screen reader.

    What you only learn by doing it: The screen-reader test is where it usually falls apart, and never in the film — it is the page. The described version is labelled “version 2 final,” the play button has no accessible name, or the embed traps keyboard focus. Doing this with the monitor physically off takes fifteen minutes and is the most persuasive fifteen minutes in the course.

Where this breaks down

The biggest trap is treating this as compliance paperwork instead of craft. Grade it pass/fail on “does a caption file exist” and you will get 94%-accurate captions and description that narrates the plot.

AI audio description is not mature enough to run unsupervised. The vision models that write first drafts confuse who is speaking, miss off-screen action, and will confidently describe a character's emotional state — the one thing the professional standard forbids.

Check whether your campus captioning agreement covers department media or only LMS content before a student's thesis film sits on a vendor's server.

None of this substitutes for accessible design. A film relying on unreadable on-screen text or on colour alone to distinguish characters cannot be fixed at the caption stage; that fix belongs in pre-production.

Provenance: the accuracy thresholds (99% high-priority, 97% medium) and the position that AI audio description requires human review come from the University of Washington's accessibility course-content report. The description writing rules are DCMP's published Description Key, the US professional standard for educational media.