Clinical Dictation
One urgent-care encounter, dictated in four SOAP passes into a live microphone. The engine segments the speech, recognizes it on-device, retrieves the exam templates being described, and fills their fields as the words land — including a spoken correction.
What you are looking at
The session screen holds four panels, and everything in the capture happens inside them.
- Verbatim — the recognized speech, unedited, grouped under the SOAP section being dictated.
- Template fills — each retrieved template with its fields; the highlighted values and green chips are what was extracted from the speech.
- Preview (note) — the assembled note, as one block or split by SOAP section, with Send and Copy.
- Template library — the provider's templates, each listed with the phrase that calls it.
Speak
continuous dictation · one open session
Recognize
speech → text · on-device · no network
Retrieve
spoken phrase → template from the library
Extract
recognized words → template fields
Assemble
template text + extracted values → note
Templates are retrieved by phrase
Each template in the library carries the phrase that calls it. Saying the phrase mid-dictation pulls that template into the session, and the surrounding speech fills it.
- Vitals and General Appearance · “insert vitals”
- Cardiac Exam · “insert cardiac”
- Respiratory Exam · “insert respiratory”
- Laceration Repair · “insert laceration”
The library in this capture holds ten templates across the objective and plan sections.
Everything in the video is a real capture — a single unedited take of live dictation against the development bench, running the same engine, models, and decode configuration deployed on the device. Recognition, template retrieval, and field extraction are local; no audio or text leaves the machine. The encounter is a scripted vignette; no patient data appears in the capture.
The note has exactly two sources: template text the provider authored and reviewed in advance, and values extracted from the recognized speech. Highlights and chips mark which values were extracted and from what. No part of the note is generated by a language model.