Demonstration 04 · Vocis

Clinical Dictation

One urgent-care encounter, dictated in four SOAP passes into a live microphone. The engine segments the speech, recognizes it on-device, retrieves the exam templates being described, and fills their fields as the words land — including a spoken correction.

What you are looking at

The session screen holds four panels, and everything in the capture happens inside them.

Speak

continuous dictation · one open session

→

Recognize

speech → text · on-device · no network

→

Retrieve

spoken phrase → template from the library

→

Extract

recognized words → template fields

→

Assemble

template text + extracted values → note

Unedited capture · live dictation → templates retrieved → fields filled → note assembled
The Verbatim panel of recognized speech beside the assembled note preview, with extracted values highlighted in the note
Left, the recognized speech as it was said. Right, the note built from it — the highlighted spans are values extracted from that speech, and the sentences around them are the provider's template text. The fills panel continues below the crop.
The Laceration Repair template block: irrigation, anaesthetic, closure and dressing options, with the dictated choice in each group highlighted
One filled template. Each field carries the options the provider authored; the value taken from the dictation is filled and the alternatives stay visible beside it. “Patient tolerated the procedure well” is template text, not something the microphone heard.
The Vitals template block with heart rate filled as 82
Before. Heart rate was dictated as 82 and filled.
The same Vitals block two seconds later, heart rate now 84
After. “Actually heart rate 84” rewrites the filled field in place. Same field, two seconds apart.

Templates are retrieved by phrase

Each template in the library carries the phrase that calls it. Saying the phrase mid-dictation pulls that template into the session, and the surrounding speech fills it.

The library in this capture holds ten templates across the objective and plan sections.

Everything in the video is a real capture — a single unedited take of live dictation against the development bench, running the same engine, models, and decode configuration deployed on the device. Recognition, template retrieval, and field extraction are local; no audio or text leaves the machine. The encounter is a scripted vignette; no patient data appears in the capture.

The note has exactly two sources: template text the provider authored and reviewed in advance, and values extracted from the recognized speech. Highlights and chips mark which values were extracted and from what. No part of the note is generated by a language model.