Voice Intelligence

Voice Intelligence turns spoken work into workspace context. Record a conversation inside Dawn, transcribe an attached audio or video file, or dictate a prompt, and the result becomes a transcript file the agent can read, summarise, search, and act on.

What Voice Intelligence covers

Voice Intelligence is the umbrella for Dawn’s spoken-work capabilities:

  • Record: capture a meeting, call, or voice note directly inside Dawn.
  • Transcribe: turn an attached audio or video file into a transcript file in the same thread.
  • Voice input: dictate a prompt in the composer instead of typing.
  • Speaker separation: identify who said what across multi-person recordings.

All paths produce the same kind of output, a transcript file in the conversation, so anything you can do with one, you can do with the others.

Recording a Conversation in Dawn

The Record control in the Dawn composer captures audio directly inside a thread. Use it when there is no existing media file to attach: a meeting happening on the device, a voice note you want to keep with the conversation, a customer call, or any spoken work you want transcribed right away.

When you record:

  • Audio is captured through the browser microphone (or another audio source the user has granted access to).
  • When you stop recording, the audio is uploaded into the current thread as a workspace file.
  • Voice Intelligence transcribes the file in the background and attaches the transcript when it is ready.
  • Speakers are separated automatically (see below).

Recording keeps the workflow inside Dawn. You do not need to start a separate recorder, save the file, and upload it back into a thread. The audio file and its transcript both stay with the conversation.

After the transcript is attached, ask Dawn for the next thing you need: a summary, action items, open questions, a customer-ready recap, a follow-up email, or a ticket.

Transcribing audio and video files

To transcribe an existing recording, attach an audio or video file to a thread and ask Dawn to transcribe it.

  • Dawn uses ElevenLabs speech-to-text, optimised for real recordings rather than studio-quality audio.
  • Longer files are split into smaller transcription chunks so a single large upload is less likely to fail because it exceeds a per-request limit.
  • The transcript comes back as a downloadable workspace file attached to the same thread.

Once the transcript is in the thread, it becomes normal Dawn context. The agent can summarise it, extract structured information, draft follow-up work, or compare it with other workspace knowledge.

Voice input in the composer

The composer also supports voice input for dictating prompts. Use the voice control, speak the prompt, and Dawn turns the audio into text for the conversation.

Voice input is for moments where speed matters or where the prompt is naturally spoken: “summarise what changed here”, “turn this into a bug report”, “draft a follow-up plan from this recording.” It is not meant to replace writing for every prompt.

Speaker separation and transcript quality

Transcripts produced by Voice Intelligence separate speakers automatically. The output identifies different voices in the recording, so a long meeting can be scanned quickly, a decision can be traced to the moment it was made, and one person’s contributions can be pulled out for follow-up.

Transcripts are not perfect (no transcription is), but they stay usable in real recording conditions: overlapping voices, side conversations, background noise, a single bad microphone, a remote participant on a weak connection. Severely degraded source audio still requires judgement when reading the result.

Limits and usage

Voice Intelligence has practical limits:

  • Very long recordings or very large files may still fail. If a recording is unusually long, splitting it before upload can help.
  • Severe background noise or very low audio quality can produce imperfect output.
  • Recording in Dawn Web requires the browser to grant microphone (and other audio source) permission.

Transcription consumes provider capacity, so Dawn meters it against the same workspace credits as text generation. A minute of transcribed audio is billed at the same credit rate as the equivalent provider call, alongside everything else the workspace runs.

Where transcripts go

Every Voice Intelligence path produces a workspace file:

  • Transcripts can be downloaded, previewed, mentioned later, or used as context for the next step.
  • They appear in the Files Browser and travel with the conversation.
  • They can be passed into a Prompt Command, a Skill, or a Workflow the same way any other file can.

Voice Intelligence is the bridge between spoken work and the rest of Dawn. The recording is the start; the transcript is the input the agent works from.