Voice Synthesis (Meetings)
Voice Synthesis is the engine that turns meeting audio into structured minutes. It records live or accepts an uploaded file, transcribes with speaker diarization, and returns attendees with talk-time share, a topic summary, action items, decisions and risks. Every extracted line carries an audio anchor that plays the exact moment it was discussed.
What does Voice Synthesis do?
Voice Synthesis turns meeting audio into structured minutes. It records live or accepts an uploaded file, runs ASR (Gemini 2.5 Flash) with speaker diarization, then composes the output: attendees with talk-time share, topic-grouped summary, action items with assignee and due date, decisions with rationale, and risks with severity.
Every extracted line carries an audio anchor — a clickable ▶ MM:SS badge that scrubs the recording to the exact moment the line was discussed. No more arguing about what someone said three weeks ago at the OAC meeting.
Action items auto-promote into the Tasks tracker with source = meeting. The original audio is retained 18 months by default and lives in the project record for future reference.
When should I use Voice Synthesis?
Use Voice Synthesis after an OAC meeting or a pre-construction kickoff, to keep a receipt of what an architect committed to on a coordination call, to turn site walk audio into punch items, RFIs and CO candidates, or to document decisions on a vendor negotiation call.
- OAC meeting just wrapped — upload the audio and have minutes in 2 minutes.
- Pre-construction kickoff with 12 people and a packed agenda; minutes alone won't capture every nuance, but anchored minutes will.
- Coordination call where an architect committed to a response date and the GC needs the receipt.
- Site walk audio — turn the verbal walk-through into structured punch items, RFIs, and CO candidates.
- Vendor negotiation call where decisions and concessions need to be precisely documented.
If the meeting is a site walk, the audio also feeds RFI Tracker, Punch List, and Change Order Review. One recording, multiple engines, all anchored.
What do I upload to Voice Synthesis?
Audio file (MP3, M4A, WAV, AAC, FLAC, OGG) or a live recording captured from the browser. Up to ~3 hours per file. Best quality with a single-channel speakerphone source or a dedicated meeting mic. Background noise (HVAC, generator) is tolerated but cuts diarization accuracy.
Optional: attendee roster (improves diarization name attribution), prior meeting minutes (for context continuity), and the agenda (helps the engine match audio segments to topics).
How do I run Voice Synthesis?
To run Voice Synthesis, open it, choose a capture mode, provide context, run the synthesis, scrub the audio with the anchors, then review the action items, decisions and risks and distribute the minutes.
Open the engine
Sidebar → Engines → Voice Synthesis (Meetings).
Choose capture mode
Record live (browser mic) or Upload audio (paste in a file).
Provide context
Optional: attendee names, agenda, prior meeting reference. The engine still works without these but accuracy on names and topic mapping is much better with them.
Run synthesis
30–90 seconds. The longer the audio, the longer the runtime — budget 10–30s per 10 minutes of recording.
Scrub the audio with anchors
Click any ▶ MM:SS badge in the minutes to jump to that exact moment in the recording. Validates each extraction in seconds.
Review action items
Each action item has assignee, due date, and source anchor. Edit if needed, then promote to the Tasks tracker.
Review decisions and risks
Decisions log captures every "we agreed to X". Risks log captures every "if Y happens, we have a problem".
Distribute
Export PDF or share the live link. Attendees get clickable audio anchors too if they have viewer access.
How do I read the Voice Synthesis results?
The result is an attendee list with talk-time, a topic-grouped summary, an action items list, a decisions log, a risks log and audio anchors.
Diarized speakers with percentage of total talk time. Tells you who dominated and who didn't speak.
Discussion summarized in topic blocks rather than chronological transcript. Each topic carries its anchor range.
Numbered list with assignee, due date, and audio anchor. Auto-promotes to the Tasks tracker on stamp.
Every committed decision with the rationale and the moment it was made. Defensive documentation gold.
Stated risks with severity (high / medium / low) and the affected scope. Feeds into project risk register.
Clickable ▶ MM:SS markers on every extracted line. Scrubs the audio to that moment in the recording.
Every control, explained
Record liveCaptures from browser mic. Pauses and resumes supported.
Upload audioDrop in a file. MP3/M4A/WAV/AAC/FLAC/OGG up to 3 hours.
Add attendeeNames attendees so diarization can attribute speakers correctly.
Run synthesisConsumes one run. 30–90 seconds.
Play anchorClick any ▶ MM:SS badge to scrub to the moment in the recording.
Promote to TasksPushes the action items into the Tasks tracker with source = meeting and the audio anchor preserved.
Export PDFFormatted minutes with topic blocks, decisions, action items, risks. Anchors render as timestamps.
Share linkRead-only share with audio anchors live for the recipient (subject to project permission).
Sources
- 18 U.S.C. 2511: Interception and disclosure of wire, oral, or electronic communications prohibitedU.S. Government Publishing Office, govinfoThe federal wiretap law, including the exception for a recording made with the consent of a party to the conversation.
- Federal Rules of EvidenceUnited States CourtsThe official federal rules on evidence, including Rule 1002 on needing the original to prove what a recording contains.