VoiVision AI
tech· VoiVision AI Engineering

Beyond 98% Accuracy: The Five Engineering Stages That Decide Meeting Minutes Quality

Transcription accuracy is only the entry ticket. Whether minutes are actually usable depends on five more stages: speaker separation, summary templates, action-item extraction, passage polishing, and a genuinely offline chain. This article walks through each stage from a buyer's acceptance-testing perspective, with the common failure modes and a checklist you can use directly.


Meeting minutes quality: five engineering stages

Many buyers focus the entire evaluation on a single number: transcription accuracy. 98% sounds impressive. Then three months after go-live, the real feedback arrives — "we still have to assign someone to clean it up."

The problem is conflating two different things. Transcription accuracy determines whether the words are right. Minutes usability determines whether the output can be used as-is. The latter still has five gates to pass.

Note: This article is organised from the buyer's acceptance-testing perspective. Capability statements follow the published product specifications; throughput figures are published values and vary with audio conditions, model size and concurrency strategy — measure them on site.

Stage 1: Speaker separation — who said this?

Minutes without speaker attribution read like this: "I think this plan works... no, the risk is too high... then let's pilot it first." Who agreed, who objected — completely invisible. Minutes like this are worse than none, because they mislead decisions.

Two concepts get conflated here, and acceptance testing must separate them:

CapabilityQuestion answeredPrerequisiteBest suited to
Speaker separationWho said this segment (Speaker 1 / Speaker 2)NoneMeetings with changing attendees
Voiceprint recognitionWhich team member is this (Alice / Bob)Registered voiceprint samplesMeetings with a stable roster

Speaker separation works through integration with the local meeting and audio system, automatically recording and labelling speakers. It outputs role labels — no need to know identities in advance, which suits meetings with external participants.

Voiceprint recognition goes further and attributes speech to real identities: it analyses voiceprints during live or file transcription, distinguishes speakers, and displays role information, with full voiceprint library management (register, rename, delete). It suits recurring meetings with a fixed roster — executive committees, investment committees, standing working groups.

Common failure modes

  • Testing with a single-speaker reading sample: a single voice cannot test separation at all. You need a real recording with interruptions and overlapping speech.
  • Whole-segment misattribution: when two voices sound alike, the model may assign an entire stretch to the wrong person — worse than failing to distinguish them.
  • No voiceprint management: if you cannot rename or delete a voiceprint after staffing changes, the feature is dead within six months.

Acceptance action

Take a real recording with three or more participants and interruptions (not a reading script) and check whether role labels are scrambled or whole segments go to the wrong person. If you are evaluating the voiceprint route, also test registering, renaming and deleting.

Stage 2: Summary generation — did the conclusions survive?

The most common summary mistake is using the wrong template.

Different meeting types need different summary structures. Four templates are provided:

Meeting typeTemplateWhat the summary should contain
Information syncSummary-orientedKey points, progress status
Decision meetingConclusion-orientedFinal decisions, disagreements, open items
Workshop / brainstormMulti-party discussionPositions held, points of contention
Upward reportingReport-orientedConclusion first, data support

Why template mismatch is the number-one killer: run a decision meeting through a summary-oriented template and the model compresses the discussion into one smooth paragraph — flattening the conclusions and disagreements together. Yet in a decision meeting, the conclusions and disagreements are the only parts that matter.

Summary capability must cover both live meetings and audio/video files. That means minutes come out when the meeting ends, and a historical recording dropped in afterwards can be summarised too.

Common failure modes

  • Judging summaries only by fluency: a fluent summary that lost the conclusions is more dangerous than a clumsy one, because it looks fine.
  • Ignoring meeting type: one template for everything.
  • No summary for file transcription: only live meetings are handled, so historical recordings cannot be archived usefully.

Acceptance action

Prepare four recordings of different meeting types, apply the matching template to each, and focus on whether conclusions and disagreements are fully preserved rather than reading for smoothness.

Stage 3: Action-item extraction — does it actually get executed?

What moves things forward after a meeting is the action list. Minutes without action items are a record, not a management tool.

The core chain is: identify action-type statements and the responsible party from the conclusions → output structured task, owner and source passage → push through a standard API into the existing workflow.

There is a hard requirement here that is easy to overlook: action items must retain their original source. If all you get is "finish the research by next Wednesday", nobody knows in what context that was said or who raised it — the accountability trail is broken.

Integration matters just as much. Extracted action items must reach the systems people already use: WeCom, DingTalk or Feishu internally, government OA systems such as Landray and Seeyon in public-sector settings. If action items can only sit inside this one system until someone moves them by hand, the feature is worth close to nothing.

Common failure modes

  • No source reference: cannot be traced back to the words and context.
  • No integration: action items cannot leave the system, which is as good as not extracting them.
  • Over-extraction: treating "let's discuss this again sometime" as an action item.

Acceptance action

Pick three action items and push them through the integration into the system actually in use (WeCom / DingTalk / Feishu / OA). Confirm that owner, task and source passage all survive the handoff.

Stage 4: Passage polishing — is it readable?

Real meetings sound like this: "so, um, we'll push that thing back a bit, because the earlier one hasn't, you know, that yet."

Raw transcription forces the reader to spend real effort reconstructing the meaning. Polishing addresses five problem types:

Problem typeHow it shows up
RedundancyFiller words, repeated lead-ins
Scrambled word orderInversions, parenthetical interruptions
Poor wordingUnclear references, misused terminology
Missing logical connectorsSentences lose their causal links
Slips and repetitionsResidual self-corrections

But there is a constraint that cannot be crossed: polishing must not drop key information or change the speaker's tone.

Why does this matter? Because over-polishing amounts to putting words in the speaker's mouth. In formal minutes — especially in government and legal settings — the record carries a quasi-evidentiary character. If AI polishes a hedged remark into a firm commitment, accountability breaks down.

Common failure modes

  • Over-polishing: turning "we could consider it" into "we agree" changes the nature of the statement entirely.
  • No polishing: handing raw colloquial speech to the reader destroys usability.
  • Tone flattening: converting a diplomatic remark into a blunt one.

Acceptance action

Compare the polished passages against the original recording sentence by sentence, checking two things: was any key information dropped, and was the tone altered.

Stage 5: The offline chain — does compliance actually hold?

However good the first four stages are, if this one collapses, the solution is simply non-compliant in public-sector and regulated settings.

There is a gap most buyers miss here: transcription is easy to run offline, but summary generation is the stage that gets overlooked.

If the architecture is "transcribe to text locally, then send the text to a cloud LLM for summarisation", then:

The audio never leaves the site, but the entire textual content of the meeting does. For MLPS Level 3, financial and government scenarios, that directly contradicts a "no data leaves the site" commitment.

So acceptance must include one specific action: re-run the whole pipeline with the network fully disconnected, watching the summary stage in particular.

A genuinely closed loop runs both the recognition and the summary models locally, with no cloud licensing required — and this touches the licence-check channel, which should also be verified air-gapped (some solutions transcribe locally but need network access for licence validation, which will likewise undermine the "no egress" promise during assessment).

For the underlying compliance requirements, see our breakdown of MLPS Level 3 requirements for meeting recording and minutes systems.

Common failure modes

  • Verifying transcription offline but not summaries: the most common hole.
  • Ignoring the licence-check channel: transcription runs locally, but every startup calls home.
  • Silent degradation instead of an error: the system quietly falls back to a reduced flow with no user-visible indication.

Acceptance action

Re-run the full pipeline air-gapped. If the summary stage errors out or degrades sharply, it depends on the cloud and the solution is non-compliant.

Appendix: A checklist you can use as-is

The five stages compressed into one table — tick each during acceptance:

#ItemMethodPass criterion
1Speaker separationReal multi-party recordingNo whole-segment misattribution
2Voiceprint recognition (if selected)Register / rename / deleteAll three actions work
3Summary template matchTest all four meeting typesConclusions and disagreements preserved
4File-transcription summaryImport a historical recordingSummary is produced
5Action-item extractionPush three items end to endTask / owner / source all present
6System integrationConnect to target OA / IMItems enter the existing workflow
7Passage polishingSentence-level comparison with audioNo information loss, no tone change
8Output formatsInspect exportsText / graphic / action items / mind map
9Offline chainRe-run air-gappedEverything works, summaries included
10ThroughputBatch import and time itAround 10:1 (measure it)

Item 9 is the one most often skipped, and the one most likely to fail. If you can only keep one acceptance action, keep this one.

For scenario-specific rollout paths, see our government intranet and financial compliance playbooks.

One honest closing note

None of these five stages can be replaced by transcription accuracy. Accurate transcription is the entry ticket, not the destination.

If a vendor only talks about accuracy numbers and never about speaker attribution, summary templates, action-item provenance or offline summarisation, the system will most likely end up being used as a transcription tool — producing raw drafts that someone still has to clean up by hand.

That is not the same thing as buying a meeting-minutes system.


Further reading: Running multilingual meetings: offline translation and synced subtitles | OA integration in practice | Choosing an on-premise ASR solution

FAQ

Q: Beyond transcription accuracy, what else should be tested when accepting a meeting-minutes system?

A: Accuracy only determines whether the words are right, not whether the minutes are usable. Five more things matter: whether speakers are correctly attributed, whether the summary template matches the meeting type, whether action items carry an owner and a task, whether colloquial speech is polished, and whether the whole chain is genuinely offline. If any one of these fails, people rework the minutes by hand — and high accuracy does not save you.

Q: Are speaker separation and voiceprint recognition the same thing?

A: No. Speaker separation answers "who said this segment" and outputs role labels such as Speaker 1 and Speaker 2, without needing to know identities in advance. Voiceprint recognition answers "which team member is this", and requires pre-registered voiceprint samples and a voiceprint library. The former labels speakers automatically through integration with the local meeting or audio system; the latter suits meetings with a stable roster and supports registering, renaming and deleting voiceprints.

Q: What types of summary templates exist, and how do I choose?

A: There are four common types. Summary-oriented suits information-sync meetings, conclusion-oriented suits decision meetings, multi-party-discussion suits brainstorming and workshops, and report-oriented suits upward reporting. Poor summary quality is most often caused not by a weak model but by a template that mismatches the meeting type — run a decision meeting through a summary template and the conclusions and disagreements get flattened away.

Q: How are action items extracted from a conversation?

A: The system first identifies action-type statements and the responsible party from the conclusions, then outputs structured items with task, owner and source passage. Keeping the original source reference is essential, otherwise an action item cannot be traced. Results are typically pushed through a standard API or integrated with WeCom, DingTalk, Feishu, or government OA systems such as Landray and Seeyon so they enter the existing workflow directly.

Q: How is colloquial speech handled in the minutes?

A: The polishing stage addresses five problem types: redundancy, scrambled word order, poor wording, missing logical connectors, and slips or repetitions. The hard constraint is that it must not drop key information or change the speaker's tone — over-polishing amounts to putting words in the speaker's mouth, which is a serious problem in formal minutes.

Q: What output formats are supported?

A: A Word record is generated automatically at the end of a meeting, alongside text minutes, graphic minutes, action items and a mind map. All of them support web viewing, editing and download, and can be archived together with the audio, translation and summary.

Q: Can the entire minutes pipeline run offline?

A: Yes, but you must verify it stage by stage during acceptance. Transcription is relatively easy to run offline; what gets overlooked is summary generation. If summaries are produced through a cloud LLM endpoint, then the audio never leaves the site but the text does — and the compliance promise still fails. A genuinely closed loop runs both the recognition and the summary models locally, with no cloud licensing required.

Q: What is the file transcription throughput?

A: The published figure is 10:1 — roughly one minute to transcribe ten minutes of audio or video, with batch import for MP3, MP4, MOV, MKV, WAV and AAC. Actual time varies with audio conditions, model size and concurrency strategy, so treat it as something to measure on site.

#Meeting minutes#Speaker diarization#Smart summary#Action items#On-premise deployment#Acceptance checklist

Book a Personalized Demo

Tell us your meeting scenario and compliance needs — get a tailored plan.

Book now
Live Chat