VoiVision AI

ASR Recognition Optimization

Dialect × Full-Pipeline Engineering

Built on four stages — dialect baseline, dialect specialization, on-device engineering, and LLM-based correction — to push enterprise meeting ASR accuracy to industry-leading levels: multi-dialect accuracy ≥ 90%, with an extra 5–10% overall gain on top of engineering.

Key Challenges

  • !Inconsistent accuracy across 10+ dialects
  • !Complex acoustics: noise, echo, far-field
  • !Industry jargon / names / abbreviations hard to cover generically
  • !Output must be readable, searchable, post-processable

Our Approach

  • Dialect baseline: covers 10+ dialects (Mandarin, Cantonese, Sichuanese, Wu, Min, Northeastern, etc.); dedicated training + independent decoder + dialect LM; measured multi-dialect accuracy ≥ 90%
  • Dialect specialization: dialect embeddings + multi-task joint training closes the gap to Mandarin; generic ASR + dialect LM hot-pluggable, new dialects without retraining; per-customer industry-term customization
  • On-device engineering: front-end AGC / AEC / ANC; OCR screen-content awareness → dynamic domain-hotword injection into ASR decoder; full pipeline on-device, E2E latency < 1s, data never leaves the premises
  • LLM-based correction: LLM context-aware correction (error fix, smoothing, punctuation, number normalization); RAG over enterprise terms / industry knowledge / historical minutes; +5–10% overall accuracy on top of engineering

OA Integration

VoiVision deeply integrates with leading office platforms — minutes straight into your workflow.

  • Higher ASR accuracy directly improves OA-synced minutes
  • More accurate extraction of decisions / todos / owners
  • Integration with WeCom / DingTalk / Feishu

Book a Personalized Demo

Tell us your meeting scenario and compliance needs — get a tailored plan.

Book now
Live Chat