ASR Recognition Optimization
Dialect × Full-Pipeline Engineering
Built on four stages — dialect baseline, dialect specialization, on-device engineering, and LLM-based correction — to push enterprise meeting ASR accuracy to industry-leading levels: multi-dialect accuracy ≥ 90%, with an extra 5–10% overall gain on top of engineering.
Key Challenges
- !Inconsistent accuracy across 10+ dialects
- !Complex acoustics: noise, echo, far-field
- !Industry jargon / names / abbreviations hard to cover generically
- !Output must be readable, searchable, post-processable
Our Approach
- ✓Dialect baseline: covers 10+ dialects (Mandarin, Cantonese, Sichuanese, Wu, Min, Northeastern, etc.); dedicated training + independent decoder + dialect LM; measured multi-dialect accuracy ≥ 90%
- ✓Dialect specialization: dialect embeddings + multi-task joint training closes the gap to Mandarin; generic ASR + dialect LM hot-pluggable, new dialects without retraining; per-customer industry-term customization
- ✓On-device engineering: front-end AGC / AEC / ANC; OCR screen-content awareness → dynamic domain-hotword injection into ASR decoder; full pipeline on-device, E2E latency < 1s, data never leaves the premises
- ✓LLM-based correction: LLM context-aware correction (error fix, smoothing, punctuation, number normalization); RAG over enterprise terms / industry knowledge / historical minutes; +5–10% overall accuracy on top of engineering
Recommended Products

VV10 Meeting Server
10 channels · edge NPU
The VV10 is a department-grade 10-channel server with an edge NPU at 1000 TOPS. Full local processing, zero data egress, unified web console — built for R&D centers and financial institutions.
Learn more →Hardware
VV50 Enterprise Server
50 channels · full private deployment
The VV50 is an enterprise 50-channel engine with CPU+NPU/GPU heterogeneous architecture and 4–16TB RAID. Supports government intranet and air-gapped deployment, meeting MLPS Level 3 and classified requirements.
Learn more →SoftwareSpeech Engine (Software Only)
Conformer E2E · real-time on CPU · fully offline
Built on NEU NLP Lab's in-house engine with Conformer + CTC/Attention E2E architecture. INT8/INT4 quantization, pruning, knowledge distillation, and operator fusion enable real-time ASR on CPU alone. LM compute reduced 70%+, accuracy improved 10–20% over traditional approaches. Native support for NVIDIA / Ascend / Cambricon / Hygon, Docker/K8s one-click deploy.
Learn more →Hardware
V06 Voice Card
Card-sized · Carry anywhere
The V06 is a business-card-sized portable recorder. One-tap recording, dual-mic capture, and automatic sync to cloud or on-prem server turn personal meetings into searchable knowledge.
Learn more →OA Integration
VoiVision deeply integrates with leading office platforms — minutes straight into your workflow.
- ✓Higher ASR accuracy directly improves OA-synced minutes
- ✓More accurate extraction of decisions / todos / owners
- ✓Integration with WeCom / DingTalk / Feishu
Book a Personalized Demo
Tell us your meeting scenario and compliance needs — get a tailored plan.
