VoiVision AI

Tech & Ecosystem

Four hardware lines · one engine · full coverage

AI Speech Engine

Built on NEU NLP Lab's in-house research, we re-architected the ASR pipeline from the traditional AM+LM+Decoder cascade into an end-to-end (E2E) neural model. Using Conformer + CTC/Attention hybrid modeling, combined with INT8/INT4 quantization, pruning, knowledge distillation, structural re-parameterization, and operator fusion, the engine achieves real-time ASR on CPU alone — no GPU required. Lightweight language modeling with adaptive context compression reduces LM compute by over 70%, while recognition accuracy improves 10–20% over traditional approaches.

  • E2E model: Conformer + CTC/Attention hybrid, eliminating acoustic/language modeling error accumulation
  • INT8/INT4 quantization & pruning: drastically reduced params for real-time CPU inference
  • Knowledge distillation & structural re-parameterization: compact model with high generalization
  • Operator fusion & cache optimization: improved CPU matrix efficiency
  • Lightweight LM: internal LM + adaptive context compression, 70% compute reduction
  • Context-enhanced training & constrained self-attention: high accuracy in low-resource conditions

Accelerator Support

  • NVIDIA: T4 / L4 / A10 / A100
  • Ascend: 310P / 910B
  • Cambricon: MLU370 / MLU590
  • Hygon DCU (CPU-only also works)

Full OA Integration

  • WeCom
  • DingTalk
  • Feishu
  • Landray
  • Seeyon

Deployment

  • Docker / K8s one-click
  • Bare-metal
  • Gov cloud / private cloud / SIEM
  • RESTful API / WebSocket

ASR × OCR Multimodal Fusion

Screen-content awareness × dynamic ASR enhancement

VoiVision links ASR with on-device OCR, using the visual content of the screen (PPT / docs / whiteboard) to boost speech recognition. The BJ66+WCB06 captures the screen via wired + wireless channels; an on-device large-model OCR reads the text frame by frame; the domain knowledge base auto-matches and injects domain terms into the ASR decoder — lifting accuracy in terminology-dense scenes like lectures, consults, and roadshows.

Four-step closed loop

  1. ① Capture: BJ66+WCB06 grabs the screen byte stream (HDMI-in + wireless)
  2. ② On-device OCR: frame-by-frame parsing of PPT / docs / whiteboard text
  3. ③ KB match: 200+ discipline hotword dictionaries auto-align
  4. ④ Hotword injection: domain terms dynamically injected into ASR decoder

Key advantages

  • Fully local private deploy — data never leaves room
  • Hot-swappable domain KBs (education / healthcare / finance / law)
  • New domains need no ASR retraining — hotwords take effect instantly
  • Hotword injection < 500ms, frame-level OCR < 500ms
Generic ASR term accuracy baseline ~72% → ASR×OCR fusion lifts to 94%+ (NEU NLP Lab lecture-hour test)

OA Integration

Seamless with mainstream office platforms · minutes into your workflow

VoiVision deeply integrates with leading office & collaboration platforms — automatic minutes sync, smart todo extraction, one-terminal multi-platform management — embedding meeting intelligence directly into existing enterprise workflows.

WeCom

  • Calendar sync · contacts
  • Approval flow · app push
  • Auto minutes sync · todo

DingTalk

  • Calendar sync · meeting mgmt
  • DING notify · approval
  • Todo · multi-dimension sheets

Feishu

  • Calendar sync · docs
  • Msg notify · meeting
  • Approval · enterprise wiki

Landray / Seeyon

  • Eco integration · SSO
  • Workflow engine link
  • Custom self-built OA

Book a Personalized Demo

Tell us your meeting scenario and compliance needs — get a tailored plan.

Book now
Live Chat