On-Premise ASR Selection: Open-Source Self-Hosting vs Cloud Private Packs vs Commercial Engines
Compare the three routes across compliance red lines, concurrency and latency, language and dialect coverage, accuracy and three-year TCO — plus the five real barriers to self-hosting Whisper, what MLPS Level 3 actually demands during selection, and a ready-to-use eight-question vendor checklist.

No concepts here — just one question: which route fits your scenario.
This breaks down on-premise ASR selection work done on government, finance and energy customer sites into a decision table, a three-year cost model, and the real barriers to open-source self-hosting.
1. The Answer Up Front: One Decision Table
If you have three minutes, this table is enough.
| Your situation | Recommended route | Why |
|---|---|---|
| You have ML engineers, scenario tolerates error (internal stand-ups, subtitles) | Open-source self-hosting (faster-whisper / Paraformer) | Zero license cost, good enough |
| Already committed to one cloud vendor, audio egress is acceptable | Cloud API | Fastest to launch, do not overbuild |
| Compliance requirements but not classified, low concurrency | Cloud vendor private pack | Fast delivery, familiar ecosystem |
| Data residency is a hard red line / Xinchuang / classified / many concurrent streams | Commercial on-premise engine | The other routes cannot reach it |
| Moving to domestic accelerators (Ascend / Cambricon / Hygon) | Commercial engine with native support | Porting it yourself is very expensive |
One priority people routinely invert: most buyers compare accuracy first, but what actually kills a project is usually whether data may leave the domain.
Two accuracy points is survivable; data leaving the domain means the project never gets approved. So the correct order is:
- Compliance red line (data residency, Xinchuang) → eliminates half the options outright
- Concurrency and latency (how many streams, real-time or not) → sets hardware specs
- Language and dialect → sets engine capability
- Accuracy → compare only among options that passed 1–3
- Total cost → last, but do not forget to count people
2. On-Premise vs Cloud: Three Ledgers, Not Just Security
The latency ledger
A cloud API's latency is: upload + queue + inference + return.
A one-hour meeting recording (16 kHz mono, roughly 115 MB) takes about 10 seconds to upload over a 100 Mbps intranet; over the public internet to a cloud endpoint you add jitter on top. For batch transcription this does not matter. For live captions it is fatal.
On-premise audio never leaves the NIC, so local inference RTF is the entire latency budget.
Rule of thumb: if you need live captions or live minutes, the cloud latency structure is structurally disadvantaged — prefer on-premise.
The cost ledger (the one most often miscalculated)
Comparing "cloud per-hour pricing" directly against "one-time server purchase" is wrong. Three-year TCO should include:
Scenario: 50 hours of meeting audio per day, business days, three-year horizon.
| Cost item | Cloud API | Open-source self-hosting | Commercial on-premise engine |
|---|---|---|---|
| Transcription fees | Per-hour billing, tens of thousands over 3 years | 0 | 0 |
| Hardware | 0 | Single-card server (one-off) | Included |
| Software license | 0 | 0 (open source) | One-time license (by concurrency tier) |
| Implementation effort | 0 | 1–2 person-months | Included |
| 3-year operations effort | 0 | 0.3–0.5 FTE | Mostly vendor-supported |
| Failure fallback | Vendor SLA | On you | Contractual SLA |
You might look at this and ask: isn't self-hosting the cheapest?
No. Self-hosting saves license fees and costs people. And that people cost is fixed — models need upgrading, operators need adapting, and new customer hardware means reconverting models. The 0.3–0.5 FTE above is conservative; with domestic accelerators and air-gapped deployment it goes higher.
In one line: for low-volume general scenarios without compliance constraints, cloud APIs usually have the lower TCO. The economic tipping point for on-premise arrives when either concurrency pressure is high or compliance forbids the cloud.
The risk ledger
Hard to quantify, but often decisive:
- Connectivity risk — on an isolated intranet, cloud options are off the table immediately
- Vendor lock-in — switching cloud APIs means redoing integration; switching on-premise engines means redeploying
- Audit responsibility — when something breaks, with cloud it is "the vendor's problem"; self-hosted, it is yours
3. The Real Barriers to Open-Source Self-Hosting
Self-hosting Whisper is a heavily searched topic, which tells you plenty of teams are considering it. We have done it. Here is the honest version.
Getting it running is not hard
# faster-whisper is the default choice today
pip install faster-whisper
# 8GB VRAM can run large-v3-turbo with INT8 quantization
python -c "
from faster_whisper import WhisperModel
m = WhisperModel('large-v3-turbo', device='cuda', compute_type='int8_float16')
segments, info = m.transcribe('meeting.wav', language='zh')
print(f'language {info.language}, probability {info.language_probability:.2f}')
for s in segments:
print(f'[{s.start:6.1f}s -> {s.end:6.1f}s] {s.text}')
"
You can have a demo running in ten minutes. Getting from demo to production is the hard part.
Five barriers you only meet in production
1. No hotwords
The most painful one in enterprise scenarios. Company names, product codenames, people's names, project abbreviations — Whisper will consistently write them as similar-sounding characters. initial_prompt gives only weak guidance, decays over long audio, and enforces nothing.
If your customer's meetings are full of proper nouns, you will not avoid this.
2. No built-in speaker separation
Whisper outputs text, not "who said it." Producing meeting minutes means adding pyannote or a NeMo speaker-separation module separately and aligning timestamps yourself. One more model, one more slice of VRAM, one more thing that can break.
3. Streaming is a structural weakness
Whisper's architecture was not designed for streaming. The common workaround is chunked streaming (feed 5–10 second windows), which costs boundary errors at split points — a sentence gets cut in half, each half is transcribed independently, and the seam produces wrong characters. Poor experience for live captions.
4. Essentially no dialects
large-v3 has a yue (Cantonese) token, but quality is nowhere near production-ready. Wu, Min Nan, Sichuanese, Central Plains Mandarin — none are supported by the official models.
5. Domestic accelerators require porting it yourself
Whisper's ONNX export plus Ascend ATC conversion does work, but you will meet: unsupported opset versions, dynamic-axis configuration, silent unsupported-operator failures, and models needing reconversion after version upgrades. See the Ascend 910B private ASR walkthrough for details.
When open-source self-hosting is the right call
Being fair about where it fits:
- You have ML engineers who can modify models
- The scenario tolerates error (internal meeting notes, video subtitles)
- General Mandarin, without heavy proper-noun loads
- No real-time requirement
- No domestic-accelerator mandate
If all five hold, self-hosting is excellent value. If two or more do not, run the people-cost numbers carefully.
4. Capability Comparison Across the Three Routes
| Dimension | Open-source self-hosting | Cloud vendor private pack | Commercial on-premise engine |
|---|---|---|---|
| License cost | 0 | Medium | High |
| Implementation effort | High (1–2 person-months) | Medium (1–2 weeks) | Low (vendor-delivered) |
| Chinese accuracy (meetings) | Medium | Medium-high | High |
| Hotwords / proper nouns | None | Yes | Yes |
| Speaker separation | Separate module required | Yes | Built in |
| Real-time streaming | Weak | Yes | Yes |
| Dialect support | Essentially none | Limited | 22 |
| Domestic accelerator support | Port it yourself | Partial | Native |
| Air-gapped operation | Package it yourself | Depends on vendor | Supported |
| MLPS / classified support | Your own paperwork | Partial | Documentation provided |
| Operations responsibility | Entirely yours | Vendor | Vendor + SLA |
The two rows most often underestimated are implementation effort and operations responsibility. They never appear on a quote, but they consume the team continuously.
5. What MLPS Level 3 Demands During Selection
This is the most common question from government and finance customers. Of the MLPS Level 3 requirements (GB/T 22239-2019) for meeting recording and minutes systems, these are the ones that shape selection:
| Control | MLPS L3 requirement | The question to ask |
|---|---|---|
| Data residency | Audio and text stored locally, no public upload | "Does any part of transcription make a public network call?" — including license checks, model downloads and telemetry |
| Access control | Identity verification, role-based permissions | "Can it integrate with our existing LDAP / OAuth?" |
| Security auditing | Traceable operations, retained logs | "Who can export minutes? Is export logged? How long are logs kept?" |
| Transport encryption | Encrypted internal communication | "Is TLS enforced on the intranet too, or only on the public side?" |
| Residual information protection | Deleted data must be unrecoverable | "After deleting a meeting, are model caches and temp files cleared too?" |
The first and the last are the ones most often missed.
Many solutions claim "no data egress," yet license validation reaches the public internet, model weights download on first launch, and telemetry phones home — any single public call puts the no-egress claim in question during an MLPS assessment. Ask this one until you get a hard answer.
Likewise, if "delete meeting" only removes a database row and leaves intermediate files and vector caches on disk, residual information protection fails.
See the MLPS Level 3 compliance breakdown for the full detail.
6. Eight Questions to Answer Before Selecting
Send this list to vendors, or ask your own team. Any option that cannot answer is out:
- Does any part of the transcription pipeline make a public network call? (license, model and telemetry all count)
- What is the sustained concurrent stream count per server? Sustained, not peak
- Are hotwords supported? Hard constraint or soft guidance?
- Is speaker separation built in, or added as a separate module? How is overlapping speech handled?
- Which dialects are supported? What is the measured accuracy? Can you share a test set?
- Are Ascend / Cambricon / Hygon supported? Natively, or ported on site?
- Is air-gapped deployment supported? What does the offline bundle contain?
- What documentation is provided for MLPS assessment? (deployment topology, data flow diagram, audit log specification)
Questions 1 and 7 are the watershed. A vendor who cannot answer them has not delivered in a real compliance environment.
7. When Not to Go On-Premise
A word against our own interest, so this does not read as a pitch.
Skip on-premise if:
- Daily volume is tiny (a few hours a day) — pay-per-use cloud is cheaper and carries no operations burden
- There is no compliance requirement — if you can use the cloud, do not spend six figures on servers for a feeling of safety
- Nobody on the team does operations — after going on-premise, model upgrades, hardware failures and performance tuning are all yours
- You need transcription once — buying a server to transcribe one batch and then idle it is a bad trade
The right trigger for on-premise is a compliance red line or sustained concurrency pressure — not "it feels safer."
8. How VoiVision Does It
VoiVision builds enterprise on-premise meeting transcription, with the full stack developed in-house from speech frontend through ASR, speaker separation and semantic structuring:
- 52 years of NLP research, founded in 1973 as the NEU NLP Lab
- 200+ papers published (20+ CCF-A), 110+ invention patents (54 granted)
- NiuTrans, the machine translation engine, used over 100,000 times worldwide
- Inference 4x faster than mainstream general-purpose LLMs on equivalent hardware
- Native support for Ascend 310P/910B, Cambricon MLU370/590, Hygon DCU and NVIDIA T4/L4/A10/A100, plus CPU-only operation
- 22 dialects, and air-gapped deployment
- ≥98% Chinese transcription accuracy, with clusters supporting 50+ concurrent streams
Use the eight questions above as-is. Any solution that cannot answer them deserves a second look, no matter how low the quote.
Need an assessment against your MLPS level and concurrency target? Book a demo, or start with the complete private deployment guide.
FAQ
Q: On-premise or cloud ASR — which is better?
A: It depends on three conditions: daily volume, compliance requirements and concurrency pressure. With low daily volume and no compliance constraint, cloud APIs usually have a lower three-year TCO and far less operational burden. When data residency is a hard red line, Xinchuang rules apply, or concurrency pressure is sustained, on-premise is the right call. The trigger for on-premise is a compliance red line or sustained concurrency — not a vague feeling that it is safer.
Q: What are the barriers to self-hosting Whisper?
A: Five main ones: no hotword support (proper nouns get written as similar-sounding characters, and initial_prompt is only weak guidance), no built-in speaker separation (you must add pyannote or a similar module and align timestamps yourself), weak streaming (chunking produces boundary errors at split points), essentially no dialect support, and domestic accelerators require porting yourself (ONNX export plus ATC conversion, with plenty of pitfalls).
Q: How do you calculate three-year TCO for on-premise ASR?
A: Do not compare license fees alone. Include per-hour transcription cost, hardware, software licensing, implementation effort, three-year operations effort and failure fallback. Open-source self-hosting saves on licensing but costs in people — implementation typically takes one to two person-months and operations around 0.3–0.5 FTE, higher when domestic accelerators and air-gapped deployment are involved.
Q: What does MLPS Level 3 require when selecting a meeting minutes system?
A: Five controls matter: data residency (audio and text stored locally, no public upload), access control (identity verification and role-based permissions), security auditing (traceable operations, retained logs), transport encryption (encrypted internal communication), and residual information protection (deleted data must be unrecoverable). The most overlooked and most critical question is whether any part of the pipeline makes a public network call.
Q: How do I verify an on-premise solution really keeps data in-domain?
A: Ask directly whether any part of the transcription pipeline makes a public network call, and explicitly include license checks, model weight downloads and telemetry. Any single public call will put the "no data egress" claim in question during an MLPS assessment.
Q: For 50 hours of meetings per day, cloud or on-premise?
A: Without compliance constraints, pay-per-use cloud usually costs meaningfully less over three years. If data residency, Xinchuang or classified requirements apply, or concurrency pressure is sustained, choose on-premise. The tipping point on both economics and compliance usually arrives when either concurrency pressure or a no-cloud compliance rule is present.
Q: Does on-premise ASR support Ascend and other domestic accelerators?
A: It depends on the route. Open-source models require you to complete ONNX export and ATC conversion yourself, which is substantial work with many pitfalls. Commercial engines that natively support Ascend 310P/910B, Cambricon MLU370/590 and Hygon DCU remove that porting effort entirely. Always confirm whether support is native or requires on-site porting.
