Running meeting transcription on Ascend 310P: version matrix, model conversion and seven field lessons
Can an edge inference card handle meeting transcription? This article documents deploying a Chinese ASR engine on the Ascend 310P — CANN environment alignment, the ATC parameters that matter when converting ONNX to OM, the four devices a container must mount, how to set dynamic dimensions, and seven lessons you only learn on site. The most dangerous one is a silent operator failure that produces all-zero output without any error.

This article documents deploying a Chinese speech transcription engine on the Ascend 310P inference card. The 310P is an edge inference card with a different positioning from the 910B, and its scenarios are closer to in-room local processing.
Data basis: figures marked "official specification" come from published product metrics; the performance table gives reference values for a typical environment. Actual numbers vary with audio conditions, model size and concurrency strategy — always defer to your own measurements.
What the 310P is good at, and what it is not
There is no shortage of articles about the Ascend 310P, but most stop at "CANN is installed and the demo runs". The real trap is that getting it running once and serving reliably are two entirely different things.
The conclusion up front:
| Dimension | 310P behaviour | Notes |
|---|---|---|
| Positioning | Edge inference card | Not a training card; do not expect to fine-tune on it |
| Split with 910B | Edge node | 910B handles central inference, 310P handles local processing |
| Pure transcription concurrency per card | Tens of streams | ASR single-model basis (reference value) |
| Full-pipeline concurrency | Drops significantly | Once speaker separation and structuring are added |
| Chinese transcription accuracy | ≥98% | Standard Mandarin meeting scenario (official specification) |
| Dialect coverage | 22 dialects | Automatic detection (official specification) |
One selection trap worth flagging: many vendors quote concurrency without stating the basis. "XX streams per card" might mean pure ASR, or it might mean the full pipeline. The two can differ several times over. Always ask which layer the number covers.
Our typical architecture is 910B at the centre, 310P at the edge: audio is transcribed locally in the meeting room, and only structured text returns to the centre. This keeps bandwidth under control and satisfies a "data does not leave the meeting room" requirement.
Environment preparation: the version matrix is the real obstacle
The pain point in the Ascend ecosystem is never compute — it is version alignment. Driver, firmware, CANN, container runtime, inference framework, model format: if any one of the six is misaligned it produces a baffling failure, and the error message frequently does not say what went wrong.
A combination we have verified and found stable:
| Component | Version | Notes |
|---|---|---|
| Driver / firmware | Tightly bound to CANN | Check the official compatibility table before upgrading |
| CANN | 8.x series | The operator library and ATC conversion tool live here |
| Container runtime | Ascend Docker Runtime | Required; ordinary Docker cannot mount the NPU |
| Inference framework | ONNX Runtime + ACL | Runs OM models |
| OS | openEuler / Kylin V10 / UnionTech UOS | Common in Xinchuang environments |
The first step is always to confirm the NPU is visible:
# Check NPU devices and status
npu-smi info
# You should see the device list with health status OK
# If no card appears here, everything downstream is wasted — fix the driver first
A field lesson: on some Kylin V10 environments,
npu-smi inforeported admpdaemon exception after the driver was installed. The cause turned out to be a conflict between the system's bundleddkmsand the driver installation script. The fix was to uninstall the system dkms and reinstall using the version bundled with the driver package.Distribution-specific issues like this are usually not findable in official documentation, and they are the most common blocker on site.
Model conversion: ONNX to OM
ASR models are typically trained in PyTorch, exported to ONNX, then converted to the Ascend OM format with ATC.
Three things to watch when exporting ONNX
torch.onnx.export(
model, dummy_input,
"asr.onnx",
opset_version=14, # Do not use the newest; lower opset versions are more compatible
do_constant_folding=True,
input_names=["audio", "audio_len"],
output_names=["logits"],
dynamic_axes={ # Critical: audio length is inherently dynamic
"audio": {0: "batch", 1: "time"},
"logits": {0: "batch", 1: "time"},
},
)
- A higher opset version is not better. Newer opsets are more likely to hit unsupported operators; 14 is a good compatibility choice.
- Dynamic axes must be declared. Otherwise you can only run fixed-length audio, which is unusable in practice.
- Enable constant folding.
do_constant_folding=Truesubstantially reduces graph size and improves conversion success.
The ATC parameters, one by one
atc --model=asr.onnx \
--framework=5 \
--output=asr_om \
--input_format=ND \
--input_shape="audio:-1,-1;audio_len:-1" \
--dynamic_dims="1,16000;1,32000;1,48000" \
--soc_version=Ascend310P \
--precision_mode=allow_fp16 \
--log=error
The three parameters most likely to cause trouble:
--soc_version: must match the actual hardware exactly. The 310P must be written asAscend310P;Ascend310orAscend910Bwill both fail at the load stage, with a very vague error.--dynamic_dims: the dynamic tiers. For audio, tier by duration in seconds. Too many tiers makes compile time explode; too few triggers frequent recompilation and jittery latency. The set above (1s / 2s / 3s) is the balance point we measured.--precision_mode:allow_fp16has negligible accuracy impact on ASR and a clear throughput gain.
Should you quantise?
| Precision | Throughput | Accuracy |
|---|---|---|
| FP16 | Baseline | Baseline |
| INT8 | Clear gain | Perceptible loss |
Conclusion: for scenarios extremely sensitive to accuracy — medical, legal — stay on FP16. For internal meetings and training records, where tolerance is higher, INT8 is the better trade.
Do not chase benchmark numbers with INT8 blindly. In meeting scenarios the accuracy loss shows up as "every name wrong, every term wrong", and the rework cost far exceeds the compute you saved.
Deployment: containerisation is the only sensible choice
Installing the environment on bare metal is practically unmaintainable in Xinchuang environments. We deliver as a container:
docker run -it --name asr_server \
--device /dev/davinci0 \
--device /dev/davinci_manager \
--device /dev/devmm_svm \
--device /dev/hisi_hdc \
-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \
-v /data/models:/models \
-p 8000:8000 \
asr-server:310p \
./asr_server --model /models/asr_om --port 8000
Missing any one of those --device flags produces an ACL initialisation failure, and the error does not tell you a device was omitted. This is the biggest blocker for newcomers.
One more lesson: do not invent a private protocol. Use standard RESTful (file transcription) plus WebSocket (real-time streaming), or integrating with a customer's OA and meeting systems becomes very painful.
Customers who already have servers can deploy the pure-software speech engine directly, gaining private transcription on domestic accelerators without buying hardware.
Seven lessons from the field
This is the most valuable part of the article. Every item was learned the hard way, not copied from documentation.
1. Dynamic shapes cause frequent recompilation
We did not set --dynamic_dims at first, so every differently-sized audio clip triggered a graph compilation and latency spiked to seconds. Once tiers were configured, latency settled in the sub-second range.
2. Long audio exhausts memory Feeding a two-hour meeting recording to the model in one pass caused an immediate OOM. You must use segmented incremental inference — cut by VAD and feed the decoder segment by segment instead of loading the whole thing. We later added a forced 60-second cap on the audio buffer, which eliminated the OOM risk entirely.
3. Silent operator failure A particular LayerNorm variant had no corresponding operator on Ascend. ATC reported no error at conversion time, and the runtime returned all zeros. This is the most dangerous class of failure, because it looks like success. You must run an equivalence check: feed the same audio through ONNX Runtime on CPU and the OM on NPU, then compare outputs. This is the only fallback — do not skip it.
4. Context leakage under multithreaded concurrency Early versions created a separate ACL context per request, exhausting contexts under high concurrency. Context pooling fixed it.
5. A driver upgrade means re-converting the model After a major CANN upgrade, an old OM model may fail to load or regress in performance. Put model conversion in CI rather than converting by hand and calling it done.
6. Offline packaging for air-gapped environments In a physically isolated environment, CANN dependencies, drivers, model weights, Chinese fonts and certificates must all be packaged in advance. We once hit it: a missing Chinese font on site, and every transcription result came out as boxes. The offline bundle must be rehearsed through the full deployment flow in an equivalent environment.
7. Misleading error messages derail diagnosis
Some Ascend errors give only a code, not a root cause. We once spent three days on a load failure with no progress, and it turned out to be a wrong soc_version. The lesson: when you hit a vague error, re-verify the version matrix from the top — it is far faster than reading logs line by line.
The full pipeline: from speech to minutes
One final point: speech transcription is only the first step.
What a real scenario needs is "minutes when the meeting ends". That requires speaker separation, semantic structuring (topics, decisions, action items) and minutes generation after the ASR stage. Layering these three on reduces concurrency significantly — which is exactly why we keep stressing that you must ask which layer a concurrency number covers.
Our approach runs the full pipeline on a single server, with audio and text never leaving the internal network. For government, finance and defence scenarios, that is a hard requirement.
When not to choose domestic accelerators
Let us be straight about this rather than pushing it.
If your situation meets either of the following, use a GPU instead:
- No mandatory Xinchuang or domestic-technology requirement → the GPU ecosystem is far more mature; do not create trouble for yourself
- No one on the team knows CANN → the operational cost will far exceed the money saved on hardware
The value of domestic accelerators is compliance, not cost-effectiveness. If compliance is not a hard constraint, the arithmetic does not work out.
Appendix: full compatibility list
| Component | Supported range |
|---|---|
| Ascend | 310P / 910B |
| Cambricon | MLU370 / MLU590 |
| Hygon | DCU |
| NVIDIA | T4 / L4 / A10 / A100 and the full range |
| CPU only | Supported (low concurrency / existing hardware reuse) |
Deployment via Docker and Kubernetes, running on bare metal, government cloud and Xinchuang environments, integrating with OA and meeting systems over standard RESTful / WebSocket interfaces.
Further reading: Private ASR on the Ascend 910B | How to choose a private-deployment ASR
FAQ
Q: How do I choose between Ascend 310P and 910B?
A: It depends on the workload shape. The 310P is an edge inference card focused on low power and local processing, suited to branch offices and in-room processing; the 910B has more compute and suits central inference and high-concurrency aggregation. The typical architecture for meeting transcription uses the 910B as the central node and 310P cards as edge nodes, so audio is transcribed locally and only structured text returns to the centre.
Q: How many concurrent meeting transcription streams can one 310P handle?
A: The definition matters. Concurrent streams for pure ASR transcription are relatively high; once speaker separation, terminology correction and minutes generation are layered on, the number drops significantly. If a vendor quotes a figure without saying whether it covers pure transcription or the full pipeline, the two can differ several times over — always ask.
Q: Why is version alignment the hardest part of Ascend deployment?
A: In the Ascend ecosystem the driver, firmware, CANN, container runtime, inference framework and model format are mutually bound. Any mismatch causes a load failure, and the error message is often vague and does not point at the version problem. The reliable approach is to lock the combination against the official compatibility table first, then install in order, without mixing driver packages from different batches.
Q: What happens if soc_version is wrong when converting ONNX to OM?
A: The model fails at the load stage, and the error does not indicate a model mismatch. For the 310P you must write Ascend310P; writing Ascend310 or Ascend910B will both fail to load. The vagueness of this error is one of the easiest ways to be misled on site.
Q: Why must I run an equivalence check after model conversion?
A: Because silent failures exist. When an operator variant has no implementation on Ascend, the ATC conversion does not report an error, but the runtime produces all-zero output. This kind of failure shows nothing unusual in the logs. The only fallback is to run the same audio through ONNX Runtime on CPU and the OM on NPU and compare the output layer by layer.
Q: Why does long audio exhaust memory?
A: Feeding an entire meeting recording to the model in one pass is a common mistake. Loading a two-hour recording directly causes an out-of-memory failure. The correct approach is to segment by voice activity detection and run incremental inference segment by segment, with a cap on the buffer, rather than loading the whole recording.
Q: Which devices must a container mount for Ascend?
A: You must mount the four devices /dev/davinci0, /dev/davinci_manager, /dev/devmm_svm and /dev/hisi_hdc, along with the driver directory. Missing any one of them triggers an ACL initialisation failure, and the error does not tell you a device is missing — this is the most common blocker for newcomers.
Q: How do I deploy offline in an air-gapped environment?
A: A physically isolated environment has no external network, so CANN dependencies, drivers, model weights, Chinese fonts and certificates must all be packaged in advance. We once hit a missing Chinese font on site and every transcription result came out as boxes. The offline bundle must be rehearsed through the complete deployment flow in an equivalent environment.
