Ascend server in hand: how to get the first model running
Community Discussion · Policy

Ascend server in hand: how to get the first model running

Warehouse RunningWarehouse RunningSep 232026/09/23 207 views

Last month our warehouse wanted to move visual verification from the cloud to on-prem, and procurement shoved an Ascend inference server at me. I opened the box, sat at the command line for half an hour, and had no idea where to start. I'm familiar with the GPU stack, but switching to Ascend basically reset all the muscle memory I'd built up over the past few years.

This post is for people just getting started like me: you've got an Ascend machine and you want to get your first model to produce output. No theory, just how I fumbled my way through it step by step.

Day one, first figure out what card you're holding. Open a terminal and type npu-smi info. This command is the equivalent of nvidia-smi on the NVIDIA side, used to check chip status. Normally it lists the card model, memory usage, and temperature. If nothing happens after you type it, the driver isn't installed yet — go back and install the driver first. Then check the model. Ascend cards come in two types: one can only do inference, the other can train. The material says it plainly: inference servers don't support training; if you want to train, you need a training-type machine. The first one I got was an inference card, and I still tried to fine-tune a model on it — wasted a whole day. Finally, confirm the software stack is installed. Huawei's stack is called CANN — think of it as the graphics driver plus a bunch of compute libraries. Any model that wants to call the chip relies on it. After installing, you need to source the environment variable script in the terminal, otherwise the program can't find the chip. From my testing, not sourcing the environment variables is the most common first pitfall — the program errors out saying it can't find the device, and nine times out of ten that's the reason.

This step takes less than an hour, but it can save you two days. When you're done, you should be able to get npu-smi info to list the card properly, and be able to say clearly whether it's an inference card or a training card.

Day three, convert the model into a format the chip can eat. Ascend doesn't directly consume PyTorch weight files — it wants OM format. Think of it as turning fresh ingredients into a meal kit: the chip opens the bag and it's ready to use. Conversion uses the official tool called ATC. First export the original model into a general intermediate format, usually ONNX — this step is done in your own original environment, nothing to do with Ascend. Then feed the ONNX to ATC, specify the input dimensions and output nodes, and run it. If it goes smoothly, you get an OM in a few minutes. Finally, feed in the simplest possible input and check whether the output is correct.

The pitfall is right here. The first time I converted, it threw a long list of unsupported operators. Ascend's operator coverage is filled in gradually — it's not like you can just grab any model and convert it. The material also mentions that Ascend has done adaptation for open-source models like ChatGLM and OpenLLaMA, but it's still not very stable, there are some runtime issues, and they're still figuring it out. Two solutions: switch to a model someone in the community has already converted, or pull out the unsupported operators and rewrite them yourself. The former is fast, the latter is grunt work. I chose the former. The goal is to convert an OM file that produces reasonable output with sample input. If it won't run, switch models — don't bang your head against it.

A week later, wire it into your own pipeline. Getting a model to run and getting it to be usable are two different things. This week I mainly did two things: wrap the inference service in an interface layer so the upper-level business code doesn't have to care what chip is underneath — this abstraction layer is a must, otherwise you'll have to rewrite everything when you swap hardware later. The other thing is stress testing — the warehouse runs 7x24, and peak vs. average differ a lot.

There's a hard limit you need to know upfront. The material states clearly that for large models like LLaMA, multi-card model-parallel inference is currently not supported — meaning you can't split a large model across stacked cards. My original plan was designed around this, and later I had to switch to a model that fits on a single card. I should have checked this during the selection phase, but I only discovered it at deployment. Checking parallel support clearly during selection is far more important than tuning later. Once you've actually run it in a warehouse, you know: getting it to run and getting it to hold up are two different things.

After the model runs, there are three directions going forward, in my recommended order: first do quantization, compress the model smaller, and see how much accuracy drops. Then wire the inference service into the real business flow and run it with real data for a week. Only then consider multi-card and multi-machine. My approach on multi-card parallelism is to switch to a smaller model — good enough for now, but I haven't found a better path yet.

2 replies

?
Ctrl + Enter to reply
Is Operator Fusion Done?

I really feel that "multi-GPU parallelism not supported" point. I tried stacking GPUs to split a large model before too, wasted a whole day. You really gotta check parallel support first when choosing hardware.

Zhe Dan Bai De

Using inference cards to fine-tune models really is a waste of effort. The one in my lab hit that same pitfall.