vLLM/Recipes
Qwen

Qwen/Qwen3-30B-A3B

Qwen3 MoE model with 30.5B total and 3.3B active parameters, validated for BF16 serving on Intel Xeon 6 and Intel Arc Pro B70 and W8A8 serving on Ascend A3.

Validated on Intel Xeon 6 (GNR), Ascend A3 W8A8 and 4x Intel Arc Pro B70 with BF16 serving

moe30.5B / 3.3B40,960 ctxvLLM 0.8.5+text
Guide

Overview

Qwen3-30B-A3B is a mixture-of-experts Qwen3 model with 30.5B total parameters and 3.3B active parameters per token. This recipe includes an Intel Xeon 6 CPU configuration validated with BF16 serving and an Ascend A3 W8A8 variant.

Prerequisites

  • Hardware: 4x Xeon 6 CPUs, or 4x Intel Arc Pro B70 (32 GB per card).
  • vLLM >= 0.8.5

pip (Intel Xeon 6 CPUs)

For Intel and AMD x86 CPUs, follow the CPU pre-built wheels installation instructions.

Docker (Intel Xeon 6 CPUs)

docker pull vllm/vllm-openai-cpu:latest-x86_64

Docker (Intel Arc Pro B70)

docker pull vllm/vllm-openai-xpu:latest

Intel Xeon 6

Choose the tensor parallel size based on the system topology and deployment requirements. Add --tensor-parallel-size <N> when needed; the recipe does not prescribe a Xeon 6 TP value or hard-code NUMA node IDs.

vllm serve Qwen/Qwen3-30B-A3B \
  --max-num-batched-tokens 16384 \
  --gpu-memory-utilization 0.8

Docker:

docker run \
  --privileged --ipc=host -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai-cpu:latest-x86_64 Qwen/Qwen3-30B-A3B \
  --tensor-parallel-size 1 \
  --max-num-batched-tokens 16384 \
  --gpu-memory-utilization 0.8 \
  --enable-auto-tool-choice \
  --tool-call-parser hermes \
  --reasoning-parser qwen3

Intel Arc Pro B70 (XPU)

Validated on 4x Intel Arc Pro B70 (32 GB per card) with the official vLLM XPU image vllm/vllm-openai-xpu:latest, BF16, TP=4.

docker run --device /dev/dri \
  -v /dev/dri/by-path:/dev/dri/by-path --shm-size=16g \
  --privileged --ipc=host -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai-xpu:latest \
  Qwen/Qwen3-30B-A3B \
  --tensor-parallel-size 4 \
  --block-size 64 \
  --enforce-eager \
  --no-enable-prefix-caching \
  --disable-sliding-window \
  --max-model-len 9472 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.9

Client Usage

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
response = client.chat.completions.create(
    model="Qwen/Qwen3-30B-A3B",
    messages=[{"role": "user", "content": "Explain tensor parallelism briefly."}],
)
print(response.choices[0].message.content)

References