在 Jetson AGX Orin 上部署 Qwen3.5-0.8B:从 TensorRT Engine 到 Vision
记录使用 TensorRT Edge-LLM 在 Jetson AGX Orin 上构建 Qwen3.5-0.8B、启动 OpenAI 兼容服务、测量生成速度,并补齐视觉 engine 完成图片识别的全过程。
在 Jetson AGX Orin 上部署 Qwen3.5-0.8B:从 TensorRT Engine 到 Vision
这次实践的目标,是把 Qwen/Qwen3.5-0.8B 真正部署到 Jetson AGX Orin 上:先跑通纯文本生成,再把视觉 encoder 补齐,最后通过局域网提供 OpenAI 兼容接口。
最终链路如下:
Qwen3.5-0.8B checkpoint
│
├── TensorRT Edge-LLM exporter
│ ├── llm/model.onnx
│ └── visual/model.onnx
│
├── llm_build
│ └── llm.engine
│
├── visual_build
│ └── visual/visual.engine
│
└── TensorRT Edge-LLM server
└── OpenAI-compatible API
一、先纠正一个容易混淆的地方
一开始看到 engine 配置中的:
model: qwen3_5_text
容易误以为 Qwen3.5-0.8B 是纯文本模型。实际上,官方 checkpoint 是多模态模型,模型配置包含 vision_config、image_token_id 和 video_token_id,官方模型卡也提供了图片输入示例。
官方资料:
这里的 qwen3_5_text 只是语言模型部分的名称。Qwen3.5 的视觉 encoder 是独立的 ONNX 和 TensorRT engine,不能只构建 llm.engine 就获得图片理解能力。
二、实际环境
目标设备是 192.168.1.30 上的 Jetson AGX Orin 64GB。实际环境如下:
| 项目 | 版本或状态 |
|---|---|
| 设备 | Jetson AGX Orin Developer Kit 64GB |
| 架构 | aarch64 |
| JetPack | 6.2 |
| L4T | R36.4.3 |
| Ubuntu | 22.04 |
| CUDA | 12.6 |
| TensorRT | 10.3.0 |
| Python | 3.10 |
| TensorRT Edge-LLM | 0.10.0 |
| 电源模式 | MAXN |
| Engine build | build-qwen35-gdn |
这次没有升级 JetPack、CUDA 或 TensorRT。导出和 engine build 分开进行:export 侧生成 ONNX,Orin 侧调用 TensorRT 本机构建 engine。
三、先构建文本 engine
Qwen3.5-0.8B 使用单 batch、16K 输入和约 32K KV cache:
cd /home/wooley/TensorRT-Edge-LLM
export EDGELLM_PLUGIN_PATH=$PWD/build-qwen35-gdn/libNvInfer_edgellm_plugin.so
mkdir -p /home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn logs
./build-qwen35-gdn/examples/llm/llm_build \
--onnxDir /home/wooley/models/Qwen3.5-0.8B/onnx/llm \
--engineDir /home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn \
--maxBatchSize 1 \
--maxInputLen 16384 \
--maxKVCacheCapacity 32768 \
2>&1 | tee logs/qwen35_08b_engine_build_32k_gdn.log
构建完成后,语言模型部分包括:
/home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn/
├── config.json
├── embedding.safetensors
├── llm.engine
├── processed_chat_template.json
├── tokenizer.json
└── tokenizer_config.json
其中 llm.engine 大约 1.45 GiB。
四、为什么第一次没有 Vision
第一次 export 日志中出现的是:
visual: no
最终只生成了 onnx/llm,所以第一次构建出来的只是文本 engine。这个过程可以正常完成文本生成,但不能处理图片。
当前 Edge-LLM 的 exporter 对支持的 VLM 会输出多个组件。为了只补充视觉部分、保留已经验证过的 onnx/llm,在 export 主机上执行:
export HF_HUB_OFFLINE=1
cd /home/wooley/TensorRT-Edge-LLM-export
source .venv/bin/activate
tensorrt-edgellm-export \
/media/wooley/AGXINSTALL/models/Qwen3.5-0.8B \
/media/wooley/AGXINSTALL/models/Qwen3.5-0.8B/onnx \
--components visual \
2>&1 | tee /media/wooley/AGXINSTALL/models/Qwen3.5-0.8B/export_visual.log
成功后,导出目录应包含:
onnx/
├── llm/
└── visual/
├── config.json
├── model.onnx
├── model.onnx.data
└── preprocessor_config.json
曾经有 Edge-LLM 高级 Python API 导出 Qwen3.5 视觉部分时缺少 model_config 参数的问题。如果遇到:
_export_visual() missing model_config
应先检查 export 主机上的 Edge-LLM checkout 是否过旧。这个问题曾被记录在 NVIDIA TensorRT Edge-LLM issue #107。
五、复制并构建视觉 engine
只需要把视觉 ONNX 目录复制到 Orin:
rsync -a --info=progress2 \
/media/wooley/AGXINSTALL/models/Qwen3.5-0.8B/onnx/visual/ \
wooley@192.168.1.30:/home/wooley/models/Qwen3.5-0.8B/onnx/visual/
在 Orin 上构建视觉 engine:
cd /home/wooley/TensorRT-Edge-LLM
export LD_LIBRARY_PATH=/home/wooley/TensorRT-Edge-LLM/build-qwen35-gdn:/usr/local/cuda/lib64:/usr/lib/aarch64-linux-gnu:$LD_LIBRARY_PATH
export EDGELLM_PLUGIN_PATH=$PWD/build-qwen35-gdn/libNvInfer_edgellm_plugin.so
./build-qwen35-gdn/examples/multimodal/visual_build \
--onnxDir /home/wooley/models/Qwen3.5-0.8B/onnx/visual \
--engineDir /home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn \
2>&1 | tee logs/qwen35_08b_visual_engine_build.log
构建成功后得到:
/home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn/visual/
├── config.json
├── preprocessor_config.json
└── visual.engine
视觉 engine 大约 200 MiB。官方流程也是使用 visual_build 将视觉 engine 写入 engine/visual/,再通过 multimodalEngineDir 交给运行时。官方视觉 engine 流程
六、启动 OpenAI 兼容服务
先启动文本 engine 时,服务的核心参数是:
python -m experimental.server \
--model /home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn \
--host 0.0.0.0 \
--port 8001 \
--max-input-len 16384 \
--max-kv-cache-capacity 32768
加入 Vision 后,关键是增加同一个 engine 根目录作为 multimodal engine 目录:
python -m experimental.server \
--model /home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn \
--multimodal-engine-dir /home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn \
--host 0.0.0.0 \
--port 8001 \
--max-input-len 16384 \
--max-kv-cache-capacity 32768
实际部署时使用 systemd 管理服务,核心 ExecStart 如下:
ExecStart=/home/wooley/TensorRT-Edge-LLM/.torch-jetpack62-venv/bin/python -m experimental.server \
--model /home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn \
--multimodal-engine-dir /home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn \
--host 0.0.0.0 \
--port 8001 \
--max-input-len 16384 \
--max-kv-cache-capacity 32768
对应服务名:
edge-llm-qwen35-0p8b-32k-gdn.service
检查服务:
systemctl status edge-llm-qwen35-0p8b-32k-gdn.service
curl http://192.168.1.30:8001/health
curl http://192.168.1.30:8001/v1/models
成功时日志中会出现:
Visual runner successfully initialized
服务返回的模型 ID 是:
engine-32k-gdn
客户端使用的 Base URL:
http://192.168.1.30:8001/v1
七、图片识别测试中的接口差异
第一次按照常见 OpenAI 多模态格式发送:
{
"type": "image_url",
"image_url": {
"url": "https://example.com/test.jpg"
}
}
当前 Edge-LLM server 返回:
Unsupported content type: image_url
这个版本的 server 使用的是:
{
"type": "image",
"image": "https://example.com/test.jpg"
}
完整测试请求:
curl http://192.168.1.30:8001/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "engine-32k-gdn",
"messages": [
{
"role": "user",
"content": [
{
"type": "image",
"image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"
},
{
"type": "text",
"text": "请仔细观察图片,只回答图片中糖果包装上动物的名称。"
}
]
}
],
"max_tokens": 16,
"temperature": 0,
"stream": false
}'
实际返回:
{
"choices": [
{
"message": {
"role": "assistant",
"content": "牛"
}
}
]
}
这次请求返回 HTTP 200,同时服务日志确认视觉 runner 已初始化。因此可以确认图片已经经过视觉 encoder,而不是被当成普通文本忽略。
八、性能测试
原生 Decode benchmark
使用 llm_bench,上下文为 2048 tokens,batch size 为 1:
./build-qwen35-gdn/examples/llm/llm_bench \
--engineDir /home/wooley/models/Qwen3.5-0.8B/engine-32k-gdn \
--mode decode \
--batchSize 1 \
--pastKVLen 2048 \
--warmup 3 \
--iterations 10 \
--profile \
2>&1 | tee logs/qwen35_08b_decode_2048.log
结果:
| 项目 | 结果 |
|---|---|
| Decode 单步平均耗时 | 16.46 ms |
| 理论 Decode 速度 | 约 60.8 tok/s |
| Past KV length | 2048 |
| Batch size | 1 |
| Engine 文件 | 1450 MiB |
| Visual engine | 约 200 MiB |
因为命令带有 --profile,日志明确说明 CUDA Graph 的端到端计时被跳过。因此这个数字更适合分析算子,不一定是服务实际最高速度。
实际 API 端到端测试
一次生成 186 tokens,总耗时 2.379 秒:
端到端速度:78.19 tok/s
这个数字包含 HTTP、tokenizer、prefill、decode、采样和输出处理,更接近 VS Code 或其他客户端的真实体验。它高于带 profiling 的原生 benchmark 是正常的,因为正式服务没有打开逐层 profiling,并且可以使用更完整的 CUDA Graph 路径。
九、最终结论
这次 Qwen3.5-0.8B 的结果可以分成两部分:
- 文本链路已经稳定:TensorRT engine、插件、Python server 和局域网 API 都正常。
- Vision 链路也已经跑通:补充视觉 ONNX、构建
visual.engine、让 server 加载 multimodal engine,并完成图片识别请求。
当前可用接口:
http://192.168.1.30:8001/v1
当前模型 ID:
engine-32k-gdn
当前配置适合作为轻量本地 Agent 和图片问答服务:单 batch、16K 输入、32K KV cache,实际 API 生成速度约 78 tok/s。需要注意的是,VS Code 或其他客户端如果只会发送标准 image_url 格式,可能需要增加一层请求格式转换;当前 Edge-LLM server 直接接受的是 type=image 加 image 字段。
后续如果继续测试更大的 Qwen3.5 或代码模型,建议保持相同的 batch、上下文和生成参数,同时分别记录纯 Decode、端到端速度、TTFT、峰值统一内存和视觉 encoder 的额外开销。
Conversation
正在读取留言…