Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 44 additions & 0 deletions CNN-examples/getting_started_resnet/int8/Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# ==============================================================================
# Dockerfile: AMD Quark Quantization Pipeline for ResNet INT8
# MLOps Decoupled Build Environment (Google Cloud Console / Linux / Docker)
# ==============================================================================
# Purpose:
# Isolates the AMD Quark Quantization and PyTorch-to-ONNX export pipeline
# from the host Ryzen AI runtime environment. This completely avoids the
# Python package dependency conflicts with 'flexml 1.8.0' (Issue #385).
#
# Volume Mount Contract:
# - /app/models : Mounted from host to persist exported ONNX & quantized models
# - /app/data : Mounted from host to cache downloaded CIFAR-10 datasets
# ==============================================================================

FROM python:3.10-slim

ENV PYTHONUNBUFFERED=1 \
DEBIAN_FRONTEND=noninteractive

WORKDIR /app

# Install system dependencies
RUN apt-get update && apt-get install -y --no-install-recommends \
ca-certificates \
curl \
wget \
git \
libgl1 \
libglib2.0-0 \
&& rm -rf /var/lib/apt/lists/*

# Install exact pinned Python quantization dependencies for reproducibility
COPY requirements_quantize.txt .
RUN pip install --no-cache-dir --upgrade pip && \
pip install --no-cache-dir -r requirements_quantize.txt

# Copy source scripts
COPY resnet_utils.py prepare_model_data.py resnet_quantize.py ./

# Create explicit directories for host volume mounts
RUN mkdir -p /app/data /app/models

# Entry command: prepare data, export ONNX, and run Quark INT8 quantization
CMD ["sh", "-c", "python prepare_model_data.py && python resnet_quantize.py && echo '=== Quantization Complete! Artifact: models/resnet_quantized.onnx ==='"]
158 changes: 154 additions & 4 deletions CNN-examples/getting_started_resnet/int8/Readme.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,162 @@
<table class="sphinxhide" width="100%">
<tr width="100%">
<td align="center"><img src="https://raw.githubusercontent.com/Xilinx/Image-Collateral/main/xilinx-logo.png" width="30%"/><h1> Ryzen™ AI ResNet Tutorial </h1>
<td align="center"><img src="https://raw.githubusercontent.com/Xilinx/Image-Collateral/main/xilinx-logo.png" width="30%"/><h1> Ryzen™ AI ResNet INT8 Tutorial </h1>
</td>
</tr>
</table>

# Getting Started Example
# Getting Started ResNet with INT8 Quantization & NPU Deployment

This tutorial uses a fine-tuned version of the ResNet model (using the CIFAR-10 dataset) to demonstrate the process of preparing, quantizing, and deploying a model using Ryzen AI Software. The tutorial features deployment using both Python and C++ ONNX runtime code.
This tutorial demonstrates how to prepare, quantize, and deploy a fine-tuned **ResNet-50** model on CIFAR-10 using **AMD Quark** and **Ryzen AI Software 1.8** with ONNX Runtime on the NPU.

For a walkthrough of this tutorial please follow the [tutorial documentation](https://ryzenai.docs.amd.com/en/latest/getstartex.html)
---

## ⚠️ Important Notice on Issue #385 & Dependency Isolation

In **Ryzen AI 1.8**, AMD unbundled the Quark quantization toolkit from the default installation. Running `pip install` for Quark inside the active `ryzen-ai-1.8.0` environment creates severe Python package dependency conflicts:
- `flexml 1.8.0` strictly requires `torch==2.6.0+cpu`, `torchvision==0.21.0+cpu`, and `onnx==1.21.0`.
- AMD Quark requires newer or unpinned versions of `torch`, `torchvision`, and `onnx`.
- Installing both into the same environment breaks `flexml`, causing NPU execution (`predict.py --ep npu`) to fail.

### MLOps Best Practice: Decoupled Build vs. Runtime Pipeline

```
┌────────────────────────────────────────────────────────┐
│ Phase 1: Build & Quantization │
│ (Docker / Google Cloud Console / Isolated Conda) │
│ │
│ - Export ResNet PyTorch checkpoint -> ONNX │
│ - AMD Quark Calibration & INT8 Quantization │
│ - Output Artifact: models/resnet_quantized.onnx │
└──────────────────────────┬─────────────────────────────┘
│
│ Transfer Artifact
▼
┌────────────────────────────────────────────────────────┐
│ Phase 2: Edge NPU Serving │
│ (Local Ryzen AI 1.8 Host) │
│ │
│ - ONNX Runtime with VitisAIExecutionProvider (flexml) │
│ - Minimal runtime dependencies (OpenCV, Pillow) │
│ - Zero pollution of flexml pinned packages │
└──────────────────────────┘
```

---

## Option A: Docker / Google Cloud Console (Recommended)

Using Docker keeps the quantization pipeline 100% reproducible and isolated from your host OS. You can run this in **Google Cloud Console (Cloud Shell / GCP VM)** or **Docker Desktop**.

### Step 1: Run Quantization in Docker / GCP Cloud Shell
Navigate to this directory and run:

**On Linux / Google Cloud Console / WSL:**
```bash
chmod +x docker_run.sh
./docker_run.sh
```

**On Windows (Docker Desktop):**
```cmd
docker_run.bat
```

**Or run Docker manually:**
```bash
# 1. Build the image
docker build -t ryzenai-resnet-quark:latest -f Dockerfile .

# 2. Run the container with volume mounts
docker run --rm \
-v "$(pwd)/models:/app/models" \
-v "$(pwd)/data:/app/data" \
ryzenai-resnet-quark:latest
```

This will automatically:
1. Download CIFAR-10 data and the pre-trained ResNet-50 checkpoint (via Git LFS URL).
2. Export `models/resnet_trained_for_cifar10.onnx`.
3. Quantize the model using Quark into `models/resnet_quantized.onnx`.

### Step 2: Transfer Quantized Model to Ryzen AI Host
If you ran Docker in Google Cloud Console:
- Download the generated file `models/resnet_quantized.onnx` to your local machine (place it in `CNN-examples/getting_started_resnet/int8/models/`).

---

## Option B: Dual Conda Environments (Local Windows without Docker)

If you prefer running locally on Windows without Docker, use two separate Conda environments.

### Step 1: Quantization Environment (Quark)
```cmd
# Create an isolated environment for Quark
conda create -n quark_env python=3.10 -y
conda activate quark_env

# Install quantization dependencies
pip install -r requirements_quantize.txt

# Export model to ONNX & Quantize
python prepare_model_data.py
python resnet_quantize.py

# Deactivate when done
conda deactivate
```

---

## Phase 2: Inference & Deployment on Ryzen AI NPU

Now that `models/resnet_quantized.onnx` is ready, run inference using your official `ryzen-ai-1.8.0` environment:

### Step 1: Activate Ryzen AI Environment
```cmd
conda activate ryzen-ai-1.8.0
```

### Step 2: Install Runtime Inference Requirements
```cmd
pip install -r requirements_infer.txt
```
*(Notice: This only installs `opencv-python` and `Pillow`, keeping `flexml 1.8.0` completely intact!)*

### Step 3: Run Inference

**On NPU:**
```cmd
python predict.py --ep npu
```

**On CPU (Fallback):**
```cmd
python predict.py --ep cpu
```

### Step 4: C++ Inference (Optional)
Navigate to `cpp/resnet_cifar`:
```cmd
cd cpp/resnet_cifar
mkdir build && cd build
cmake .. -A x64 -DCMAKE_PREFIX_PATH="%RYZEN_AI_INSTALLATION_PATH%"
cmake --build . --config Release
cd Release
resnet_cifar.exe -m ..\..\..\models\resnet_quantized.onnx -d ..\..\..\data\cifar-10-batches-bin
```

---

## File Overview

| File | Purpose |
| :--- | :--- |
| `Dockerfile` | Container configuration for offline Quark quantization |
| `docker_run.sh` / `docker_run.bat` | One-click script to build & run Docker quantization |
| `requirements_quantize.txt` | Dependencies for Model Export & Quark INT8 quantization |
| `requirements_infer.txt` | Dependencies for NPU runtime inference on Ryzen AI host |
| `prepare_model_data.py` | Downloads CIFAR-10 data, pre-trained weights, and exports ONNX |
| `resnet_quantize.py` | Runs AMD Quark calibration and produces `resnet_quantized.onnx` |
| `predict.py` | Runs inference on Ryzen AI NPU using ONNX Runtime VitisAI EP |
| `resnet_utils.py` | Helper functions for NPU device detection and paths |
37 changes: 37 additions & 0 deletions CNN-examples/getting_started_resnet/int8/docker_run.bat
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
@echo off
REM ==============================================================================
REM AMD Quark Quantization Pipeline Runner (Windows Docker)
REM ==============================================================================
setlocal enabledelayedexpansion

echo ==================================================================
echo Building Docker image: ryzenai-resnet-quark
echo ==================================================================
docker build -t ryzenai-resnet-quark:latest -f Dockerfile .
if %errorlevel% neq 0 (
echo Error: Docker build failed!
exit /b %errorlevel%
)

echo ==================================================================
echo Running Quantization Pipeline in Container
echo ==================================================================
if not exist "models" mkdir models
if not exist "data" mkdir data

docker run --rm ^
-v "%cd%\models:/app/models" ^
-v "%cd%\data:/app/data" ^
ryzenai-resnet-quark:latest

if %errorlevel% neq 0 (
echo Error: Quantization failed!
exit /b %errorlevel%
)

echo ==================================================================
echo SUCCESS!
echo Quantized model generated at: %cd%\models\resnet_quantized.onnx
echo You can now run predict.py in your ryzen-ai conda environment.
echo ==================================================================

30 changes: 30 additions & 0 deletions CNN-examples/getting_started_resnet/int8/docker_run.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
#!/bin/bash
# ==============================================================================
# AMD Quark Quantization Pipeline Runner (Docker / Google Cloud Console)
# ==============================================================================
set -e

SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
cd "$SCRIPT_DIR"

echo "=================================================================="
echo " Building Docker image: ryzenai-resnet-quark"
echo "=================================================================="
docker build -t ryzenai-resnet-quark:latest -f Dockerfile .

echo "=================================================================="
echo " Running Quantization Pipeline in Container"
echo " Mounts local 'models' and 'data' directories"
echo "=================================================================="
mkdir -p models data
docker run --rm \
-v "${PWD}/models:/app/models" \
-v "${PWD}/data:/app/data" \
ryzenai-resnet-quark:latest

echo "=================================================================="
echo " SUCCESS!"
echo " Quantized model generated at: ${PWD}/models/resnet_quantized.onnx"
echo " You can now transfer this model to your Ryzen AI PC for NPU inference."
echo "=================================================================="

68 changes: 65 additions & 3 deletions CNN-examples/getting_started_resnet/int8/prepare_model_data.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,13 +3,71 @@
# Licensed under the MIT License.
# --------------------------------------------------------------------------
import argparse
import hashlib
import os
import random
import ssl
import sys
import tarfile
import urllib.request

ssl._create_default_https_context = ssl._create_unverified_context
EXPECTED_CHECKPOINT_SHA256 = "6fbc917577ccf951e7c6039e3d85e95b11162abc06f9a69fe9ef26a5178d00cd"
CHECKPOINT_LFS_URL = os.environ.get(
"RYZENAI_RESNET_CHECKPOINT_URL",
"https://media.githubusercontent.com/media/amd/RyzenAI-SW/main/CNN-examples/getting_started_resnet/int8/models/resnet_trained_for_cifar10.pt"
)


def compute_sha256(filepath):
sha256_hash = hashlib.sha256()
with open(filepath, "rb") as f:
for byte_block in iter(lambda: f.read(65536), b""):
sha256_hash.update(byte_block)
return sha256_hash.hexdigest().lower()


def is_git_lfs_pointer(filepath):
"""Detects Git-LFS pointer signature: 'version https://git-lfs.github.com/spec/v1'"""
if not filepath.exists():
return False, None
if filepath.stat().st_size > 1024:
return False, None
try:
content = filepath.read_text(encoding="utf-8")
if "version https://git-lfs.github.com/spec/v1" in content:
for line in content.splitlines():
if line.startswith("oid sha256:"):
return True, line.split("oid sha256:")[1].strip()
return True, None
except Exception:
pass
return False, None


def verify_or_download_checkpoint(model_pt_path):
"""Ensures authoritative pre-trained checkpoint is present and verified by SHA-256."""
is_lfs, pointer_sha = is_git_lfs_pointer(model_pt_path)

# Check for consistency if local pointer exists
if is_lfs and pointer_sha and pointer_sha.lower() != EXPECTED_CHECKPOINT_SHA256.lower():
raise RuntimeError(
f"Security Error: Repository Git-LFS pointer digest ({pointer_sha}) does not match "
f"authoritative checkpoint digest ({EXPECTED_CHECKPOINT_SHA256})."
)

if not model_pt_path.exists() or is_lfs:
print("Pre-trained checkpoint is an LFS pointer or missing. Downloading from Git LFS...")
print(f"Downloading from: {CHECKPOINT_LFS_URL}")
urllib.request.urlretrieve(CHECKPOINT_LFS_URL, model_pt_path, reporthook=download_progress)

actual_sha = compute_sha256(model_pt_path)
if actual_sha != EXPECTED_CHECKPOINT_SHA256.lower():
raise RuntimeError(
f"Security/Integrity Error: SHA-256 checksum mismatch for {model_pt_path}!\n"
f"Expected: {EXPECTED_CHECKPOINT_SHA256}\n"
f"Actual : {actual_sha}\n"
f"The downloaded file may be corrupted or compromised."
)
print(f"Verified checkpoint SHA-256 digest: {actual_sha} (Integrity OK)")


def download_progress(block_num, block_size, total_size):
Expand Down Expand Up @@ -171,9 +229,13 @@ def main():
print(f"Extracting {label}...")
with tarfile.open(path) as f:
f.extractall(data_dir)
model_pt_path = models_dir / "resnet_trained_for_cifar10.pt"
if args.train:
prepare_model(args.num_epochs, models_dir, data_dir)
model = torch.load(str(models_dir / "resnet_trained_for_cifar10.pt"), weights_only=False)
else:
verify_or_download_checkpoint(model_pt_path)

model = torch.load(str(model_pt_path), weights_only=False)
export_to_onnx(model, models_dir)


Expand Down
25 changes: 21 additions & 4 deletions CNN-examples/getting_started_resnet/int8/requirements.txt
Original file line number Diff line number Diff line change
@@ -1,5 +1,22 @@
amd-quark==0.11
torchvision==0.23.0
opencv-python==4.11.0.86
numpy==1.26.4
# ==============================================================================
# Ryzen AI ResNet INT8 Tutorial Requirements
# ==============================================================================
# WARNING / IMPORTANT NOTICE (Issue #385):
# AMD Quark is unbundled from Ryzen AI 1.8. Installing 'amd-quark' or
# conflicting 'torchvision' directly into your active 'ryzen-ai-1.8.0' environment
# will break 'flexml 1.8.0' (which requires torch==2.6.0+cpu & onnx==1.21.0).
#
# Please use the decoupled requirement files based on the stage:
#
# 1. QUANTIZATION STAGE (Docker / Google Cloud Console / Isolated Conda):
# Use: requirements_quantize.txt
# Or run the pre-configured Docker container:
# bash docker_run.sh (Linux / Google Cloud Console / WSL)
# docker_run.bat (Windows)
#
# 2. RUNTIME INFERENCE STAGE (Ryzen AI Host Environment):
# In your active 'ryzen-ai-1.8.0' environment, run:
# pip install -r requirements_infer.txt
# ==============================================================================

-r requirements_infer.txt
Loading