Model HQ

Hardware optimized

AMD Supported Models

A complete catalog of 325 AI models optimized for enhanced performance across AMD CPUs, GPUs, and NPUs.

Hover the Encyclopedia tab on the right, or tap any icon to learn what a model type does.
325
Optimized models
ONNX
Runtime
CPU · GPU · NPU
Hardware targets
Why AMD

AMD optimization features

These models are optimized for enhanced performance on AMD hardware devices.

Performance benefits

  • Optimized for AMD CPU, GPU, and NPU architectures
  • Enhanced inference speed with ONNX Runtime and VITIS AI acceleration
  • Efficient execution of AI models with optimized hardware utilization
  • Streamlined AI inference for AMD Ryzen AI and Radeon platforms

Supported hardware

  • AMD x86_64 CPU
  • AMD Ryzen AI
  • AMD GPUs and NPUs
  • Windows 11 devices
Catalog

CPU Models

GGUF and tool models that run on the CPU alone — no GPU or NPU required.

Agentic

28 models
slim-boolean-phi-3-gguf3.8B
slim-extract-phi-3-gguf3.8B
slim-extract-qwen-1.5b-gguf1.5B
slim-extract-qwen-nano-gguf0.5B
slim-sa-ner-phi-3-gguf3.8B
slim-summary-phi-3-gguf3.8B
slim-xsum-phi-3-gguf3.8B
slim-boolean-tool1.1B
slim-category-tool1.1B
slim-emotions-tool1.1B
slim-extract-tool1.1B
slim-intent-tool1.1B
slim-ner-tool1.1B
slim-nli-tool1.1B
slim-q-gen-phi-3-tool3.8B
slim-q-gen-tiny-tool1.1B
slim-qa-gen-phi-3-tool3.8B
slim-qa-gen-tiny-tool1.1B
slim-ratings-tool1.1B
slim-sa-ner-tool1.1B
slim-sentiment-tool1.1B
slim-sql-tool1.1B
slim-summary-tiny-tool1.1B
slim-summary-tool1.1B
slim-tags-3b-tool3B
slim-tags-tool1.1B
slim-topics-tool1.1B
slim-xsum-tool1.1B

Coding

5 models
qwen2.5-coder-0.5b-instruct-generic-cpu:4-foundry0.5B
qwen2.5-coder-1.5b-instruct-generic-cpu:4-foundry1.5B
qwen2.5-coder-14b-instruct-generic-cpu:4-foundry14B
qwen2.5-coder-7b-instruct-generic-cpu:4-foundry7B
qwen-2.5-7b-coder-gguf7B

General Chat

14 models
gpt-oss-20b-generic-cpu:1-foundryNA
deepseek-r1-distill-qwen-14b-generic-cpu:4-foundry14B
deepseek-r1-distill-qwen-7b-generic-cpu:4-foundry7B
qwen3-0.6b-generic-cpu:4-foundry0.6B
qwen3-1.7b-generic-cpu:2-foundry1.7B
qwen3-14b-generic-cpu:2-foundry14B
qwen3-4b-generic-cpu:3-foundry4B
qwen3-8b-generic-cpu:2-foundry8B
Phi-4-mini-reasoning-generic-cpu:3-foundry3.8B
Phi-4-generic-cpu:2-foundry14B
qwen3.5-0.8b-generic-cpu:2-foundry0.8B
qwen3.5-2b-generic-cpu:2-foundry2B
qwen3.5-4b-generic-cpu:2-foundry4B
qwen3.5-9b-generic-cpu:2-foundry9B

General Chat - GGUF

34 models
deepseek-qwen-14b-gguf14B
deepseek-qwen-7b-gguf7B
dragon-llama-3.1-gguf8B
gemma-3-12b-gguf12B
gemma-3-4b-gguf4B
gemma-4-26b-gguf26B
gemma-4-2b-gguf2B
gemma-4-4b-gguf4B
gpt-oss-20b-gguf20B
granite-4-micro-gguf1.1B
llama-2-7b-chat-gguf7B
liquidai-lfm2-2.6b-gguf2.6B
ministral-3-14b-gguf14B
minicpm-2.6-gguf4B
mistral-small-3.2-24b-gguf24B
olmo-13b-gguf13B
openhermes-2.5-mistral-7b-gguf7B
phi-3-gguf3.8B
phi-3.5-gguf3.8B
phi-4-gguf14B
phi-4-mini-gguf3.8B
phi-4-mini-reasoning-gguf3.8B
qwen-3.5-27b-gguf27B
qwen-3.5-35b-a3b-gguf35B
qwen-3.5-4b-gguf4B
qwen-3.5-9b-gguf9B
qwen2.5-32b-gguf32B
qwen2.5-72b-gguf72B
qwen3-1.7b-gguf1.7B
qwen3-14b-gguf14B
qwen3-8b-gguf8B
starling-lm-7b-alpha-gguf7B
tiny-llama-chat-gguf1.1B
zephyr-7b-beta-gguf7B

Instruct

9 models
qwen2.5-0.5b-instruct-generic-cpu:4-foundry0.5B
qwen2.5-1.5b-instruct-generic-cpu:4-foundry1.5B
qwen2.5-14b-instruct-generic-cpu:4-foundry14B
qwen2.5-7b-instruct-generic-cpu:4-foundry7B
Phi-4-mini-instruct-generic-cpu:5-foundry3.8B
mistralai-Mistral-7B-Instruct-v0-2-generic-cpu:3-foundry7B
Phi-3-mini-128k-instruct-generic-cpu:3-foundry3.8B
Phi-3-mini-4k-instruct-generic-cpu:3-foundry3.8B
Phi-3.5-mini-instruct-generic-cpu:2-foundry3.8B

Instruct - GGUF

12 models
gemma-2-27b-instruct-gguf27B
gemma-2-9b-instruct-gguf9B
llama-3-8b-instruct-gguf8B
llama-3.1-instruct-gguf8B
llama-3.2-1b-instruct-gguf1.1B
llama-3.2-3b-instruct-gguf3B
mistral-7b-instruct-v0.3-gguf7B
qwen-2.5-14b-instruct-gguf14B
qwen2-0.5b-instruct-gguf0.5B
qwen2-1.5b-instruct-gguf1.5B
qwen2-7B-instruct-gguf7B
qwen3-4b-instruct-gguf4B

Question-answer

13 models
bling-qwen-0.5b-gguf0.5B
bling-tiny-llama-onnx1.1B
bling-phi-3-onnx3.8B
bling-qwen-1.5b-ov1.5B
bling-tiny-llama-ov1.1B
bling-phi-3-ov3.8B
bling-qwen-mini-tool1.5B
dragon-qwen-7b-gguf7B
dragon-mistral-0.3-gguf7B
dragon-yi-9b-gguf9B
dragon-llama-answer-tool7B
dragon-mistral-answer-tool7B
dragon-yi-answer-tool6B

Vision

4 models
qwen2.5-vl-3b-instruct-gguf3B
qwen3-vl-30b-gguf30B
qwen3-vl-4b-gguf4B
qwen3-vl-8b-gguf8B
Catalog

GPU/CPU/NPU Models

ONNX and OpenVINO (OV) models that run on the available GPU, CPU, or NPU.

Agentic - ONNX

13 models
slim-boolean-phi-3-onnx3.8B
slim-emotions-onnx1.1B
slim-extract-phi-3-onnx3.8B
slim-extract-tiny-onnx1.1B
slim-intent-onnx1.1B
slim-ner-onnx1.1B
slim-ratings-onnx1.1B
slim-sentiment-onnx1.1B
slim-sql-onnx1.1B
slim-summary-phi-3-onnx3.8B
slim-summary-tiny-onnx1.1B
slim-tags-onnx1.1B
slim-topics-onnx1.1B

Agentic - OV

22 models
slim-boolean-phi-3-ov3.8B
slim-category-ov1.1B
slim-emotions-ov1.1B
slim-extract-phi-3-ov3.8B
slim-extract-qwen-0.5b-ov0.5B
slim-extract-qwen-1.5b-ov1.5B
slim-extract-tiny-ov1.1B
slim-intent-ov1.1B
slim-ner-ov1.1B
slim-q-gen-tiny-ov1.1B
slim-qa-gen-tiny-ov1.1B
slim-ratings-ov1.1B
slim-sa-ner-phi-3-ov3.8B
slim-sentiment-ov1.1B
slim-sql-ov1.1B
slim-sql-phi-3-ov3.8B
slim-sql-qwen-base-ov1.5B
slim-summary-phi-3-ov3.8B
slim-summary-tiny-ov1.1B
slim-tags-ov1.1B
slim-topics-ov1.1B
slim-xsum-phi-3-ov3.8B

Coding

2 models
codegemma-7b-it-ov7B
qwen2.5-coder-7b-instruct-ov7B

Embedding - ONNX

9 models
all-mini-lm-l6-v2-onnx0.02B
bge-base-en-v1.5-onnx0.11B
bge-large-en-v1.5-onnx0.33B
bge-small-en-v1.5-onnx0.03B
gte-base-onnx0.11B
gte-large-onnx0.33B
gte-small-onnx0.03B
industry-bert-contracts-onnx0.11B
industry-bert-insurance-onnx0.11B

Embedding - OV

14 models
all-mini-lm-l6-v2-ov0.02B
all-mpnet-base-v2-ov0.11B
bge-base-en-v1.5-ov0.11B
bge-large-en-v1.5-ov0.33B
bge-small-en-v1.5-ov0.03B
gte-base-ov0.11B
gte-large-ov0.33B
gte-small-ov0.03B
industry-bert-asset-management-ov0.11B
industry-bert-contracts-ov0.11B
industry-bert-insurance-ov0.11B
industry-bert-loans-ov0.11B
industry-bert-sec-ov0.11B
paraphrase-multilingual-MiniLM-L12-v2-ov0.12B

General Chat

14 models
gpt-oss-20b-generic-gpu:1-foundry20B
deepseek-r1-distill-qwen-14b-generic-gpu:4-foundry14B
qwen3-0.6b-generic-gpu:2-foundry0.6B
qwen3-1.7b-generic-gpu:2-foundry1.7B
qwen3-14b-generic-gpu:2-foundry14B
qwen3-4b-generic-gpu:2-foundry4B
qwen3-8b-generic-gpu:2-foundry8B
Phi-4-mini-reasoning-generic-gpu:3-foundry3.8B
deepseek-r1-distill-qwen-7b-generic-gpu:4-foundry7B
Phi-4-generic-gpu:2-foundry14B
qwen3.5-0.8b-generic-gpu:2-foundry0.8B
qwen3.5-2b-generic-gpu:2-foundry2B
qwen3.5-4b-generic-gpu:2-foundry4B
qwen3.5-9b-generic-gpu:2-foundry9B

General Chat - OV

30 models
qwen2-0.5b-chat-ov0.5B
qwen3-1.7b-ov1.7B
qwen3-14b-ov14B
qwen3-4b-ov4B
qwen3-8b-ov8B
dolphin-2.9.4-llama3.1-8b-ov8B
llama-2-13b-chat-ov13B
llama-2-chat-ov7B
tiny-llama-chat-ov1.1B
phi-3-ov3.8B
phi-4-ov14B
phi-4-mini-ov3.8B
dolphin-2.9.3-mistral-7b-32k-ov7B
teknium-open-hermes-2.5-mistral-ov7B
zephyr-mistral-7b-chat-ov7B
yi-1.5-34b-ov34B
yi-6b-1.5v-chat-ov6B
yi-9b-chat-ov9B
gemma-2-27b-ov27B
gemma-2b-it-ov2B
gemma-7b-it-ov7B
stablelm-2-12b-chat-ov12B
stablelm-2-zephyr-1_6b-ov1.6B
stablelm-zephyr-3b-ov3B
dreamgen-wizardlm-2-7b-ov7B
granite-4-micro-ov1.1B
intel-neural-chat-7b-v3-2-ov7B
openchat-3.6-8b-20240522-ov8B
tiny-dolphin-2.8-1.1b-ov1.1B
lcm-dreamshaper-ov1.1B

Instruct

13 models
qwen2.5-0.5b-instruct-generic-gpu:4-foundry0.5B
qwen2.5-1.5b-instruct-generic-gpu:4-foundry1.5B
qwen2.5-14b-instruct-generic-gpu:4-foundry14B
qwen2.5-7b-instruct-generic-gpu:4-foundry7B
qwen2.5-coder-0.5b-instruct-generic-gpu:4-foundry0.5B
qwen2.5-coder-1.5b-instruct-generic-gpu:4-foundry1.5B
qwen2.5-coder-14b-instruct-generic-gpu:4-foundry14B
qwen2.5-coder-7b-instruct-generic-gpu:4-foundry7B
Phi-4-mini-instruct-generic-gpu:5-foundry3.8B
mistralai-Mistral-7B-Instruct-v0-2-generic-gpu:2-foundry7B
Phi-3-mini-128k-instruct-generic-gpu:2-foundry3.8B
Phi-3-mini-4k-instruct-generic-gpu:2-foundry3.8B
Phi-3.5-mini-instruct-generic-gpu:2-foundry3.8B

Instruct - ONNX

8 models
llama-3.1-instruct-onnx8B
llama-3.2-1b-instruct-onnx1.1B
llama-3.2-3b-instruct-onnx3B
llama-2-chat-onnx7B
tiny-llama-chat-onnx1.1B
phi-3-onnx3.8B
mistral-7b-instruct-v0.3-onnx7B
gemma-2b-it-onnx2B

Instruct - OV

15 models
qwen2-1.5b-instruct-ov1.5B
qwen2-7b-instruct-ov7B
qwen2.5-0.5b-instruct-ov0.5B
qwen2.5-1.5b-instruct-ov1.5B
qwen2.5-14b-instruct-ov14B
qwen2.5-32b-instruct-ov32B
qwen2.5-3b-instruct-ov3B
qwen2.5-72b-instruct-ov72B
llama-3.1-instruct-ov8B
llama-3.2-1b-instruct-ov1.1B
llama-3.2-3b-instruct-ov3B
mistral-7b-instruct-v0.2-ov7B
mistral-7b-instruct-v0.3-ov7B
mistral-nemo-instruct-2407-ov12B
mistral-small-instruct-2409-ov22B

Language Detector

1 models
xlm-roberta-language-detector-ov0.28B

Math

1 models
mathstral-7b-ov7B

Prompt Safety

7 models
protectai-prompt-injection-onnx0.3B
unitary-toxic-roberta-onnx0.1B
valurank-bias-onnx0.1B
malicious-url-detector-ov0.1B
protectai-prompt-injection-ov0.3B
unitary-toxic-roberta-ov0.1B
valurank-bias-ov0.1B

Question-answer

8 models
dragon-mistral-0.3-onnx7B
dragon-qwen-7b-ov7B
dragon-llama2-ov7B
nvidia-llama3-chatqa-1.5-8b-ov8B
dragon-mistral-0.3-ov7B
dragon-mistral-ov7B
dragon-yi-6b-ov6B
dragon-yi-9b-ov9B

Re-ranker

4 models
jina-reranker-tiny-onnx0.1B
jina-reranker-turbo-onnx0.6B
jina-reranker-v1-tiny-en-ov0.1B
jina-reranker-v1-turbo-en-ov0.6B

Text-to-speech

1 models
speech-t5-tts-ov0.6B

Vision

5 models
phi-3-vision-onnx4.2B
qwen2.5-vl-3b-ov3B
qwen2.5-vl-7b-ov7B
qwen2-vl-2b-instruct-ov2B
qwen2-vl-7b-instruct-ov7B
Catalog

NPU Models

Models designed to run on Neural Processing Units for efficient, low-power inference.

Agentic

10 models
slim-emotions-npu-ov1.1B
slim-extract-tiny-npu-ov1.1B
slim-intent-npu-ov1.1B
slim-ner-npu-ov1.1B
slim-ratings-npu-ov1.1B
slim-sentiment-npu-ov1.1B
slim-sql-npu-ov1.1B
slim-summary-tiny-npu-ov1.1B
slim-tags-npu-ov1.1B
slim-topics-npu-ov1.1B

General Chat

4 models
DeepSeek-R1-Distill-Qwen-7B-vitis-npu:2-foundry7B
Phi-4-mini-reasoning-vitis-npu:2-foundry3.8B
mistral-7b-v0.3-npu-ov7B
yi-9b-npu-ov9B

Instruct

10 models
qwen2.5-0.5b-instruct-vitis-npu:3-foundry0.5B
qwen2.5-7b-instruct-vitis-npu:2-foundry7B
qwen2.5-coder-0.5b-instruct-vitis-npu:2-foundry0.5B
qwen2.5-coder-1.5b-instruct-vitis-npu:2-foundry1.5B
qwen2.5-coder-7b-instruct-vitis-npu:2-foundry7B
Phi-4-mini-instruct-vitis-npu:2-foundry3.8B
Mistral-7B-Instruct-v0-2-vitis-npu:2-foundry7B
phi-3-mini-128k-instruct-vitis-npu:2-foundry3.8B
Phi-3-mini-4k-instruct-vitis-npu:2-foundry3.8B
phi-4-mini-instruct-vitis-npu:2-foundry3.8B
Catalog

Cloud Models

Models that run on remote servers over the internet, increasing speed and complex capabilities.

Cloud

15 models
gpt-5.2-proNA
gpt-5.2NA
gpt-5-miniNA
gpt-5-nanoNA
gpt-4.1NA
claude-opus-4-5NA
claude-haiku-4-5NA
claude-sonnet-4-5NA
claude-sonnet-4-20250514NA
claude-opus-4-20250514NA
gemini-3-pro-previewNA
gemini-3-flash-previewNA
gemini-2.5-proNA
gemini-2.5-flashNA
gemini-2.5-flash-liteNA
Next steps

Getting started with AMD models

  1. 01Ensure you have a device with compatible AMD hardware.
  2. 02Select models optimized for AMD from the Models section.
  3. 03The system automatically applies AMD optimizations when available.
  4. 04Monitor performance improvements in the system metrics.

Not sure what your hardware supports?

Check the system requirements to find the right models for your machine.

Check system requirements

For AMD-specific optimization questions, contact our technical support team at support@aibloks.com.

Reference

Encyclopedia