Model HQ

Hardware optimized

Intel Supported Models

A complete catalog of 219 AI models optimized for Intel processors with the OpenVINO runtime — enhanced performance across Intel CPUs, GPUs, and NPUs.

Hover the Encyclopedia tab on the right, or tap any icon to learn what a model type does.
219
Optimized models
OpenVINO
Runtime
CPU · GPU · NPU
Hardware targets
Why Intel

Intel optimization features

OpenVINO-optimized models deliver faster inference with a smaller footprint on Intel hardware.

Performance benefits

  • Optimized for Intel CPU, GPU, and NPU architectures
  • Enhanced inference speed with the OpenVINO runtime
  • Reduced memory footprint and power consumption
  • Hardware-specific optimizations for the latest Intel chips

Supported hardware

  • Intel Core (Meteor Lake, Lunar Lake, Arrow Lake)
  • Intel Xeon processors
  • Intel Arc GPUs
  • Intel Neural Processing Units (NPUs)
Catalog

CPU Models

GGUF and tool models that run on the CPU alone — no GPU or NPU required.

Agentic

29 models
slim-extract-phi-3-gguf3.8B
slim-extract-qwen-1.5b-gguf1.5B
slim-extract-qwen-nano-gguf0.5B
slim-summary-phi-3-gguf1.1B
slim-xsum-phi-3-gguf3.8B
slim-boolean-phi-3-gguf3.8B
slim-sa-ner-phi-3-gguf3.8B
slim-q-gen-phi-3-tool3.8B
slim-qa-gen-phi-3-tool3.8B
slim-ner-tool1.1B
slim-sentiment-tool1.1B
slim-emotions-tool1.1B
slim-ratings-tool1.1B
slim-intent-tool1.1B
slim-nli-tool1.1B
slim-topics-tool1.1B
slim-tags-tool1.1B
slim-sql-tool1.1B
slim-category-tool1.1B
slim-xsum-tool1.1B
slim-extract-tool2.8B
slim-extract-tiny-tool1.1B
slim-summary-tiny-tool1.1B
slim-boolean-tool2.8B
slim-sa-ner-tool2.8B
slim-tags-3b-tool2.8B
slim-summary-tool2.8B
slim-q-gen-tiny-tool1.1B
slim-qa-gen-tiny-tool1.1B

Coding

1 models
qwen-2.5-7b-coder-gguf7B

General Chat

18 models
qwen3-1.7b-gguf1.7B
qwen3-8b-gguf8B
qwen3-14b-gguf14B
qwen2.5-32b-gguf32B
deepseek-qwen-14b-gguf14B
deepseek-qwen-7b-gguf7B
dragon-llama-3.1-gguf8B
phi-3-gguf3.8B
phi-3.5-gguf3.8B
phi-4-gguf14B
phi-4-mini-gguf4B
phi-4-mini-reasoning-gguf4B
mistral-3.2-24b-gguf24B
openhermes-2.5-mistral-7b-gguf7.3B
zephyr-7b-beta-gguf7.3B
starling-lm-7b-alpha-gguf7B
gemma-3-4b-gguf4B
gemma-3-12b-gguf12B

Instruct

8 models
qwen2-7B-instruct-gguf7B
qwen3-4b-instruct-gguf4B
qwen2-1.5b-instruct-gguf1.5B
qwen2-0.5b-instruct-gguf0.5B
qwen-2.5-14b-instruct-gguf14B
mistral-7b-instruct-v0.3-gguf7.3B
gemma-2-9b-instruct-gguf9B
gemma-2-27b-instruct-gguf27B

Question-answer

18 models
bling-qwen-0.5b-gguf0.5B
dragon-qwen-7b-gguf7B
bling-phi-3-gguf3.8B
bling-phi-3.5-gguf3.8B
dragon-mistral-0.3-gguf7.3B
dragon-yi-9b-gguf8.8B
bling-stablelm-3b-gguf2.8B
bling-qwen-mini-tool1.5B
bling-answer-tool1.1B
dragon-mistral-answer-tool7.3B
dragon-yi-answer-tool5.8B
dragon-llama-answer-tool7B
bling-tiny-llama-onnx1.1B
bling-phi-3-onnx3.8B
bling-phi-3-ov3.8B
bling-qwen-500m-ov0.5B
bling-qwen-1.5b-ov1.5B
bling-tiny-llama-ov1.1B
Catalog

GPU/CPU/NPU Models

OpenVINO (OV) and ONNX models that run on the available GPU, CPU, or NPU.

Agentic

35 models
slim-sentiment-onnx1.1B
slim-extract-tiny-onnx1.1B
slim-summary-tiny-onnx1.1B
slim-sql-onnx1.1B
slim-emotions-onnx1.1B
slim-topics-onnx1.1B
slim-ner-onnx1.1B
slim-intent-onnx1.1B
slim-tags-onnx1.1B
slim-ratings-onnx1.1B
slim-summary-phi-3-onnx3.8B
slim-extract-phi-3-onnx3.8B
slim-boolean-phi-3-onnx3.8B
slim-sentiment-ov1.1B
slim-xsum-phi-3-ov3.8B
slim-boolean-phi-3-ov3.8B
slim-sa-ner-phi-3-ov3.8B
slim-summary-phi-3-ov3.8B
slim-sql-qwen-base-ov2B
slim-sql-phi-3-ov3.8B
slim-extract-phi-3-ov3.8B
slim-extract-tiny-ov1.1B
slim-extract-qwen-0.5b-ov0.5B
slim-extract-qwen-1.5b-ov1.5B
slim-summary-tiny-ov1.1B
slim-sql-ov1.1B
slim-emotions-ov1.1B
slim-topics-ov1.1B
slim-ner-ov1.1B
slim-intent-ov1.1B
slim-category-ov1.1B
slim-tags-ov1.1B
slim-ratings-ov1.1B
slim-q-gen-tiny-ov1.1B
slim-qa-gen-tiny-ov1.1B

Coding

2 models
qwen2.5-coder-7b-instruct-ov7B
codegemma-7b-it-ov7B

Embedding

14 models
industry-bert-contracts-ovNA
industry-bert-insurance-ovNA
industry-bert-asset-management-ovNA
industry-bert-sec-ovNA
industry-bert-loans-ovNA
all-mini-lm-l6-v2-ovNA
all-mpnet-base-v2-ovNA
paraphrase-multilingual-MiniLM-L12-v2-ovNA
gte-small-ovNA
gte-base-ovNA
gte-large-ovNA
bge-small-en-v1.5-ovNA
bge-base-en-v1.5-ovNA
bge-large-en-v1.5-ovNA

General Chat

26 models
llama-2-chat-onnx7B
qwen2-0.5b-chat-ov0.5B
phi-3-ov3.8B
phi-4-ov14B
phi-4-mini-ov4B
qwen3-8b-ov8B
qwen3-1.7b-ov1.7B
qwen3-4b-ov4B
qwen3-14b-ov14B
llama-2-chat-ov7B
llama-2-13b-chat-ov13B
tiny-llama-chat-ov1.1B
dolphin-2.9.3-mistral-7b-32k-ov7.3B
dolphin-2.9.4-llama3.1-8b-ov8B
zephyr-mistral-7b-chat-ov7.3B
teknium-open-hermes-2.5-mistral-ov7.3B
yi-9b-chat-ov8.8B
yi-6b-1.5v-chat-ov5.8B
stablelm-zephyr-3b-ov2.8B
stablelm-2-zephyr-1_6b-ov1.6B
stablelm-2-12b-chat-ov12B
intel-neural-chat-7b-v3-2-ov7B
tiny-dolphin-2.8-1.1b-ov1.1B
dreamgen-wizardlm-2-7b-ov7B
openchat-3.6-8b-20240522-ov8B
lcm-dreamshaper-ovNA

Instruct

25 models
llama-3.1-instruct-onnx8B
llama-3.2-1b-instruct-onnx1.3B
llama-3.2-3b-instruct-onnx3B
phi-3-onnx3.8B
mistral-7b-instruct-v0.3-onnx7.3B
gemma-2b-it-onnx2B
qwen2-vl-2b-instruct-ov2B
qwen2-vl-7b-instruct-ov7B
qwen2-7b-instruct-ov7B
qwen2-1.5b-instruct-ov1.5B
qwen2.5-1.5b-instruct-ov1.5B
qwen2.5-0.5b-instruct-ov0.5B
qwen2.5-3b-instruct-ov3B
qwen2.5-14b-instruct-ov14B
qwen2.5-32b-instruct-ov32B
qwen2.5-72b-instruct-ov72B
llama-3.1-instruct-ov8B
llama-3.2-3b-instruct-ov3B
llama-3.2-1b-instruct-ov1.1B
mistral-7b-instruct-v0.3-ov7.3B
mistral-small-instruct-2409-ov22B
mistral-nemo-instruct-2407-ov12B
mistral-7b-instruct-v0.2-ov7.3B
gemma-7b-it-ov7B
gemma-2b-it-ov2B

Math

1 models
mathstral-7b-ov7.3B

Prompt Safety

8 models
protectai-prompt-injection-onnxNA
valurank-bias-onnxNA
unitary-toxic-roberta-onnxNA
protectai-prompt-injection-ovNA
malicious-url-detector-ovNA
xlm-roberta-language-detector-ovNA
valurank-bias-ovNA
unitary-toxic-roberta-ovNA

Question-answer

6 models
dragon-mistral-0.3-onnx7.3B
dragon-qwen-7b-ov7B
dragon-llama2-ov7B
nvidia-llama3-chatqa-1.5-8b-ov8B
dragon-mistral-ov7.3B
dragon-mistral-0.3-ov7.3B

Re-ranker

4 models
jina-reranker-tiny-onnxNA
jina-reranker-turbo-onnxNA
jina-reranker-v1-tiny-en-ovNA
jina-reranker-v1-turbo-en-ovNA

Text-to-speech

1 models
speech-t5-tts-ovNA

Vision

2 models
phi-3-vision-onnx3.8B
llama-11b-vision-instruct-ov11B
Catalog

NPU Models

Models designed to run on Neural Processing Units for efficient, low-power inference.

NPU Models

21 models
bling-tiny-llama-npu-ov1.1B
llama-3.1-8b-instruct-npu-ov8B
llama-3.2-3b-instruct-npu-ov3B
llama-3.2-1b-instruct-npu-ov1.1B
phi-3-npu-ov3.8B
phi-4-mini-npu-ov4B
phi-4-npu-ov14B
mistral-7b-v0.3-npu-ov7.3B
yi-9b-npu-ov8.8B
slim-sentiment-npu-ov1.1B
slim-extract-tiny-npu-ov1.1B
slim-summary-tiny-npu-ov1.1B
slim-sql-npu-ov1.1B
slim-emotions-npu-ov1.1B
slim-topics-npu-ov1.1B
slim-ner-npu-ov1.1B
slim-intent-npu-ov1.1B
slim-tags-npu-ov1.1B
slim-ratings-npu-ov1.1B
llama-3.2-3b-onnx-qnn3B
phi-3.5-onnx-qnnNA
Next steps

Getting started with Intel models

  1. 01Ensure you have an Intel processor with OpenVINO support.
  2. 02Select models with the -ov suffix from the Models section.
  3. 03The system automatically applies Intel optimizations when available.
  4. 04Monitor performance improvements in the system metrics.

Not sure what your hardware supports?

Check the system requirements to find the right models for your machine.

Check system requirements

For Intel-specific optimization questions, contact our technical support team at support@aibloks.com.

Reference

Encyclopedia