Model HQ
DocumentationHardware optimized
Intel Supported Models
A complete catalog of 219 AI models optimized for Intel processors with the OpenVINO runtime — enhanced performance across Intel CPUs, GPUs, and NPUs.
Hover the Encyclopedia tab on the right, or tap any icon to learn what a model type does.
219
Optimized models
OpenVINO
Runtime
CPU · GPU · NPU
Hardware targets
Why Intel
Intel optimization features
OpenVINO-optimized models deliver faster inference with a smaller footprint on Intel hardware.
Performance benefits
- Optimized for Intel CPU, GPU, and NPU architectures
- Enhanced inference speed with the OpenVINO runtime
- Reduced memory footprint and power consumption
- Hardware-specific optimizations for the latest Intel chips
Supported hardware
- Intel Core (Meteor Lake, Lunar Lake, Arrow Lake)
- Intel Xeon processors
- Intel Arc GPUs
- Intel Neural Processing Units (NPUs)
Catalog
CPU Models
GGUF and tool models that run on the CPU alone — no GPU or NPU required.
Agentic
29 modelsslim-extract-phi-3-gguf3.8Bslim-extract-qwen-1.5b-gguf1.5Bslim-extract-qwen-nano-gguf0.5Bslim-summary-phi-3-gguf1.1Bslim-xsum-phi-3-gguf3.8Bslim-boolean-phi-3-gguf3.8Bslim-sa-ner-phi-3-gguf3.8Bslim-q-gen-phi-3-tool3.8Bslim-qa-gen-phi-3-tool3.8Bslim-ner-tool1.1Bslim-sentiment-tool1.1Bslim-emotions-tool1.1Bslim-ratings-tool1.1Bslim-intent-tool1.1Bslim-nli-tool1.1Bslim-topics-tool1.1Bslim-tags-tool1.1Bslim-sql-tool1.1Bslim-category-tool1.1Bslim-xsum-tool1.1Bslim-extract-tool2.8Bslim-extract-tiny-tool1.1Bslim-summary-tiny-tool1.1Bslim-boolean-tool2.8Bslim-sa-ner-tool2.8Bslim-tags-3b-tool2.8Bslim-summary-tool2.8Bslim-q-gen-tiny-tool1.1Bslim-qa-gen-tiny-tool1.1BCoding
1 modelsqwen-2.5-7b-coder-gguf7BGeneral Chat
18 modelsqwen3-1.7b-gguf1.7Bqwen3-8b-gguf8Bqwen3-14b-gguf14Bqwen2.5-32b-gguf32Bdeepseek-qwen-14b-gguf14Bdeepseek-qwen-7b-gguf7Bdragon-llama-3.1-gguf8Bphi-3-gguf3.8Bphi-3.5-gguf3.8Bphi-4-gguf14Bphi-4-mini-gguf4Bphi-4-mini-reasoning-gguf4Bmistral-3.2-24b-gguf24Bopenhermes-2.5-mistral-7b-gguf7.3Bzephyr-7b-beta-gguf7.3Bstarling-lm-7b-alpha-gguf7Bgemma-3-4b-gguf4Bgemma-3-12b-gguf12BInstruct
8 modelsqwen2-7B-instruct-gguf7Bqwen3-4b-instruct-gguf4Bqwen2-1.5b-instruct-gguf1.5Bqwen2-0.5b-instruct-gguf0.5Bqwen-2.5-14b-instruct-gguf14Bmistral-7b-instruct-v0.3-gguf7.3Bgemma-2-9b-instruct-gguf9Bgemma-2-27b-instruct-gguf27BQuestion-answer
18 modelsbling-qwen-0.5b-gguf0.5Bdragon-qwen-7b-gguf7Bbling-phi-3-gguf3.8Bbling-phi-3.5-gguf3.8Bdragon-mistral-0.3-gguf7.3Bdragon-yi-9b-gguf8.8Bbling-stablelm-3b-gguf2.8Bbling-qwen-mini-tool1.5Bbling-answer-tool1.1Bdragon-mistral-answer-tool7.3Bdragon-yi-answer-tool5.8Bdragon-llama-answer-tool7Bbling-tiny-llama-onnx1.1Bbling-phi-3-onnx3.8Bbling-phi-3-ov3.8Bbling-qwen-500m-ov0.5Bbling-qwen-1.5b-ov1.5Bbling-tiny-llama-ov1.1BCatalog
GPU/CPU/NPU Models
OpenVINO (OV) and ONNX models that run on the available GPU, CPU, or NPU.
Agentic
35 modelsslim-sentiment-onnx1.1Bslim-extract-tiny-onnx1.1Bslim-summary-tiny-onnx1.1Bslim-sql-onnx1.1Bslim-emotions-onnx1.1Bslim-topics-onnx1.1Bslim-ner-onnx1.1Bslim-intent-onnx1.1Bslim-tags-onnx1.1Bslim-ratings-onnx1.1Bslim-summary-phi-3-onnx3.8Bslim-extract-phi-3-onnx3.8Bslim-boolean-phi-3-onnx3.8Bslim-sentiment-ov1.1Bslim-xsum-phi-3-ov3.8Bslim-boolean-phi-3-ov3.8Bslim-sa-ner-phi-3-ov3.8Bslim-summary-phi-3-ov3.8Bslim-sql-qwen-base-ov2Bslim-sql-phi-3-ov3.8Bslim-extract-phi-3-ov3.8Bslim-extract-tiny-ov1.1Bslim-extract-qwen-0.5b-ov0.5Bslim-extract-qwen-1.5b-ov1.5Bslim-summary-tiny-ov1.1Bslim-sql-ov1.1Bslim-emotions-ov1.1Bslim-topics-ov1.1Bslim-ner-ov1.1Bslim-intent-ov1.1Bslim-category-ov1.1Bslim-tags-ov1.1Bslim-ratings-ov1.1Bslim-q-gen-tiny-ov1.1Bslim-qa-gen-tiny-ov1.1BCoding
2 modelsqwen2.5-coder-7b-instruct-ov7Bcodegemma-7b-it-ov7BEmbedding
14 modelsindustry-bert-contracts-ovNAindustry-bert-insurance-ovNAindustry-bert-asset-management-ovNAindustry-bert-sec-ovNAindustry-bert-loans-ovNAall-mini-lm-l6-v2-ovNAall-mpnet-base-v2-ovNAparaphrase-multilingual-MiniLM-L12-v2-ovNAgte-small-ovNAgte-base-ovNAgte-large-ovNAbge-small-en-v1.5-ovNAbge-base-en-v1.5-ovNAbge-large-en-v1.5-ovNAGeneral Chat
26 modelsllama-2-chat-onnx7Bqwen2-0.5b-chat-ov0.5Bphi-3-ov3.8Bphi-4-ov14Bphi-4-mini-ov4Bqwen3-8b-ov8Bqwen3-1.7b-ov1.7Bqwen3-4b-ov4Bqwen3-14b-ov14Bllama-2-chat-ov7Bllama-2-13b-chat-ov13Btiny-llama-chat-ov1.1Bdolphin-2.9.3-mistral-7b-32k-ov7.3Bdolphin-2.9.4-llama3.1-8b-ov8Bzephyr-mistral-7b-chat-ov7.3Bteknium-open-hermes-2.5-mistral-ov7.3Byi-9b-chat-ov8.8Byi-6b-1.5v-chat-ov5.8Bstablelm-zephyr-3b-ov2.8Bstablelm-2-zephyr-1_6b-ov1.6Bstablelm-2-12b-chat-ov12Bintel-neural-chat-7b-v3-2-ov7Btiny-dolphin-2.8-1.1b-ov1.1Bdreamgen-wizardlm-2-7b-ov7Bopenchat-3.6-8b-20240522-ov8Blcm-dreamshaper-ovNAInstruct
25 modelsllama-3.1-instruct-onnx8Bllama-3.2-1b-instruct-onnx1.3Bllama-3.2-3b-instruct-onnx3Bphi-3-onnx3.8Bmistral-7b-instruct-v0.3-onnx7.3Bgemma-2b-it-onnx2Bqwen2-vl-2b-instruct-ov2Bqwen2-vl-7b-instruct-ov7Bqwen2-7b-instruct-ov7Bqwen2-1.5b-instruct-ov1.5Bqwen2.5-1.5b-instruct-ov1.5Bqwen2.5-0.5b-instruct-ov0.5Bqwen2.5-3b-instruct-ov3Bqwen2.5-14b-instruct-ov14Bqwen2.5-32b-instruct-ov32Bqwen2.5-72b-instruct-ov72Bllama-3.1-instruct-ov8Bllama-3.2-3b-instruct-ov3Bllama-3.2-1b-instruct-ov1.1Bmistral-7b-instruct-v0.3-ov7.3Bmistral-small-instruct-2409-ov22Bmistral-nemo-instruct-2407-ov12Bmistral-7b-instruct-v0.2-ov7.3Bgemma-7b-it-ov7Bgemma-2b-it-ov2BMath
1 modelsmathstral-7b-ov7.3BPrompt Safety
8 modelsprotectai-prompt-injection-onnxNAvalurank-bias-onnxNAunitary-toxic-roberta-onnxNAprotectai-prompt-injection-ovNAmalicious-url-detector-ovNAxlm-roberta-language-detector-ovNAvalurank-bias-ovNAunitary-toxic-roberta-ovNAQuestion-answer
6 modelsdragon-mistral-0.3-onnx7.3Bdragon-qwen-7b-ov7Bdragon-llama2-ov7Bnvidia-llama3-chatqa-1.5-8b-ov8Bdragon-mistral-ov7.3Bdragon-mistral-0.3-ov7.3BRe-ranker
4 modelsjina-reranker-tiny-onnxNAjina-reranker-turbo-onnxNAjina-reranker-v1-tiny-en-ovNAjina-reranker-v1-turbo-en-ovNAText-to-speech
1 modelsspeech-t5-tts-ovNAVision
2 modelsphi-3-vision-onnx3.8Bllama-11b-vision-instruct-ov11BCatalog
NPU Models
Models designed to run on Neural Processing Units for efficient, low-power inference.
NPU Models
21 modelsbling-tiny-llama-npu-ov1.1Bllama-3.1-8b-instruct-npu-ov8Bllama-3.2-3b-instruct-npu-ov3Bllama-3.2-1b-instruct-npu-ov1.1Bphi-3-npu-ov3.8Bphi-4-mini-npu-ov4Bphi-4-npu-ov14Bmistral-7b-v0.3-npu-ov7.3Byi-9b-npu-ov8.8Bslim-sentiment-npu-ov1.1Bslim-extract-tiny-npu-ov1.1Bslim-summary-tiny-npu-ov1.1Bslim-sql-npu-ov1.1Bslim-emotions-npu-ov1.1Bslim-topics-npu-ov1.1Bslim-ner-npu-ov1.1Bslim-intent-npu-ov1.1Bslim-tags-npu-ov1.1Bslim-ratings-npu-ov1.1Bllama-3.2-3b-onnx-qnn3Bphi-3.5-onnx-qnnNANext stepsCheck system requirements
Getting started with Intel models
- 01Ensure you have an Intel processor with OpenVINO support.
- 02Select models with the
-ovsuffix from the Models section. - 03The system automatically applies Intel optimizations when available.
- 04Monitor performance improvements in the system metrics.
Not sure what your hardware supports?
Check the system requirements to find the right models for your machine.
For Intel-specific optimization questions, contact our technical support team at support@aibloks.com.
Reference
