Model HQ

Hardware optimized

NVIDIA Supported Models

A complete catalog of 122 AI models optimized optimized for NVIDIA GPUs and accelerators.

Hover the Encyclopedia tab on the right, or tap any icon to learn what a model type does.
122
Optimized models
CUDA
Runtime
NVIDIA GPU
Hardware target
Why NVIDIA

NVIDIA optimization features

These models are optimized for enhanced performance on NVIDIA GB10 devices (DGX Spark).

Performance benefits

  • Optimized for NVIDIA GPUs such as GB10, GeForce, and RTX
  • Enhanced AI inference performance with NVIDIA GPUs
  • High-throughput execution for demanding AI workloads
  • Hardware-specific optimizations for latest NVIDIA GPUs

Supported hardware

  • NVIDIA DGX Spark
  • NVIDIA GB10 GPUs
  • NVIDIA GeForce GPUs
  • NVIDIA RTX GPUs
Catalog

All Supported Models

Models optimized for NVIDIA GPUs and accelerators

Agentic

29 models
slim-ner-tool1.1B
slim-sentiment-tool1.1B
slim-emotions-tool1.1B
slim-ratings-tool1.1B
slim-intent-tool1.1B
slim-nli-tool1.1B
slim-topics-tool1.1B
slim-tags-tool1.1B
slim-sql-tool1.1B
slim-category-tool1.1B
slim-xsum-tool1.1B
slim-extract-tool1.1B
slim-extract-phi-3-gguf3.8B
slim-extract-qwen-1.5b-gguf1.5B
slim-extract-qwen-nano-gguf0.5B
slim-extract-tiny-tool1.1B
slim-summary-tiny-tool1.1B
slim-summary-phi-3-gguf3.8B
slim-xsum-phi-3-gguf3.8B
slim-boolean-tool1.1B
slim-boolean-phi-3-gguf3.8B
slim-sa-ner-phi-3-gguf3.8B
slim-sa-ner-tool1.1B
slim-tags-3b-tool3B
slim-summary-tool1.1B
slim-q-gen-phi-3-tool3.8B
slim-q-gen-tiny-tool1.1B
slim-qa-gen-tiny-tool1.1B
slim-qa-gen-phi-3-tool3.8B

Cloud

15 models
claude-opus-4-5NA
claude-haiku-4-5NA
claude-sonnet-4-5NA
claude-sonnet-4-20250514NA
claude-opus-4-20250514NA
gemini-3-pro-previewNA
gemini-3-flash-previewNA
gemini-2.5-proNA
gemini-2.5-flashNA
gemini-2.5-flash-liteNA
gpt-5.2-proNA
gpt-5.2NA
gpt-5-miniNA
gpt-5-nanoNA
gpt-4.1NA

Coding

1 models
qwen-2.5-7b-coder-gguf7B

Embedding

8 models
all-mini-lm-L6-v20.02B
all-mpnet-base-v20.1B
industry-bert-insurance0.1B
industry-bert-contracts0.1B
industry-bert-asset-management0.1B
industry-bert-sec0.1B
industry-bert-loans0.1B
nomic-ai/nomic-embed-text-v10.1B

General Chat

34 models
llama-2-7b-chat-gguf7B
dragon-llama-3.1-gguf8B
tiny-llama-chat-gguf1.1B
qwen3-1.7b-gguf1.7B
qwen3-8b-gguf8B
qwen3-14b-gguf14B
qwen-3.5-4b-gguf4B
qwen-3.5-9b-gguf9B
qwen-3.5-27b-gguf27B
qwen-3.5-35b-a3b-gguf35B
qwen2.5-32b-gguf32B
qwen2.5-72b-gguf72B
deepseek-qwen-14b-gguf14B
deepseek-qwen-7b-gguf7B
phi-3-gguf3.8B
phi-3.5-gguf3.8B
phi-4-gguf14B
phi-4-mini-gguf3.8B
phi-4-mini-reasoning-gguf3.8B
mistral-small-3.2-24b-gguf24B
ministral-3-14b-gguf14B
openhermes-2.5-mistral-7b-gguf7B
zephyr-7b-beta-gguf7B
starling-lm-7b-alpha-gguf7B
gemma-3-4b-gguf4B
gemma-3-12b-gguf12B
gemma-4-4b-gguf4B
gemma-4-2b-gguf2B
gemma-4-26b-gguf26B
gpt-oss-20b-gguf20B
olmo-13b-gguf13B
granite-4-micro-gguf1.1B
liquidai-lfm2-2.6b-gguf2.6B
minicpm-2.6-gguf4B

Instruct

13 models
gemma-2-9b-instruct-gguf9B
gemma-2-27b-instruct-gguf27B
llama-3.1-instruct-gguf8B
llama-3-8b-instruct-gguf8B
llama-3.2-1b-instruct-gguf1.1B
llama-3.2-3b-instruct-gguf3B
mistral-7b-instruct-v0.3-gguf7B
qwen2.5-vl-3b-instruct-gguf3B
qwen2-7B-instruct-gguf7B
qwen3-4b-instruct-gguf4B
qwen2-1.5b-instruct-gguf1.5B
qwen2-0.5b-instruct-gguf0.5B
qwen-2.5-14b-instruct-gguf14B

Re-ranker

6 models
jina-reranker-tiny-ppt0.1B
jina-reranker-turbo-ppt0.6B
jina-reranker-tiny-onnx0.1B
jina-reranker-turbo-onnx0.6B
jina-reranker-v1-turbo-en0.6B
jina-reranker-v1-tiny-en0.1B

Speech-to-text

1 models
whisper-cpp-base-english0.07B

Question-answer

11 models
bling-qwen-mini-tool1.5B
dragon-qwen-7b-gguf7B
bling-phi-3-gguf3.8B
bling-phi-3.5-gguf3.8B
dragon-mistral-0.3-gguf7B
dragon-yi-9b-gguf9B
dragon-yi-answer-tool6B
bling-stablelm-3b-gguf3B
bling-answer-tool1.1B
dragon-llama-answer-tool7B
dragon-mistral-answer-tool7B

Vision

4 models
qwen2.5-vl-3b-instruct-gguf3B
qwen3-vl-8b-gguf8B
qwen3-vl-4b-gguf4B
qwen3-vl-30b-gguf30B
Next steps

Getting started with NVIDIA models

  1. 01Ensure you have a device with compatible NVIDIA GPU.
  2. 02Select models optimized for NVIDIA from the Models section.
  3. 03The system automatically applies NVIDIA optimizations when available.
  4. 04Monitor performance improvements in the system metrics.

Not sure what your hardware supports?

Check the system requirements to find the right models for your machine.

Check system requirements

For NVIDIA-specific optimization questions, contact our technical support team at support@aibloks.com.

Reference

Encyclopedia