Skip to content

NVIDIA GeForce RTX 4090

NVIDIA · 24GBGDDR6X · Can run 84 models

BuyAmazon
ManufacturerNVIDIA
VRAM24 GB
Memory TypeGDDR6X
ArchitectureAda Lovelace
CUDA Cores16,384
Tensor Cores512
Bandwidth1008 GB/s
TDP450W
MSRP$1,599
ReleasedOct 12, 2022

AI Notes

The RTX 4090 remains one of the best GPUs for local AI inference. Its 24GB of GDDR6X VRAM can run 13B models at full precision and 30B+ models with quantization. The massive tensor core count delivers class-leading inference throughput among consumer GPUs.

Compatible Models

ModelParametersBest QuantVRAM UsedFitEst. Speed
Qwen 3 0.6B600MQ4_K_M2.5 GBRuns~403 tok/s
Qwen 3.5 0.8B800MQ4_K_M1.5 GBRuns~672 tok/s
Gemma 3 1B1BQ8_02 GBRuns~504 tok/s
Llama 3.2 1B1BQ8_03 GBRuns~336 tok/s
DeepSeek R1 1.5B1.5BQ8_03 GBRuns~336 tok/s
SmolLM2 1.7B1.7BQ8_02.7 GBRuns~373 tok/s
Gemma 2 2B2BQ8_04 GBRuns~252 tok/s
Gemma 3n E2B2BQ4_K_M3.3 GBRuns~305 tok/s
Gemma 4 E2B2BQ4_K_M4 GBRuns~252 tok/s
Qwen 3.5 2B2BQ4_K_M3 GBRuns~336 tok/s
Llama 3.2 3B3BQ8_05 GBRuns~202 tok/s
StarCoder2 3B3BQ4_K_M3.5 GBRuns~288 tok/s
Phi-3 Mini 3.8B3.8BQ8_05.8 GBRuns~174 tok/s
Phi-4 Mini 3.8B3.8BQ4_K_M4.5 GBRuns~224 tok/s
Gemma 3 4B4BQ4_K_M5 GBRuns~202 tok/s
Gemma 3n E4B4BQ4_K_M4.5 GBRuns~224 tok/s
Gemma 4 E4B4BQ4_K_M6 GBRuns~168 tok/s
Qwen 3 4B4BQ4_K_M4.5 GBRuns~224 tok/s
Qwen 3.5 4B4BQ4_K_M4.5 GBRuns~224 tok/s
Yi 1.5 6B6BQ4_K_M5 GBRuns~202 tok/s
Codestral Mamba 7B7BQ4_K_M6.9 GBRuns~146 tok/s
DeepSeek R1 7B7BQ8_09 GBRuns~112 tok/s
Falcon 3 7B7BQ4_K_M6.8 GBRuns~148 tok/s
InternLM 2.5 7B7BQ4_K_M5.5 GBRuns~183 tok/s
Mistral 7B7BQ8_09 GBRuns~112 tok/s
OpenChat 3.5 7B7BQ4_K_M6.9 GBRuns~146 tok/s
Qwen 2.5 7B7BQ8_09 GBRuns~112 tok/s
Qwen 2.5 Coder 7B7BQ8_09 GBRuns~112 tok/s
Qwen 2.5 VL 7B7BQ4_K_M7 GBRuns~144 tok/s
StarCoder2 7B7BQ4_K_M5.5 GBRuns~183 tok/s
WizardLM 2 7B7BQ4_K_M6.9 GBRuns~146 tok/s
Aya Expanse 8B8BQ4_K_M6.5 GBRuns~155 tok/s
Cogito 8B8BQ4_K_M7.5 GBRuns~134 tok/s
DeepSeek R1 8B8BQ4_K_M7.5 GBRuns~134 tok/s
Dolphin 3 8B8BQ4_K_M6 GBRuns~168 tok/s
Granite 3.3 8B8BQ8_010 GBRuns~101 tok/s
Llama 3.1 8B8BQ8_010 GBRuns~101 tok/s
Nemotron 3 Nano 8B8BQ4_K_M7.5 GBRuns~134 tok/s
Nous Hermes 2 8B8BQ4_K_M6 GBRuns~168 tok/s
Qwen 3 8B8BQ4_K_M7.5 GBRuns~134 tok/s
Gemma 2 9B9BQ8_011 GBRuns~92 tok/s
Qwen 3.5 9B9BQ4_K_M7.5 GBRuns~134 tok/s
Yi 1.5 9B9BQ4_K_M6.5 GBRuns~155 tok/s
Yi Coder 9B9BQ4_K_M8 GBRuns~126 tok/s
Falcon 3 10B10BQ4_K_M8.5 GBRuns~119 tok/s
Llama 3.2 Vision 11B11BQ4_K_M8.5 GBRuns~119 tok/s
Gemma 3 12B12BQ4_K_M10.5 GBRuns~96 tok/s
Mistral Nemo 12B12BQ4_K_M9.5 GBRuns~106 tok/s
DeepSeek R1 14B14BQ4_K_M9.9 GBRuns~102 tok/s
Phi-4 14B14BQ4_K_M9.9 GBRuns~102 tok/s
Phi-4 Reasoning 14B14BQ4_K_M11 GBRuns~92 tok/s
Qwen 2.5 14B14BQ4_K_M9.9 GBRuns~102 tok/s
Qwen 2.5 Coder 14B14BQ4_K_M12 GBRuns~84 tok/s
Qwen 3 14B14BQ4_K_M12 GBRuns~84 tok/s
StarCoder2 15B15BQ8_017 GBRuns~59 tok/s
InternLM 2.5 20B20BQ4_K_M12 GBRuns~84 tok/s
gpt-oss 20B21BMXFP415 GBRuns~67 tok/s
Codestral 22B22BQ4_K_M14.7 GBRuns~69 tok/s
Devstral 24B24BQ4_K_M17 GBRuns~59 tok/s
Magistral Small 24B24BQ4_K_M17 GBRuns~59 tok/s
Mistral Small 3.1 24B24BQ4_K_M18 GBRuns~56 tok/s
Gemma 4 26B26BQ4_K_M20 GBRuns~50 tok/s
Gemma 2 27B27BQ4_K_M17.7 GBRuns~57 tok/s
Gemma 3 27B27BQ4_K_M20 GBRuns~50 tok/s
Qwen 3.5 27B27BQ4_K_M19 GBRuns~53 tok/s
Qwen 3.6 27B27BQ4_K_M20 GBRuns~50 tok/s
Nous Hermes 2 34B34BQ4_K_M19 GBRuns~53 tok/s
Qwen 3.5 35B A3B35BQ4_K_M12 GBRuns~84 tok/s
Qwen 3 30B-A3B (MoE)30BQ4_K_M22 GBRuns (tight)~46 tok/s
Gemma 4 31B31BQ4_K_M22 GBRuns (tight)~46 tok/s
Aya Expanse 32B32BQ4_K_M22 GBRuns (tight)~46 tok/s
Cogito 32B32BQ4_K_M21.5 GBRuns (tight)~47 tok/s
DeepSeek R1 32B32BQ4_K_M20.7 GBRuns (tight)~49 tok/s
Qwen 2.5 32B32BQ4_K_M20.7 GBRuns (tight)~49 tok/s
QwQ 32B32BQ4_K_M21.5 GBRuns (tight)~47 tok/s
Laguna XS 2.133BQ4_K_M22 GBRuns (tight)~46 tok/s
WizardCoder 33B33BQ4_K_M22 GBRuns (tight)~46 tok/s
Yi 1.5 34B34BQ4_K_M21 GBRuns (tight)~48 tok/s
Command R 35B35BQ4_K_M22.5 GBRuns (tight)~45 tok/s
Qwen 2.5 Coder 32B32BQ4_K_M23 GBCPU Offload~13 tok/s
Qwen 3 32B32BQ4_K_M23 GBCPU Offload~13 tok/s
Qwen 3.6 35B-A3B35BQ4_K_M27 GBCPU Offload~11 tok/s
Dolphin Mixtral 8x7B47BQ4_K_M26 GBCPU Offload~12 tok/s
Mixtral 8x7B47BQ4_K_M29.7 GBCPU Offload~10 tok/s
30 model(s) are too large for this hardware.