The model library
Every model below is open-weight, downloads directly from its publisher, and runs fully on your hardware. Footprints are for the recommended quantization; smaller and larger variants are one dropdown away.
Everyday chat
| Model | Params | Fits in | License | Good at |
| Llama 3.3 Instruct | 70B | 48 GB | Llama Community | The strongest generalist most desktops can hold |
| Llama 3.1 Instruct | 8B | 6 GB | Llama Community | Fast, dependable daily driver |
| Qwen3 Instruct | 14B | 10 GB | Apache-2.0 | Multilingual chat, strong tool calling |
| Gemma 3 | 12B | 9 GB | Gemma Terms | Concise answers on modest hardware |
Reasoning
| Model | Params | Fits in | License | Good at |
| DeepSeek-R1 Distill Qwen | 32B | 20 GB | MIT | Long-chain reasoning with visible thinking |
| DeepSeek-R1 Distill Llama | 8B | 6 GB | MIT | Reasoning on laptop-class memory |
| Qwen3 Thinking | 30B-A3B | 19 GB | Apache-2.0 | Mixture-of-experts speed at large-model quality |
Code
| Model | Params | Fits in | License | Good at |
| Qwen2.5-Coder | 32B | 20 GB | Apache-2.0 | Repository-scale completion and review |
| Qwen2.5-Coder | 7B | 5 GB | Apache-2.0 | Inline completion at editor latency |
Vision & speech
| Model | Params | Fits in | License | Good at |
| Llama 3.2 Vision | 11B | 8 GB | Llama Community | Screenshots, documents, photos |
| Whisper large-v3 | 1.5B | 3 GB | MIT | Transcription in 90+ languages, fully offline |
Model names and weights belong to their publishers (Meta, Alibaba, DeepSeek, Google, OpenAI) and are governed by their original licenses, shown in full inside AX before the first download of each model. The library is curated, not exhaustive: any GGUF file imports directly, and a Hugging Face link pasted into search does the rest. Memory figures assume the recommended quantization with 8K context; longer context costs more.