The model library

Every model below is open-weight, downloads directly from its publisher, and runs fully on your hardware. Footprints are for the recommended quantization; smaller and larger variants are one dropdown away.

Everyday chat

ModelParamsFits inLicenseGood at
Llama 3.3 Instruct70B48 GBLlama CommunityThe strongest generalist most desktops can hold
Llama 3.1 Instruct8B6 GBLlama CommunityFast, dependable daily driver
Qwen3 Instruct14B10 GBApache-2.0Multilingual chat, strong tool calling
Gemma 312B9 GBGemma TermsConcise answers on modest hardware

Reasoning

ModelParamsFits inLicenseGood at
DeepSeek-R1 Distill Qwen32B20 GBMITLong-chain reasoning with visible thinking
DeepSeek-R1 Distill Llama8B6 GBMITReasoning on laptop-class memory
Qwen3 Thinking30B-A3B19 GBApache-2.0Mixture-of-experts speed at large-model quality

Code

ModelParamsFits inLicenseGood at
Qwen2.5-Coder32B20 GBApache-2.0Repository-scale completion and review
Qwen2.5-Coder7B5 GBApache-2.0Inline completion at editor latency

Vision & speech

ModelParamsFits inLicenseGood at
Llama 3.2 Vision11B8 GBLlama CommunityScreenshots, documents, photos
Whisper large-v31.5B3 GBMITTranscription in 90+ languages, fully offline

Model names and weights belong to their publishers (Meta, Alibaba, DeepSeek, Google, OpenAI) and are governed by their original licenses, shown in full inside AX before the first download of each model. The library is curated, not exhaustive: any GGUF file imports directly, and a Hugging Face link pasted into search does the rest. Memory figures assume the recommended quantization with 8K context; longer context costs more.