Models that run on your device

48 downloadable variants across 26 model families - from 2 GB laptop models to frontier-class mixtures of experts. Everything here runs fully offline, keeps your conversations on your machine, and needs no account for local use.

Muse GlimmerRecommended

Meta · Aug 2026

Meta's new agent-first model - built for reliable tool use, working in project folders, and recovering from its own mistakes. Sees images too. Brand-new: early support, expect rough edges.

30Bfrom 12.4 GBruns from 16 GB RAM128K contextsees imagesMLXApache 2.0

Qwen 3.8Recommended

Alibaba's Qwen team · Aug 2026

Alibaba's newest. Frontier coding, math, and reasoning with built-in thinking; a hybrid attention design keeps long documents fast. The strongest model here for 24GB-class graphics cards.

27Bfrom 16.5 GBruns from 32 GB RAM256K contextsees imagesMLXApache 2.0

Qwen 3.6Recommended

Alibaba's Qwen team · Apr 2026

Top-tier coding, math, and reasoning. The 35B MoE runs fast for its size — only 3B parameters active per token.

27B · 35B-A3B (MoE)from 12.3 GBruns from 24 GB RAM256K contextMLXApache 2.0

Gemma 4Recommended

Google · Apr 2026

Google's latest open model. Strong writing, analysis, and instruction-following. E2B/E4B run on modest machines; the 26B is a fast MoE.

E2B · E4B · 12B · 26B-A4B (MoE) · 31Bfrom 3.1 GBruns from 8 GB RAM128K contextsees imagesMLXApache 2.0 (Gemma 4 terms)

Ministral 3

Mistral AI · Dec 2025

Mistral's latest small model family. Fast inference, great for on-device use.

3B · 8B · 14Bfrom 2.2 GBruns from 8 GB RAM256K contextMLXApache 2.0

Phi-4 MiniRecommended

Microsoft · Feb 2025

Microsoft's compact model. Best option for machines with limited RAM.

3.8Bfrom 2.5 GBruns from 8 GB RAM128K contextMLXMIT

GPT-OSS (OpenAI)

OpenAI · Aug 2025

OpenAI's open-weight model. Strong reasoning, agentic tasks, and function calling.

20B (3.6B active) · 120B (5.1B active)from 11.6 GBruns from 16 GB RAM128K contextApache 2.0

LFM2.5Recommended

Liquid AI · May 2026

Liquid AI's small mixture-of-experts - 8B total, 1.5B active per token, so it runs fast on ordinary machines. Strong instruction following and tool use for its size; 10 languages.

8B-A1B (MoE)from 5.2 GBruns from 12 GB RAM125K contextLFM Open License v1.0

Granite 4.0 TinyRecommended

IBM · Sep 2025

IBM's small hybrid mixture-of-experts - 7B total, about 1B active per token, with a 1M-token context. Apache-2.0. A fast, permissive everyday model for long documents.

7B-A1B (MoE)from 4.2 GBruns from 8 GB RAM1M contextApache 2.0

Nemotron 3.5 LightningRecommended

NVIDIA · Aug 2026

NVIDIA's newest open reasoning model - a 30B mixture-of-experts with 3B active per token and a 256K context. Strong reasoning, coding, and tool use; runs on a 32 GB machine with the experts in main memory.

30B-A3B (MoE)from 18.9 GBruns from 32 GB RAM256K contextOpenMDW 1.1

Ling-mini 2.0

Ant Group (inclusionAI) · Sep 2025

Ant Group's mid-size mixture-of-experts - 16B total, 1.4B active per token. MIT-licensed, 128K context; a quick all-rounder for 24 GB machines.

16B-A1.4B (MoE)from 9.9 GBruns from 24 GB RAM128K contextMIT

DeepSeek V4 Flash

DeepSeek · Jul 2026

DeepSeek's V4 Flash - a 284B mixture-of-experts with 13B active per token and a 1M-token context. Frontier-class reasoning and agentic work on a workstation with a big graphics card and 128 GB or more of main memory.

284B-A13B (MoE)from 96.9 GBruns from 128 GB RAM1M contextMIT

GLM-5.2

Zhipu AI · Jun 2026

Zhipu's GLM-5.2 - a 753B mixture-of-experts with a 1M-token context. For very large workstations: a big graphics card and 384 GB or more of main memory, experts in main memory.

753B (MoE)from 253.9 GBruns from 384 GB RAM1M contextMIT

DeepSeek R1 0528

DeepSeek · May 2025

Powerful chain-of-thought reasoning. Excels at complex problem solving, math, and logic.

8Bfrom 5.0 GBruns from 16 GB RAM128K contextMIT

Devstral Small 2

Mistral AI · Dec 2025

Mistral's agentic coding model. Tops open models on SWE-bench at its size — purpose-built for software engineering and code agents.

24Bfrom 14.3 GBruns from 16 GB RAM384K contextApache 2.0

Ornith 1.5

DeepReinforce · Aug 2026

DeepReinforce's agentic coding model, self-improved with RL - 1.5 extends the loop to generating its own training tasks. State-of-the-art among open coders at its size, purpose-built for terminal coding agents and tool use. The 9B runs on modest machines; the 35B is a mixture-of-experts (3B active per token) that runs fast for its size and adds vision.

9B · 35B-A3B (MoE)from 5.6 GBruns from 16 GB RAM256K contextMLXMIT

GLM-4.7 Flash

Zhipu (Z.ai) · Jan 2026

Zhipu's fast GLM model — a strong all-rounder tuned for agentic coding, reasoning, and tool use. The lighter "Flash" tier of the GLM family; the 30B MoE runs only 3B parameters per token.

30B-A3B (MoE)from 18.3 GBruns from 32 GB RAM198K contextMLXMIT

Qwen3-Coder

Alibaba's Qwen team · Jul 2025

Alibaba's agentic coding model — a 30B MoE with only 3B active per token. Tuned for repository-scale work and tool use (Qwen Code, Cline).

30B-A3B (MoE)from 18.6 GBruns from 32 GB RAM256K contextMLXApache 2.0

Qwen 3.5 Opus Distilled

Jackrong · Mar 2026

Qwen 3.5 distilled from Claude Opus reasoning traces. Enhanced chain-of-thought capabilities.

4B · 9B · 27Bfrom 2.7 GBruns from 8 GB RAM256K contextApache 2.0

Qwen 3.5 Uncensored

HauhauCS · Mar 2026

Qwen 3.5 with safety guardrails removed. No content filtering or refusals.

2B · 4B · 9B · 27Bfrom 1.2 GBruns from 4 GB RAM256K contextApache 2.0

MedGemma

Google · Jan 2026

Google's open medical models - discuss your own health records, lab results, and medical images (X-rays, skin photos, scans) privately on your device. For understanding and preparing questions, not diagnosis.

4B (v1.5) · 27Bfrom 2.4 GBruns from 8 GB RAM128K contextsees imagesMLXHealth AI Developer Foundations terms

Qwen 3.8 Distilled

empero-ai · Aug 2026

Empero's full-parameter distillations of the Qwen 3.8 flagship into small, fast sizes - a big knowledge jump over same-size models (the 9B scores near the giants on broad knowledge), with step-by-step reasoning. Text only.

2B · 4B · 9Bfrom 1.3 GBruns from 4 GB RAM256K contextApache 2.0

Qwen 3.8 Uncensored

orcarouter · Aug 2026

The Qwen 3.8 flagship with refusal behavior substantially reduced (openly documented - reduced, not eliminated). Same frontier coding, math, and reasoning, and it can see images with its vision add-on. For 24GB-class graphics cards.

27Bfrom 16.5 GBruns from 32 GB RAM256K contextsees imagesApache 2.0

Qwythos 9B

empero-ai · Jul 2026

A community Qwen 3.5-based merge — uncensored and multimodal (it can see images you attach), with step-by-step reasoning and tool use. A creative, unfiltered generalist. v2 trains out the repetition loops of the original.

9Bfrom 5.4 GBruns from 16 GB RAM1M contextsees imagesMLXApache 2.0

Hemmingway 1

Altworld · Sep 2026

Writes like a person. Built on Qwen 3.8 for the writing you do every day - messages, emails, the awkward note to a colleague - and gives you the text, not three options and a preamble. Free for personal use; business use needs an agreement with Altworld.

27Bfrom 17.4 GBruns from 32 GB RAM256K contextMLXCreative Commons BY-NC 4.0 (non-commercial)

Skyfall 31B

TheDrummer · Feb 2026

A creative-writing and roleplay model built on Mistral's Magistral Small. Vivid, in-character storytelling that keeps long scenes and many characters straight. No content filtering or refusals.

31Bfrom 19.0 GBruns from 32 GB RAM128K contextApache 2.0

Looking for frontier models instead? Browse the online catalog - one account, pay per use.