Model
Fit
Speed
Memory
Scores / Run
Runs comfortably
9.8GB
262k ctx
Runs comfortably
16
tok/s
ⓘ estimated
10
/ 16GB
88
/
89
/
85
Runs comfortably
16 tok/s
ⓘ estimated
10/16GB
6.2GB
262k ctx
Runs comfortably
26
tok/s
ⓘ estimated
6
/ 16GB
88
/
89
/
85
Runs comfortably
26 tok/s
ⓘ estimated
6/16GB
Runs comfortably
16 tok/s
ⓘ estimated
10/16GB
Runs comfortably
21 tok/s
ⓘ estimated
8/16GB
Runs comfortably
24 tok/s
ⓘ estimated
7/16GB
Runs comfortably
27 tok/s
ⓘ estimated
6/16GB
Runs comfortably
19 tok/s
ⓘ estimated
9/16GB
Runs comfortably
18 tok/s
ⓘ estimated
9/16GB
Runs comfortably
12 tok/s
ⓘ estimated
13/16GB
Runs comfortably
22 tok/s
ⓘ estimated
7/16GB
Runs comfortably
30 tok/s
ⓘ estimated
5/16GB
Runs comfortably
19 tok/s
ⓘ estimated
9/16GB
Runs comfortably
12 tok/s
ⓘ estimated
14/16GB
Runs comfortably
21 tok/s
ⓘ estimated
8/16GB
Runs comfortably
22 tok/s
ⓘ estimated
7/16GB
Runs comfortably
12 tok/s
ⓘ estimated
13/16GB
Runs comfortably
61 tok/s
ⓘ estimated
3/16GB
Runs comfortably
39 tok/s
ⓘ estimated
4/16GB
Runs comfortably
63 tok/s
ⓘ estimated
3/16GB
8.0GB
128k ctx
Runs comfortably
20
tok/s
ⓘ estimated
8
/ 16GB
52
/
55
/
62
Runs comfortably
20 tok/s
ⓘ estimated
8/16GB
Llama 3.1 8B Instruct
Q4_K_M
4.5GB
128k ctx
Runs comfortably
35
tok/s
ⓘ estimated
5
/ 16GB
52
/
55
/
62
Runs comfortably
35 tok/s
ⓘ estimated
5/16GB
Runs comfortably
32 tok/s
ⓘ estimated
5/16GB
Runs comfortably
53 tok/s
ⓘ estimated
3/16GB
3.5GB
128k ctx
Runs comfortably
45
tok/s
ⓘ estimated
4
/ 16GB
35
/
38
/
50
Runs comfortably
45 tok/s
ⓘ estimated
4/16GB
Llama 3.2 3B Instruct
Q4_K_M
2.0GB
128k ctx
Runs comfortably
79
tok/s
ⓘ estimated
2
/ 16GB
35
/
38
/
50
Runs comfortably
79 tok/s
ⓘ estimated
2/16GB
Runs tight
Runs tight
81 tok/s
ⓘ estimated
13/16GB
Runs tight
104 tok/s
ⓘ estimated
16/16GB
Runs tight
11 tok/s
ⓘ estimated
15/16GB
Runs tight
11 tok/s
ⓘ estimated
15/16GB
Runs tight
10 tok/s
ⓘ estimated
16/16GB
Runs with CPU offload
Qwen3.8-Flash-Next
EXL3_2BPW
62.7GB
0k cap
Runs with CPU offload
76
tok/s
ⓘ estimated
63
/ 16GB
92
/
92
/
89
Runs with CPU offload
76 tok/s
ⓘ estimated
63/16GB
Qwen3.8-Flash-Next
EXL3_3BPW
85.0GB
0k cap
Runs with CPU offload
56
tok/s
ⓘ estimated
85
/ 16GB
92
/
92
/
89
Runs with CPU offload
56 tok/s
ⓘ estimated
85/16GB
Ornith-1.5-35B-A3B
Q5_K_M
25.3GB
0k cap
Runs with CPU offload
75
tok/s
ⓘ estimated
25
/ 16GB
91
/
89
/
86
Runs with CPU offload
75 tok/s
ⓘ estimated
25/16GB
Ornith-1.5-35B-A3B
Q4_K_M
21.7GB
0k cap
Runs with CPU offload
87
tok/s
ⓘ estimated
22
/ 16GB
91
/
89
/
86
Runs with CPU offload
87 tok/s
ⓘ estimated
22/16GB
Qwen3.6 35B A3B
Q4_K_M
20.0GB
0k cap
Runs with CPU offload
92
tok/s
ⓘ estimated
20
/ 16GB
90
/
91
/
87
Runs with CPU offload
92 tok/s
ⓘ estimated
20/16GB
Pokee-Isaac 28B
Q4_K_M
16.0GB
0k cap
Runs with CPU offload
10
tok/s
ⓘ estimated
16
/ 16GB
90
/
88
/
88
Runs with CPU offload
10 tok/s
ⓘ estimated
16/16GB
17.6GB
0k cap
Runs with CPU offload
9
tok/s
ⓘ estimated
18
/ 16GB
88
/
89
/
85
Runs with CPU offload
9 tok/s
ⓘ estimated
18/16GB
Runs with CPU offload
9 tok/s
ⓘ estimated
18/16GB
Runs with CPU offload
9 tok/s
ⓘ estimated
18/16GB
Runs with CPU offload
10 tok/s
ⓘ estimated
17/16GB
Runs with CPU offload
6 tok/s
ⓘ estimated
25/16GB
Runs with CPU offload
8 tok/s
ⓘ estimated
20/16GB
Runs with CPU offload
9 tok/s
ⓘ estimated
18/16GB
Xing4.0-29B-A4B
Q4_K_M
20.1GB
256k ctx
Runs with CPU offload
57
tok/s
ⓘ estimated
20
/ 16GB
75
/
73
/
72
Runs with CPU offload
57 tok/s
ⓘ estimated
20/16GB
21.5GB
0k cap
Runs with CPU offload
7
tok/s
ⓘ estimated
22
/ 16GB
72
/
77
/
80
Runs with CPU offload
7 tok/s
ⓘ estimated
22/16GB
Runs with CPU offload
9 tok/s
ⓘ estimated
17/16GB
Mixtral 8x7B Instruct
Q4_K_M
26.5GB
0k cap
Runs with CPU offload
22
tok/s
ⓘ estimated
27
/ 16GB
62
/
65
/
68
Runs with CPU offload
22 tok/s
ⓘ estimated
27/16GB
Mixtral 8x7B Instruct
Q3_K_M
19.0GB
0k cap
Runs with CPU offload
30
tok/s
ⓘ estimated
19
/ 16GB
62
/
65
/
68
Runs with CPU offload
30 tok/s
ⓘ estimated
19/16GB
Runs with CPU offload
9 tok/s
ⓘ estimated
18/16GB
Coding alternatives
Hosted coding assistants for comparison - per seat
Cursor
Hobby
AI-first code editor. Uses frontier models (GPT-4, Claude). $20/user Business...
$0.0
GitHub Copilot
Individual
AI coding assistant by GitHub. GPT-4o and Claude-based completions. $19/seat ...
$10.0
GitHub Copilot
Business
AI coding assistant by GitHub. GPT-4o and Claude-based completions. $19/seat ...
$19.0
Cursor
Pro
AI-first code editor. Uses frontier models (GPT-4, Claude). $20/user Business...
$20.0
GitHub Copilot
Enterprise
AI coding assistant by GitHub. GPT-4o and Claude-based completions. $19/seat ...
$39.0
Cursor
Business
AI-first code editor. Uses frontier models (GPT-4, Claude). $20/user Business...
$40.0
Per-token monthly figures are estimates at typical usage. Prices verified periodically - always confirm current pricing on the provider's website.
Sign in with GitHub to save this rig and get weekly model updates
Sign in with GitHub