8 nodes / 1024GB aggregate 12.5 GB/s link
96 models / sorted by best_quality
Model
Fit
Speed
Memory
Scores / Run
Runs comfortably
93.1GB
1000k ctx
Runs comfortably
29
tok/s
ⓘ estimated
93
/ 1024GB
93
/
92
/
91
Runs comfortably
29 tok/s
ⓘ estimated
93/1024GB
200.0GB
1000k ctx
Runs comfortably
0.22
tok/s
distributed - link bound
200
/ 1024GB
93
/
92
/
91
Runs comfortably
0.22 tok/s
distributed - link bound
200/1024GB
120.0GB
1000k ctx
Runs comfortably
22
tok/s
ⓘ estimated
120
/ 1024GB
93
/
92
/
91
Runs comfortably
22 tok/s
ⓘ estimated
120/1024GB
109.0GB
1000k ctx
Runs comfortably
24
tok/s
ⓘ estimated
109
/ 1024GB
93
/
92
/
91
Runs comfortably
24 tok/s
ⓘ estimated
109/1024GB
510.3GB
1000k ctx
Runs comfortably
0.16
tok/s
distributed - link bound
510
/ 1024GB
93
/
92
/
90
Runs comfortably
0.16 tok/s
distributed - link bound
510/1024GB
DeepSeek V4.1 Flash
NVFP4
260.0GB
1000k ctx
Runs comfortably
0.32
tok/s
distributed - link bound
260
/ 1024GB
93
/
92
/
90
Runs comfortably
0.32 tok/s
distributed - link bound
260/1024GB
162.0GB
1000k ctx
Runs comfortably
0.33
tok/s
distributed - link bound
162
/ 1024GB
93
/
91
/
90
Runs comfortably
0.33 tok/s
distributed - link bound
162/1024GB
155.0GB
1000k ctx
Runs comfortably
0.34
tok/s
distributed - link bound
155
/ 1024GB
93
/
91
/
90
Runs comfortably
0.34 tok/s
distributed - link bound
155/1024GB
Ornith-1.5-397B
Q4_K_M
244.0GB
262k ctx
Runs comfortably
0.01
tok/s
distributed - link bound
244
/ 1024GB
93
/
91
/
89
Runs comfortably
0.01 tok/s
distributed - link bound
244/1024GB
Ornith-1.5-397B
Q5_K_M
286.0GB
262k ctx
Runs comfortably
0.01
tok/s
distributed - link bound
286
/ 1024GB
93
/
91
/
89
Runs comfortably
0.01 tok/s
distributed - link bound
286/1024GB
Ornith-1.5-397B
Q6_K
331.0GB
262k ctx
Runs comfortably
0.01
tok/s
distributed - link bound
331
/ 1024GB
93
/
91
/
89
Runs comfortably
0.01 tok/s
distributed - link bound
331/1024GB
Ornith-1.5-397B
Q8_0
429.0GB
262k ctx
Runs comfortably
0.01
tok/s
distributed - link bound
429
/ 1024GB
93
/
91
/
89
Runs comfortably
0.01 tok/s
distributed - link bound
429/1024GB
72.5GB
262k ctx
Runs comfortably
43
tok/s
ⓘ estimated
73
/ 1024GB
92
/
92
/
89
Runs comfortably
43 tok/s
ⓘ estimated
73/1024GB
510.0GB
1000k ctx
Runs comfortably
0.09
tok/s
distributed - link bound
510
/ 1024GB
91
/
92
/
88
Runs comfortably
0.09 tok/s
distributed - link bound
510/1024GB
410.0GB
1000k ctx
Runs comfortably
0.11
tok/s
distributed - link bound
410
/ 1024GB
91
/
92
/
88
Runs comfortably
0.11 tok/s
distributed - link bound
410/1024GB
365.0GB
1000k ctx
Runs comfortably
0.12
tok/s
distributed - link bound
365
/ 1024GB
91
/
92
/
88
Runs comfortably
0.12 tok/s
distributed - link bound
365/1024GB
240.0GB
1000k ctx
Runs comfortably
0.19
tok/s
distributed - link bound
240
/ 1024GB
91
/
92
/
88
Runs comfortably
0.19 tok/s
distributed - link bound
240/1024GB
DeepSeek V3 0324
Q4_K_M
368.0GB
128k ctx
Runs comfortably
0.12
tok/s
distributed - link bound
368
/ 1024GB
92
/
88
/
87
Runs comfortably
0.12 tok/s
distributed - link bound
368/1024GB
DeepSeek V3 0324
Q2_K
188.0GB
128k ctx
Runs comfortably
0.23
tok/s
distributed - link bound
188
/ 1024GB
92
/
88
/
87
Runs comfortably
0.23 tok/s
distributed - link bound
188/1024GB
Ornith-1.5-35B-A3B
BF16
71.1GB
262k ctx
Runs comfortably
25
tok/s
ⓘ estimated
71
/ 1024GB
91
/
89
/
86
Runs comfortably
25 tok/s
ⓘ estimated
71/1024GB
Ornith-1.5-35B-A3B
Q8_0
37.8GB
262k ctx
Runs comfortably
48
tok/s
ⓘ estimated
38
/ 1024GB
91
/
89
/
86
Runs comfortably
48 tok/s
ⓘ estimated
38/1024GB
Ornith-1.5-35B-A3B
Q6_K
29.2GB
262k ctx
Runs comfortably
62
tok/s
ⓘ estimated
29
/ 1024GB
91
/
89
/
86
Runs comfortably
62 tok/s
ⓘ estimated
29/1024GB
Ornith-1.5-35B-A3B
Q5_K_M
25.3GB
262k ctx
Runs comfortably
71
tok/s
ⓘ estimated
25
/ 1024GB
91
/
89
/
86
Runs comfortably
71 tok/s
ⓘ estimated
25/1024GB
Ornith-1.5-35B-A3B
Q4_K_M
21.7GB
262k ctx
Runs comfortably
83
tok/s
ⓘ estimated
22
/ 1024GB
91
/
89
/
86
Runs comfortably
83 tok/s
ⓘ estimated
22/1024GB
Qwen3.6 35B A3B
Q4_K_M
20.0GB
262k ctx
Runs comfortably
87
tok/s
ⓘ estimated
20
/ 1024GB
90
/
91
/
87
Runs comfortably
87 tok/s
ⓘ estimated
20/1024GB
Runs comfortably
47 tok/s
ⓘ estimated
37/1024GB
Pokee-Isaac 28B
Q4_K_M
16.0GB
10000k ctx
Runs comfortably
9
tok/s
ⓘ estimated
16
/ 1024GB
90
/
88
/
88
Runs comfortably
9 tok/s
ⓘ estimated
16/1024GB
Pokee-Isaac 28B
Q8_0
28.0GB
10000k ctx
Runs comfortably
5
tok/s
ⓘ estimated
28
/ 1024GB
90
/
88
/
88
Runs comfortably
5 tok/s
ⓘ estimated
28/1024GB
Pokee-Isaac 28B
FP16
56.0GB
10000k ctx
Runs comfortably
3
tok/s
ⓘ estimated
56
/ 1024GB
90
/
88
/
88
Runs comfortably
3 tok/s
ⓘ estimated
56/1024GB
Qwen3 235B A22B
Q4_K_M
125.0GB
128k ctx
Runs comfortably
13
tok/s
ⓘ estimated
125
/ 1024GB
88
/
90
/
85
Runs comfortably
13 tok/s
ⓘ estimated
125/1024GB
Runs comfortably
24 tok/s
ⓘ estimated
68/1024GB
17.6GB
262k ctx
Runs comfortably
9
tok/s
ⓘ estimated
18
/ 1024GB
88
/
89
/
85
Runs comfortably
9 tok/s
ⓘ estimated
18/1024GB
9.8GB
262k ctx
Runs comfortably
15
tok/s
ⓘ estimated
10
/ 1024GB
88
/
89
/
85
Runs comfortably
15 tok/s
ⓘ estimated
10/1024GB
6.2GB
262k ctx
Runs comfortably
24
tok/s
ⓘ estimated
6
/ 1024GB
88
/
89
/
85
Runs comfortably
24 tok/s
ⓘ estimated
6/1024GB
15.2GB
262k ctx
Runs comfortably
10
tok/s
ⓘ estimated
15
/ 1024GB
88
/
89
/
85
Runs comfortably
10 tok/s
ⓘ estimated
15/1024GB
Runs comfortably
9 tok/s
ⓘ estimated
16/1024GB
Runs comfortably
2 tok/s
ⓘ estimated
61/1024GB
Runs comfortably
5 tok/s
ⓘ estimated
33/1024GB
Runs comfortably
8 tok/s
ⓘ estimated
18/1024GB
Runs comfortably
8 tok/s
ⓘ estimated
18/1024GB
Runs comfortably
20 tok/s
ⓘ estimated
50/1024GB
Runs comfortably
40 tok/s
ⓘ estimated
25/1024GB
Gemma 4 26B A4B
Q4_K_M
13.0GB
256k ctx
Runs comfortably
76
tok/s
ⓘ estimated
13
/ 1024GB
85
/
88
/
84
Runs comfortably
76 tok/s
ⓘ estimated
13/1024GB
Gemma 4 26B A4B
NVFP4
14.0GB
256k ctx
Runs comfortably
71
tok/s
ⓘ estimated
14
/ 1024GB
85
/
88
/
84
Runs comfortably
71 tok/s
ⓘ estimated
14/1024GB
Runs comfortably
5 tok/s
ⓘ estimated
28/1024GB
Runs comfortably
9 tok/s
ⓘ estimated
17/1024GB
Runs comfortably
4 tok/s
ⓘ estimated
34/1024GB
Runs comfortably
5 tok/s
ⓘ estimated
29/1024GB
Runs comfortably
6 tok/s
ⓘ estimated
25/1024GB
Runs comfortably
8 tok/s
ⓘ estimated
20/1024GB
Nemotron 3.5 Lightning
Q4_K_M
32.0GB
1000k ctx
Runs comfortably
41
tok/s
ⓘ estimated
32
/ 1024GB
78
/
76
/
74
Runs comfortably
41 tok/s
ⓘ estimated
32/1024GB
Runs comfortably
8 tok/s
ⓘ estimated
18/1024GB
58.0GB
1000k ctx
Runs comfortably
23
tok/s
ⓘ estimated
58
/ 1024GB
78
/
76
/
74
Runs comfortably
23 tok/s
ⓘ estimated
58/1024GB
60.0GB
1000k ctx
Runs comfortably
22
tok/s
ⓘ estimated
60
/ 1024GB
78
/
76
/
74
Runs comfortably
22 tok/s
ⓘ estimated
60/1024GB
Runs comfortably
26 tok/s
ⓘ estimated
6/1024GB
Runs comfortably
15 tok/s
ⓘ estimated
10/1024GB
Runs comfortably
23 tok/s
ⓘ estimated
7/1024GB
Runs comfortably
20 tok/s
ⓘ estimated
8/1024GB
Runs comfortably
49 tok/s
ⓘ estimated
31/1024GB
Runs comfortably
98 tok/s
ⓘ estimated
16/1024GB
Runs comfortably
10 tok/s
ⓘ estimated
15/1024GB
Runs comfortably
5 tok/s
ⓘ estimated
29/1024GB
Runs comfortably
18 tok/s
ⓘ estimated
9/1024GB
Xing4.0-29B-A4B
Q4_K_M
20.1GB
256k ctx
Runs comfortably
54
tok/s
ⓘ estimated
20
/ 1024GB
75
/
73
/
72
Runs comfortably
54 tok/s
ⓘ estimated
20/1024GB
Runs comfortably
17 tok/s
ⓘ estimated
62/1024GB
Llama 3.3 70B Instruct
Q4_K_M
42.0GB
128k ctx
Runs comfortably
4
tok/s
ⓘ estimated
42
/ 1024GB
72
/
77
/
80
Runs comfortably
4 tok/s
ⓘ estimated
42/1024GB
74.0GB
128k ctx
Runs comfortably
2
tok/s
ⓘ estimated
74
/ 1024GB
72
/
77
/
80
Runs comfortably
2 tok/s
ⓘ estimated
74/1024GB
21.5GB
128k ctx
Runs comfortably
7
tok/s
ⓘ estimated
22
/ 1024GB
72
/
77
/
80
Runs comfortably
7 tok/s
ⓘ estimated
22/1024GB
Llama 3.3 70B Instruct
Q3_K_M
30.0GB
128k ctx
Runs comfortably
5
tok/s
ⓘ estimated
30
/ 1024GB
72
/
77
/
80
Runs comfortably
5 tok/s
ⓘ estimated
30/1024GB
Runs comfortably
10 tok/s
ⓘ estimated
15/1024GB
Runs comfortably
17 tok/s
ⓘ estimated
9/1024GB
Runs comfortably
12 tok/s
ⓘ estimated
13/1024GB
Runs comfortably
21 tok/s
ⓘ estimated
7/1024GB
Runs comfortably
5 tok/s
ⓘ estimated
28/1024GB
Runs comfortably
9 tok/s
ⓘ estimated
17/1024GB
Runs comfortably
18 tok/s
ⓘ estimated
9/1024GB
Runs comfortably
29 tok/s
ⓘ estimated
5/1024GB
Runs comfortably
11 tok/s
ⓘ estimated
14/1024GB
Runs comfortably
20 tok/s
ⓘ estimated
8/1024GB
Mixtral 8x7B Instruct
Q4_K_M
26.5GB
32k ctx
Runs comfortably
21
tok/s
ⓘ estimated
27
/ 1024GB
62
/
65
/
68
Runs comfortably
21 tok/s
ⓘ estimated
27/1024GB
Mixtral 8x7B Instruct
Q3_K_M
19.0GB
32k ctx
Runs comfortably
29
tok/s
ⓘ estimated
19
/ 1024GB
62
/
65
/
68
Runs comfortably
29 tok/s
ⓘ estimated
19/1024GB
Mistral Nemo 12B
Q8_0
13.2GB
128k ctx
Runs comfortably
11
tok/s
ⓘ estimated
13
/ 1024GB
58
/
60
/
62
Runs comfortably
11 tok/s
ⓘ estimated
13/1024GB
Mistral Nemo 12B
Q4_K_M
7.3GB
128k ctx
Runs comfortably
21
tok/s
ⓘ estimated
7
/ 1024GB
58
/
60
/
62
Runs comfortably
21 tok/s
ⓘ estimated
7/1024GB
Runs comfortably
58 tok/s
ⓘ estimated
3/1024GB
Runs comfortably
60 tok/s
ⓘ estimated
3/1024GB
Runs comfortably
37 tok/s
ⓘ estimated
4/1024GB
8.0GB
128k ctx
Runs comfortably
19
tok/s
ⓘ estimated
8
/ 1024GB
52
/
55
/
62
Runs comfortably
19 tok/s
ⓘ estimated
8/1024GB
16.0GB
128k ctx
Runs comfortably
9
tok/s
ⓘ estimated
16
/ 1024GB
52
/
55
/
62
Runs comfortably
9 tok/s
ⓘ estimated
16/1024GB
Llama 3.1 8B Instruct
Q4_K_M
4.5GB
128k ctx
Runs comfortably
33
tok/s
ⓘ estimated
5
/ 1024GB
52
/
55
/
62
Runs comfortably
33 tok/s
ⓘ estimated
5/1024GB
Runs comfortably
30 tok/s
ⓘ estimated
5/1024GB
Runs comfortably
50 tok/s
ⓘ estimated
3/1024GB
3.5GB
128k ctx
Runs comfortably
43
tok/s
ⓘ estimated
4
/ 1024GB
35
/
38
/
50
Runs comfortably
43 tok/s
ⓘ estimated
4/1024GB
Llama 3.2 3B Instruct
Q4_K_M
2.0GB
128k ctx
Runs comfortably
75
tok/s
ⓘ estimated
2
/ 1024GB
35
/
38
/
50
Runs comfortably
75 tok/s
ⓘ estimated
2/1024GB
Runs comfortably
8 tok/s
ⓘ estimated
18/1024GB
Runs comfortably
5 tok/s
ⓘ estimated
33/1024GB
Runs comfortably
3 tok/s
ⓘ estimated
60/1024GB
Run GLM-5.3-Flash in the cloud
No hardware? Host it - per-token API or flat subscription
InferenceNet
per-token
$0.11
DeepInfra
per-token
$0.18
Relace
per-token
$0.19
GMICloud
per-token
$0.22
OpenInference
per-token
$0.24
Sail Research
per-token
$0.27
Novita
per-token
$0.27
Inceptron
per-token
$0.29
Decart
per-token
$0.31
Phala
per-token
$0.31
StreamLake
per-token
$0.34
CoreWeave
per-token
$0.36
AtlasCloud
per-token
$0.36
Fireworks
per-token
$0.36
Friendli
per-token
$0.36
Venice
per-token
$0.36
Io Net
per-token
$0.36
Z.AI
per-token
$0.36
Modal
per-token
Serverless GPU inference platform. Serves open-weight models including reason...
$0.36
Z.ai
per-token
Official API and platform for Zhipu AI's open-source GLM models. GLM-5.2 avai...
pricing may be outdated
$0.36
SiliconFlow
per-token
$0.36
DigitalOcean
per-token
$0.36
Together
per-token
$0.36
Reka
per-token
$0.36
Parasail
per-token
$0.36
BaseTen
per-token
$0.36
Near AI
per-token
$0.36
Crusoe
per-token
$0.36
NextBit
per-token
$0.4
Morph
per-token
$0.4
Wafer
per-token
$0.7
Cloudflare
per-token
$0.72
Z.ai
Coding Plan Lite
subscription
Official API and platform for Zhipu AI's open-source GLM models. GLM-5.2 avai...
$10.0
OpenCode Go
Go ($5 first month)
subscription
Low-cost subscription for open coding models. $5 first month, then $10/month....
$10.0
Z.ai
Coding Plan Pro
subscription
Official API and platform for Zhipu AI's open-source GLM models. GLM-5.2 avai...
$30.0
Z.ai
Coding Plan Max
subscription
Official API and platform for Zhipu AI's open-source GLM models. GLM-5.2 avai...
$80.0
Coding alternatives
Hosted coding assistants for comparison - per seat
Cursor
Hobby
AI-first code editor. Uses frontier models (GPT-4, Claude). $20/user Business...
$0.0
GitHub Copilot
Individual
AI coding assistant by GitHub. GPT-4o and Claude-based completions. $19/seat ...
$10.0
GitHub Copilot
Business
AI coding assistant by GitHub. GPT-4o and Claude-based completions. $19/seat ...
$19.0
Cursor
Pro
AI-first code editor. Uses frontier models (GPT-4, Claude). $20/user Business...
$20.0
GitHub Copilot
Enterprise
AI coding assistant by GitHub. GPT-4o and Claude-based completions. $19/seat ...
$39.0
Cursor
Business
AI-first code editor. Uses frontier models (GPT-4, Claude). $20/user Business...
$40.0
Per-token monthly figures are estimates at typical usage. Prices verified periodically - always confirm current pricing on the provider's website.
Sign in with GitHub to save this rig and get weekly model updates
Sign in with GitHub