03 · HARDWARE PATHS

On-Premise Hardware Paths

Every hardware build is delivered pre-loaded with vLLM, SGLang, and Open WebUI, fully load-tested on your model of choice. Zero software assembly required on your end.

Desktop Private AI$6,500

DGX Spark

Private AI on a desk. RAG, sensitive documents, small-team inference.

Memory: 128GB Unified Memory (~119–121 GiB usable after display reservation)
Speed: 80 to 140 tok/s (est. aggregate across concurrent sessions)
Capacity: 8 to 32 sessions
Power: 240W peak
Frontier Open Models: Quantized derivatives only (official weights require Dual Spark)
Supported Workloads
  • Document Search & RAG
  • Legal & Contract Review
  • Small Team Inference
  • Local Coding Assistant
Ask About This Build
Multi-System DeploymentsFrom $12,000

Dual DGX Spark

Frontier open models on-premises, including DeepSeek-V4-Flash-0731, Qwen3.8-Flash-Next 4-bit, and 200B to 400B MoEs, with double the concurrency.

Memory: 256GB combined across two-node ConnectX link
Speed: 160 to 280 tok/s (est. aggregate across concurrent sessions)
Capacity: 16 to 64 sessions
Power: 480W peak
Frontier Open Models: Yes (Official weights & MoE, Qwen3.8-Flash-Next 4-bit)
Supported Workloads
  • DeepSeek-V4-Flash-0731
  • Qwen3.8-Flash-Next 4-bit
  • 200B-400B Frontier MoEs
  • Multi-department AI
Ask About This Build
Multi-System DeploymentsCustom

Quad DGX Spark

Maximum Spark-tier concurrency, 300B-class frontier MoEs at high precision including GLM-5.3-Flash FP8, larger context headroom.

Memory: 512GB combined across four nodes
Speed: 320 to 560 tok/s (est. aggregate across concurrent sessions)
Capacity: 32 to 128 sessions
Power: 960W peak
Frontier Open Models: Yes (GLM-5.3-Flash FP8, GLM 5.2 & DeepSeek V4)
Supported Workloads
  • GLM-5.3-Flash FP8, GLM 5.2 & DeepSeek V4
  • 300B+ MoE Models at FP8/BF16
  • 1M Context Ingestion
  • Enterprise Agent Pipeline
Ask About This Build
RTX PRO Workstations and ServersCustom

1× RTX PRO 6000

Custom workstation. More GPU bandwidth, creative workloads, component flexibility.

Memory: 96GB VRAM pool
Speed: 50 to 90 tok/s (est.)
Capacity: 4 to 16 sessions
Power: 300W peak
Frontier Open Models: Quantized derivatives only
Supported Workloads
  • Creative & Video AI
  • Custom Component Builds
  • 96GB Dedicated VRAM
  • Local Workstation Fine-Tuning
Ask About This Build
RTX PRO Workstations and ServersCustom

2× RTX PRO 6000

Dual-GPU server. Larger models, higher concurrency, heavier RAG, image/video work.

Memory: 192GB VRAM pool
Speed: 100 to 180 tok/s (est.)
Capacity: 8 to 32 sessions
Power: 600W peak
Frontier Open Models: Yes (DeepSeek-V4-Flash-0731)
Supported Workloads
  • DeepSeek-V4-Flash-0731
  • High-bandwidth Video & Image AI
  • Dual GPU Tensor Parallelism
  • 192GB VRAM Pool
Ask About This Build
RTX PRO Workstations and ServersCustom

4× RTX PRO 6000

Small local cluster. Runs 300B-class frontier MoEs at high precision with headroom for long context.

Memory: 384GB VRAM pool
Speed: 200 to 360 tok/s (est.)
Capacity: 16 to 64 sessions
Power: 1200W peak
Frontier Open Models: Yes (300B-class MoE at FP8/BF16)
Supported Workloads
  • 300B Frontier MoEs
  • Long Context Window Ingestion
  • Rack or Tower Deployment
  • 384GB VRAM Pool
Ask About This Build
Private Multi-Node ClustersScoped

Private Cluster

Production-scale inference for GLM-5.3 (744B), GLM 5.2, and DeepSeek V4 Pro, video generation, agent workflows.

Memory: Custom multi-node VRAM pool
Speed: 500+ tok/s (est. aggregate)
Capacity: 100+ sessions
Power: Dedicated PDU
Frontier Open Models: Yes (All Models, including GLM-5.3 744B)
Supported Workloads
  • Production Scale Inference
  • GLM-5.3 (744B) & GLM 5.2
  • Multi-Node Slurm / K8s
  • Enterprise Agent Orchestration
Ask About This Build

Some frontier open models exceed single-site deployment, for example Kimi K3 at 2.8T total parameters and roughly 1.56TB of weights. We size multi-node builds for these on request.

Throughput and concurrency figures are estimates and depend heavily on model, quantization, context length, and workload. Concurrency in particular is model dependent: architectures using recurrent state allocate memory per session slot rather than per token. We measure your actual numbers during the Private AI Assessment.
INTERACTIVE COMPARISON

Side-by-Side Hardware Comparison

Select any two hardware configurations to compare memory headroom, token speed, power draw, and user capacity.

SpecificationDGX SparkDual DGX Spark
Memory Pool128GB Unified Memory (~119–121 GiB usable after display reservation)256GB combined across two-node ConnectX link
Est. Throughput80 to 140 tok/s (est. aggregate across concurrent sessions)160 to 280 tok/s (est. aggregate across concurrent sessions)
User Concurrency8 to 32 sessions16 to 64 sessions
Peak Power Draw240W peak480W peak
Frontier Open ModelsQuantized derivatives only (official weights require Dual Spark)Yes (Official weights & MoE, Qwen3.8-Flash-Next 4-bit)
Price & Inclusions$6,500 (includes deployment, configuration, and load testing)From $12,000 (includes deployment, configuration, and load testing)
Primary FitPrivate AI on a desk. RAG, sensitive documents, small-team inference.Frontier open models on-premises, including DeepSeek-V4-Flash-0731, Qwen3.8-Flash-Next 4-bit, and 200B to 400B MoEs, with double the concurrency.

Not Sure Which Hardware Path Fits Your Team?

Send us your prompt lengths and estimated concurrent users. We will generate a workload readout and recommend the exact build.