On-Premise Hardware Paths
Every hardware build is delivered pre-loaded with vLLM, SGLang, and Open WebUI, fully load-tested on your model of choice. Zero software assembly required on your end.
DGX Spark
Private AI on a desk. RAG, sensitive documents, small-team inference.
- Document Search & RAG
- Legal & Contract Review
- Small Team Inference
- Local Coding Assistant
Dual DGX Spark
Frontier open models on-premises, including DeepSeek-V4-Flash-0731, Qwen3.8-Flash-Next 4-bit, and 200B to 400B MoEs, with double the concurrency.
- DeepSeek-V4-Flash-0731
- Qwen3.8-Flash-Next 4-bit
- 200B-400B Frontier MoEs
- Multi-department AI
Quad DGX Spark
Maximum Spark-tier concurrency, 300B-class frontier MoEs at high precision including GLM-5.3-Flash FP8, larger context headroom.
- GLM-5.3-Flash FP8, GLM 5.2 & DeepSeek V4
- 300B+ MoE Models at FP8/BF16
- 1M Context Ingestion
- Enterprise Agent Pipeline
1× RTX PRO 6000
Custom workstation. More GPU bandwidth, creative workloads, component flexibility.
- Creative & Video AI
- Custom Component Builds
- 96GB Dedicated VRAM
- Local Workstation Fine-Tuning
2× RTX PRO 6000
Dual-GPU server. Larger models, higher concurrency, heavier RAG, image/video work.
- DeepSeek-V4-Flash-0731
- High-bandwidth Video & Image AI
- Dual GPU Tensor Parallelism
- 192GB VRAM Pool
4× RTX PRO 6000
Small local cluster. Runs 300B-class frontier MoEs at high precision with headroom for long context.
- 300B Frontier MoEs
- Long Context Window Ingestion
- Rack or Tower Deployment
- 384GB VRAM Pool
Private Cluster
Production-scale inference for GLM-5.3 (744B), GLM 5.2, and DeepSeek V4 Pro, video generation, agent workflows.
- Production Scale Inference
- GLM-5.3 (744B) & GLM 5.2
- Multi-Node Slurm / K8s
- Enterprise Agent Orchestration
Some frontier open models exceed single-site deployment, for example Kimi K3 at 2.8T total parameters and roughly 1.56TB of weights. We size multi-node builds for these on request.
Side-by-Side Hardware Comparison
Select any two hardware configurations to compare memory headroom, token speed, power draw, and user capacity.
| Specification | DGX Spark | Dual DGX Spark |
|---|---|---|
| Memory Pool | 128GB Unified Memory (~119–121 GiB usable after display reservation) | 256GB combined across two-node ConnectX link |
| Est. Throughput | 80 to 140 tok/s (est. aggregate across concurrent sessions) | 160 to 280 tok/s (est. aggregate across concurrent sessions) |
| User Concurrency | 8 to 32 sessions | 16 to 64 sessions |
| Peak Power Draw | 240W peak | 480W peak |
| Frontier Open Models | Quantized derivatives only (official weights require Dual Spark) | Yes (Official weights & MoE, Qwen3.8-Flash-Next 4-bit) |
| Price & Inclusions | $6,500 (includes deployment, configuration, and load testing) | From $12,000 (includes deployment, configuration, and load testing) |
| Primary Fit | Private AI on a desk. RAG, sensitive documents, small-team inference. | Frontier open models on-premises, including DeepSeek-V4-Flash-0731, Qwen3.8-Flash-Next 4-bit, and 200B to 400B MoEs, with double the concurrency. |
Not Sure Which Hardware Path Fits Your Team?
Send us your prompt lengths and estimated concurrent users. We will generate a workload readout and recommend the exact build.