INFERENCEKF Inference
Dedicated inference / CN region

Long-context inference, kept deliberately small.

KF Inference operates a focused language-model endpoint on two NVIDIA DGX Spark systems. One model, conservative capacity, transparent limits.

ModelDeepSeek V4 Flash
Context262K tokens
InterfaceOpenAI compatible
AvailabilityCanary

Operating stance

Capacity is intentionally capped below laboratory limits. When the service is full, it returns a prompt 429 response instead of hiding latency in a long queue.

Streaming, usage reporting, JSON output, reasoning content and tool calls are validated on the current deployment.

Infrastructure

Inference runs on a two-node NVIDIA GB10 deployment with approximately 256 GB of combined unified memory. The public edge reaches the model over a dedicated private WireGuard link.

No shell, container, notebook or file-system access is offered to API consumers.