Service state
Pre-production
Public provider integration and external uptime monitoring are being prepared. No production availability is claimed yet.
Validated capacity
Measured September 4, 2026 on a dedicated 64 GB GA100-class accelerator using Qwen3.8-27B-MTP-Q8_0.
Single stream64K context56.05 tok/s mean3 / 3 stable
Four concurrentShort prompt83.22 tok/s aggregate4 / 4 stable
Four concurrent~60K input each239,987 input tokens4 / 4 stable
Eight concurrent16K each99.76 tok/s aggregateStable
Launch profile
Initial production profile: four concurrent requests, a 65,536-token per-request context ceiling, streaming text output, and request-content logging disabled. Four simultaneous ~60K-token inputs completed in 357.38 seconds at the slowest request.
Planned monitoring
External health probes, latency and error-rate alerts, GPU memory and thermal monitoring, controlled restart, and capacity-aware admission.