83.22tok/s
Independent open-model infrastructure
Dense-model inference,
measured honestly.
A dedicated Qwen3.8 27B endpoint engineered for practical throughput, predictable concurrency, and transparent operating limits.
4×60Ktokens
Actual long-input test · 4/4 complete
64GB
Dedicated Ampere memory
Current serving profile
One model. A clear envelope.
- Model
- Qwen3.8-27B-MTP
- Quantization
- Q8_0 / int8
- Interface
- OpenAI-compatible, streaming
- Context ceiling
- 65,536 tokens / request
- Prompt retention
- None by default
- Availability
- Pre-production validation