GPT OSS 20B running in a Trusted Execution Environment (TEE). OpenAI's compact open-weight 21B MoE model with 3.6B active parameters, optimized for lower-latency inference, with hardware attestation evidence available for independent verification.
Beta
Modalities
In / out price
$0.03 / $0.14
per 1M
Context
131K
Max output
16K
Released
Mar 18, 2026
Updated
Aug 7, 2026
Providers
3 routesThe same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below.
| Provider | Privacy | Input / 1M | Output / 1M | Cache read / 1M | Context | Max output | Uptime | Latency | Throughput |
|---|---|---|---|---|---|---|---|---|---|
| E2EE | $0.05 | $0.19 | — | 128K | 33K | — | — | — | |
| Private | $0.07 | $0.30 | — | 128K | 16K | — | — | — | |
| Private | $0.03 | $0.14 | — | 131K | 16K | — | — | — |
Privacy
3 / 3End-to-end encrypted routing
E2EE tier
- Anonymous
- Private
- E2EE
Your client encrypts the prompt before it leaves your environment. E2EE includes TEE isolation; only the verified enclave can decrypt the request.
Features
3Streaming
Reasoning
Web search