Phala AI

inference provider · 8 models

Access 8 models served through Phala AI on AnonRouter's privacy-first gateway, including DeepSeek V4 Flash 0731, Gemma 4 26B A4B Uncensored, and GPT OSS 20B. Phala AI routes through an aggregator gateway that decrypts each request before forwarding it to the model, so Phala can see prompt content; AnonRouter strips the caller's identity before the request reaches it, and these routes carry no enclave attestation.

Models

8

Modalities

1

Text

From (input)

$0.04

per 1M tokens

Max context

1M

Private routes

8 / 8

not anonymous-only

Catalog by modality

8 routes
Text8

Phala AI models8

DeepSeek
DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
Phala AIPrivate

DeepSeek V4 Flash is an efficiency-optimized 284B-parameter Mixture-of-Experts model with 13B active parameters and a 1M-token context window. Tuned for fast inference and high-throughput workloads while maintaining strong reasoning and coding performance.

Private|1M context|$0.44/M input|$1.32/M output
Gemma
Gemma 4 26B A4B Uncensored
google/gemma-4-26b-a4b-uncensored
Phala AIPrivate

Gemma 4 26B A4B Uncensored running in a Trusted Execution Environment (TEE). An uncensored variant of Google's Gemma 4 MoE model with 25.2B total / 3.8B active parameters, supporting multimodal input across text and images, with hardware attestation evidence available for independent verification.

Private|66K context|$0.15/M input|$0.70/M output
OpenAI
GPT OSS 20B
openai/gpt-oss-20b
Phala AIPrivate

GPT OSS 20B running in a Trusted Execution Environment (TEE). OpenAI's compact open-weight 21B MoE model with 3.6B active parameters, optimized for lower-latency inference, with hardware attestation evidence available for independent verification.

Private|131K context|$0.04/M input|$0.15/M output
Qwen
Qwen 2.5 7B
qwen/qwen-2.5-7b
Phala AIPrivate

Qwen 2.5 7B Instruct running in a Trusted Execution Environment (TEE). A compact model with strong coding, math, and multilingual capabilities supporting 29+ languages, with hardware attestation evidence available for independent verification.

Private|33K context|$0.10/M input|$0.20/M output
Qwen
Qwen3.6 35B A3B Uncensored
qwen/qwen3.6-35b-a3b-uncensored
Phala AIPrivate

Qwen3.6 35B A3B Uncensored running in a Trusted Execution Environment (TEE). An uncensored variant of Alibaba's Qwen3.6 MoE model with 35B total parameters and ~3B active, supporting 262K context and multimodal input across text, images, and video, with hardware attestation evidence available for independent verification.

Private|131K context|$0.30/M input|$1.50/M output
Zhipu
GLM 5.3
z-ai/glm-5.3
Phala AIPrivate

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Private|1M context|$1.40/M input|$4.40/M output
Zhipu
GLM 5.3 Flash
z-ai/glm-5.3-flash
Phala AIPrivate

GLM-5.3 Flash is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.

Private|1M context|$0.15/M input|$0.50/M output
DeepInfra
Muse Glimmer 30B
meta-models/muse-glimmer-30b
Phala AIPrivate

Muse Glimmer is a 30B multimodal agentic model distilled from Muse Spark — reasoning, tool use, and failure recovery in a single model that runs locally on consumer hardware.

Private|131K context|$0.30/M input|$1.10/M output
Provider documentation2 sources
Phala AI provider — AnonRouter