Kling

Inkling

kuaishou
Chat

Inkling is a general-purpose multimodal model from Thinking Machines Lab that accepts text, image, and audio inputs and generates text. It is a 66-layer sparse MoE (975B total / 41B active) with hybrid local/global attention, 1M context, and variable thinking effort — suited for chat, coding, tool use, and agentic workflows. Video input is not supported on Venice.

Beta

Modalities

In / out price

$2.3375 / $5.85

per 1M

Context

1M

Max output

66K

Released

Jul 16, 2026

Updated

Jul 17, 2026

Privacy

2 / 3

Private routing

Private tier

  1. Anonymous
  2. Private
  3. E2EE
  • Prompt and response content is not retained after the request.
  • Request content is visible to the inference runtime while it is being processed.

Moderation

Not specified

No model-level moderation classification is recorded in this catalog.

Data retention

Zero retention

The provider reports that request content is discarded after inference.

Features

7
Streaming
Tool calling
Reasoning
Vision
JSON/schema
Web search
Code optimized

Routing

2

Anonrouter hosted

Routed through AnonRouter's gateway with metadata-only logging.

Provider direct

Requests egress directly to the provider runtime.

USD per 1M tokens from the Venice models API.

Providers & sources

1 route

Uptime

30d