Venice

Inkling

Chat

Inkling is a general-purpose multimodal model from Thinking Machines Lab that accepts text, image, and audio inputs and generates text. It is a 66-layer sparse MoE (975B total / 41B active) with hybrid local/global attention, 1M context, and variable thinking effort — suited for chat, coding, tool use, and agentic workflows. Video input is not supported on Venice.

Beta

Modalities

In / out price

$2.3375 / $5.85

per 1M

Context

1M

Max output

66K

Released

Jul 16, 2026

Updated

Jul 17, 2026

Providers

1 route

The same model can have different pricing and privacy guarantees depending on who serves it. The model's headline privacy uses the strongest available route below. Discovered routes awaiting approval remain listed as not live.

ProviderPrivacyInput / 1MOutput / 1MCache read / 1MContextMax outputUptimeLatencyThroughput
VeniceVenice
Private$2.3375$5.851M66K

Privacy

2 / 4

Private routing

Private tier

  1. Anonymous
  2. Private
  3. TEE
  4. E2EE

Prompt and response content is not retained after the request. Request content is visible to the inference runtime while it is being processed. Learn more

Features

7
Streaming
Tool calling
Reasoning
Vision
JSON/schema
Web search
Code optimized

Uptime