engy MiniMax H3 is now live!

Pricing

Per-token, pay as you go. No subscriptions and no minimums; you pay only for the tokens you use.

model input↑ output cached vs official ctx
MiniMax-H3new
$0.03/second — 32K
deepseek-v4.1-flashnewZDR
$0.30$0.04 $1.20$0.08 $0.006$0.008 −90% 328K
deepseek-v4-flash-0731ZDR
$0.045 $0.09 $0.009 — 1M
qwen3.6-35b-a3b
$0.375$0.045 $2.25$0.30 $0.015 −87% 208K
qwen3.8-27bZDR
$0.50$0.045 $3.00$0.32 $0.015 −90% 1M
glm-5.3-flash(Ox Alpha)newZDR
$0.15$0.135 $0.50$0.45 $0.03$0.027 −10% 262K
glm-5.2ZDR
$1.40$0.68 $4.40$1.50 $0.26$0.18 −59% 262K
glm-5.3newZDR
$1.40$0.98 $4.40$3.08 $0.26$0.18 −30% 328K
kimi-k3ZDR
$3.00$1.95 $15.00$9.75 $0.30$0.195 −35% 1M
$ per 1M tokens (/second = $ per second of output) · $1.40 = the model developer's own API price, checked 2026-09-29 · vs official uses a 3:1 input:output blend

Prompt-cache hits bill at the cached rate automatically, with no config and no cache_control markers. Agentic workloads (coding assistants, multi-turn tools) typically hit 90%+ cache on repeated prefixes, so effective input cost is usually far below the headline rate.

ZDR marks models under zero data retention: served only on hardware engy operates, prompts and outputs never stored or trained on. MiniMax-H3 (not covered by engy's zero-data-retention configuration) and qwen3.6-35b-a3b (routed across permissionless subnet miners, on hardware without confidential computing) carry no such mark; see terms and privacy.

These prices are a recorded Engy snapshot from 10 October 2026; this replica is not connected to the billing engine. The API reports the same numbers at https://api.engy.ai/v1/models.