Getting started
Point any OpenAI-compatible client at https://api.engy.ai/v1 with your API key. Everything below is the same endpoint from a different client.
Claude Code
In ~/.claude/settings.json, then run claude:
{
"env": {
"ANTHROPIC_BASE_URL": "https://api.engy.ai",
"ANTHROPIC_AUTH_TOKEN": "$ENGY_API_KEY",
"ANTHROPIC_MODEL": "kimi-k3[1m]",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "kimi-k3",
"CLAUDE_CODE_AUTO_MODE_SERVER": "0"
}
}
No /v1 on the URL; keep the haiku model set.
The [1m]suffix is how Claude Code learns the context window. It assumes 200K tokens for any model it does not recognise and refuses bigger prompts with “Prompt is too long” before sending them; kimi-k3 serves 1M, so the suffix unlocks the rest. It is stripped on the wire, so the request still names kimi-k3. Any model from the models list works the same way; add [1m] only where the list shows a 1M window.
The last line is for auto mode. Claude Code 2.1.278 and later asks the Anthropic API to run auto mode's safety classifier inside each request. That check is done by Anthropic's servers, which are not in the path here, so without this setting Claude Code shows a “this session isn't eligible” notice before its first checked action and then falls back to its own classifier calls, billed as normal tokens on your key. CLAUDE_CODE_AUTO_MODE_SERVER=0 skips the ask and the notice. Auto mode itself works the same either way.
OpenAI API
curl https://api.engy.ai/v1/chat/completions \
-H "Authorization: Bearer $ENGY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.2","messages":[{"role":"user","content":"hello"}]}'
from openai import OpenAI
client = OpenAI(base_url="https://api.engy.ai/v1", api_key="$ENGY_API_KEY")
client.chat.completions.create(model="glm-5.2",
messages=[{"role":"user","content":"hello"}])
Cursor
Cursor Settings → Models → API Keys, under OpenAI API Key:
OpenAI API Key $ENGY_API_KEY Override OpenAI Base URL https://api.engy.ai/v1 ← enable the toggle, then Verify
Under Models, + Add model → engy/glm-5.2, enable it, and select it in chat.
Cursor needs the slash. It wants a custom model named in vendor/model form, so the bare glm-5.2 the rest of this page uses is the one spelling it will not take. Prefix it with engy/ and any id from the models list works — engy/qwen3.8-27b, engy/deepseek-v4-flash-0731, and so on. Same model, same price; answers echo back whichever spelling you sent, and the id after the slash still has to be exact.
Custom endpoints drive the chat/plan panel; Tab and Composer stay on Cursor's own models. Requests are relayed by Cursor's servers, so it works from any machine. Needs a paid Cursor plan: on the free plan Cursor shows "Named models unavailable" for any custom model (free is Auto-only, that is Cursor's gating, not this API).
Codex
In ~/.codex/config.toml, then export ENGY_API_KEY=… and run codex:
model = "glm-5.2" model_provider = "engy" [model_providers.engy] name = "engy" base_url = "https://api.engy.ai/v1" env_key = "ENGY_API_KEY" wire_api = "responses"
Hermes
In ~/.hermes/config.yaml, then run hermes chat:
model: provider: "custom" default: "glm-5.2" base_url: "https://api.engy.ai/v1" api_key: "$ENGY_API_KEY"
Reasoning effort
Steer how long a model thinks with OpenAI's reasoning_effort, or OpenRouter's reasoning object. Both work on /v1/chat/completions, and you do not need to know which one a given model prefers.
curl https://api.engy.ai/v1/chat/completions \
-H "Authorization: Bearer $ENGY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.2",
"reasoning_effort":"high",
"messages":[{"role":"user","content":"Prove 2 is irrational."}]}'
The OpenRouter form is equivalent: "reasoning": {"effort": "high"}.
The tiers
Accepted values, lowest to highest: none, minimal, low, medium, high, xhigh, max. none and minimalboth mean “do not think”. Case does not matter. A value outside this list is forwarded unchanged, and depending on the model it is either refused with a 400 or ignored.
What each model hears
This is the part worth knowing: every model's chat template speaks its own dialect, so the same word does not mean the same thing everywhere. engy translates your tier into what the target model actually understands, rather than passing it through and letting the template decide. Sending high straight to Qwen3.8, for instance, is a 400 from its template — engy maps it to xhigh instead.
| you send | GLM-5.2 | GLM-5.3 | Qwen3.8 | Kimi-K3, DeepSeek V4 | DeepSeek V4.1 |
|---|---|---|---|---|---|
| low | high | low | low | low | low |
| medium | high | low | medium | low | high |
| high | high | high | xhigh | high | high |
| xhigh | max | max | xhigh | max | xhigh |
| max | max | max | xhigh | max | max |
The GLM-5.3 column covers GLM-5.3-Flash, and the DeepSeek V4.1 column covers V4.1-Flash; “DeepSeek V4” means the V4 models without the .1. Two consequences to plan around. GLM-5.2 collapses low/medium/high onto one level, so below xhigh you are choosing between only two settings. Qwen3.8 tops out at xhigh, which is also its default, so max asks for nothing extra. Any other model, Qwen3.6 included, gets the GLM-5.2 column.
Turning thinking off
"reasoning_effort":"none" (or "minimal", or "reasoning":{"enabled":false}) turns thinking off. engy drops the effort hint and sets the model's own template switch: thinking = false for Kimi-K3 and DeepSeek V4.1, enable_thinking = false for every other model (GLM-5.2, Qwen, DeepSeek V4). If you set that switch yourself in chat_template_kwargs, your value wins. Kimi-K3 and DeepSeek V4.1 also accept enable_thinking, which engy renames to thinking for you.
One exception: GLM-5.3 and GLM-5.3-Flash cannot turn thinking off. Their template has no switch, so engy runs them at their lowest level, low, instead. They still think briefly, so leave room for it in max_tokens.
If thinking stays on, the setting is probably not reaching us: some client libraries drop reasoning_effort for models they do not recognise. Sending the switch directly works with any client, for example in Python extra_body={"chat_template_kwargs": {"enable_thinking": False}}.
On /v1/messages, send "thinking": {"type": "disabled"}. engy translates it into each model's own switch exactly as above, including the GLM-5.3 exception (it runs at low). A budget_tokens value is not used: the model thinks at its default level.
Image input
Vision models take images through the standard OpenAI image_url content part, so any client that already speaks multimodal chat works unchanged. Which models accept one is live fleet state, not a fixed list: the pricing table carries a modalities column, and https://api.engy.ai/v1/models reports the same thing as input_modalities per model. Send an image to a text-only model and you get a 400 naming the modality it does not support, before the request is dispatched.
Send an image
IMG="data:image/png;base64,$(base64 -w0 photo.png)" # macOS: base64 -i photo.png
curl https://api.engy.ai/v1/chat/completions \
-H "Authorization: Bearer $ENGY_API_KEY" \
-H "Content-Type: application/json" \
-d @- <<JSON
{"model": "kimi-k3",
"messages": [{"role": "user", "content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "$IMG"}}]}]}
JSON
Text and images mix freely in one turn, and across turns — put as many image_url parts in the content list as you need, up to the limits below. Output is always text.
Inline, never linked
The image must be a base64 data:image/…;base64, URI. An http(s) URL is rejected with a 400 rather than fetched, as is a data: URI that is not an image or not base64, and this is deliberate: the gateway ships media inline to the miner, and miners run on egress-restricted GPU clusters that genuinely cannot reach most of the public internet. Fetching buyer-supplied URLs would also make the fleet an SSRF gadget. A clear error at the edge beats a confusing miner-side failure, so download the image yourself and inline it.
Limits
| limit | value | over it |
|---|---|---|
| media per request | 12 MiB decoded, all parts summed |
413
|
| media parts per request | 100 |
413
|
| video | not supported |
400
|
The budget is on decoded bytes, so a base64 payload roughly a third larger than that still fits. It is sized well under the miner control-plane frame limit: one oversized request would otherwise take down a miner connection that is also carrying its heartbeats.
Both limits count every media part in the request, audio included. For billing headroom each media part reserves 1,024 prompt tokens up front, about a mid-size image's encoder cost. That only sizes the pre-flight reservation; you are settled on the tokens the model actually reports.
Anthropic-style blocks
On the Anthropic-compatible endpoint, a {"type":"image"} block with a base64 source is translated into the OpenAI part shape for you, media type included. A url source is passed through as-is and then meets the same data:-only rule above, so inline it there too. Images are read from user turns only; an image block in an assistant turn is dropped.
Raw prompts and logprobs
/v1/completions takes a raw prompt instead of a message list: no chat template, no system prompt, and no reasoning parser rewriting <think>, so the model sees the exact bytes you send. Use it to continue a prompt from mid-turn, or to score text you supply. It does not stream: stream: true is refused with a 400.
Continue a raw prompt
The model picks up where your prompt stops. Build the prompt with the model's own chat template (tokenizer.apply_chat_template(messages, add_generation_prompt=True)) rather than by hand; the string below is exactly what GLM-5.2's template renders for one user turn:
curl https://api.engy.ai/v1/completions \
-H "Authorization: Bearer $ENGY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.2",
"prompt":"[gMASK]<sop><|system|>Reasoning Effort: Max<|user|>Prove 2 is irrational.<|assistant|><think>",
"max_tokens":512,"temperature":0}'
Score a span
To read the log-probability of text you supplied, send logprob_start_len: the token offset where your span begins. The response carries token_logprobs and token_ids, one entry per scored token. A call scores up to 1024 tokens (the current default) over any prefix length. Pass prompt as token ids so the offsets are the ones you computed, and always send max_tokens=1:
from openai import OpenAI
client = OpenAI(base_url="https://api.engy.ai/v1", api_key="$ENGY_API_KEY")
ids = tokenizer.encode(prefix + action) # your own tokenizer
span = len(tokenizer.encode(action)) # tokens you want scored
r = client.completions.create(model="glm-5.2", prompt=ids,
max_tokens=1, temperature=0,
extra_body={"logprob_start_len": len(ids) - span})
lp = r.choices[0].logprobs
print(sum(x for x in lp.token_logprobs if x is not None))
Memory scales with the tokens you SCORE, not with prompt length, which is what makes a short span over a 20k-token prefix practical. Prefill still costs what prefill costs: a 64-token span over a 14,900-token prefix measured 1.2 s to 4.2 s end to end on production, depending on what else the backend was serving.
The 1024 is on the SCORED span, not the prefix, which is unbounded up to the context window. 1024 passes, 1025 is refused with a 400 quoting the length it computed, so split longer spans across calls. The limit is set per model, not per account. Send the prompt as token ids: a string is not tokenised at the edge, only estimated from its length, so the span cannot be measured, and any string prompt estimated over the limit is refused however small the span.
Score a whole prompt
echo returns a log-probability for every prompt token. The logits tensor behind it grows with prompt length until it will not fit, so echo is capped at the same 1024 tokens, applied here to the whole prompt. Past that the call is refused with a 400 pointing at logprob_start_len rather than sent to a backend that cannot serve it:
curl https://api.engy.ai/v1/completions \
-H "Authorization: Bearer $ENGY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.2","prompt":"The capital of France is",
"max_tokens":1,"temperature":0,"echo":true,"logprobs":1}'
Gotchas
Always send max_tokens=1 on a scoring call. Leave it out and the request still succeeds, but the model first generates up to its own output limit, tens of thousands of tokens on most models, billed as output, for exactly the same log-probabilities a 1-token call returns in under a second.
max_tokens and logprobs must be at least 1 whenever you send them: 0 is refused with a 400 on both paths. Leaving logprobs out on the echo path returns a normal 200 with no log-probabilities at all, the one failure here that looks like success.
logprob_start_len is an absolute offset from the START of the prompt, so to score the last n tokens send prompt_tokens - n, not n. Get it backwards and the call scores the wrong span without complaint, unless that span is over the limit, in which case the 400 quotes the length it computed. Check that you got back n log-probabilities.
A span call returns only token_logprobs and token_ids, with no text_offset, so align on the token_ids. On the echo path text_offset comes back as -1 throughout.
With max_tokens=1 a scoring call is prefill plus one token, so it bills essentially as input. Log-probabilities can differ slightly between backends serving the same model, so contact us before a run where the numbers are compared to each other.
For generated tokens, /v1/chat/completions takes the standard logprobs and top_logprobs (up to 30), streaming included. It does not score the prompt; that is what /v1/completions is for.
Data retention
Zero data retention is the default for every account: engy does not store your prompts or model outputs. You can choose to opt out, for one request or for your whole account. engy then keeps the prompt and model output of each request your opt-out covers, for its own research only (never sold or published). Details are in the privacy policy.
Opt out for one request
Add OpenRouter's provider object with "data_collection": "allow". It works on /v1/chat/completions, /v1/completions, /v1/messages and /v1/responses, and is removed before the request reaches the model.
curl https://api.engy.ai/v1/chat/completions \
-H "Authorization: Bearer $ENGY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"kimi-k3",
"provider":{"data_collection":"allow"},
"messages":[{"role":"user","content":"hello"}]}'
What is kept is filed under the request id returned in x_engy.request_id. From the OpenAI Python client, pass the object through extra_body={"provider": {"data_collection": "allow"}}.
Opt out for your account
For clients that cannot add fields to each request, the opt-out can be set on the whole account: email us from the address on your account and we will switch it. Every request from every key on the account is then covered, until you ask us to switch it back.
Keeping a request out
A request that sends "data_collection": "deny" or "zdr": true is never kept, even when the account has opted out. Any other value of data_collection is refused with a 400, so a typo cannot be read as either answer.
Agent API
Read-only, programmatic access to balance, usage, and models for agents. See the Agent API docs.