Shared computedocs
Quickstart
Docs
Connect a supported OpenAI-compatible client to CueCloud hosted inference. Model-specific rates are confirmed during onboarding. Full reference on the API page.
1. Point at Cue Cloud
- Base URL:
https://api.cuecloud.io/v1 - Auth:
Authorization: Bearer cue_…from the console. - POST
/chat/completionswith the model ID supplied for your enabled workspace.
Requests carry Authorization: Bearer cue_… from the console. Do not share keys across the team.
2. First request
curl https://api.cuecloud.io/v1/chat/completions \
-H "Authorization: Bearer cue_••••••••" \
-H "Content-Type: application/json" \
-d '{
"model": "YOUR_ENABLED_MODEL_ID",
"messages": [
{ "role": "user", "content": "fix the failing test" }
]
}'3. Stream
Set stream: true for SSE token chunks. The normal path for IDE agents. Non-streaming completions work for scripts.
curl https://api.cuecloud.io/v1/chat/completions \
-H "Authorization: Bearer cue_••••••••" \
-H "Content-Type: application/json" \
-N \
-d '{
"model": "YOUR_ENABLED_MODEL_ID",
"stream": true,
"messages": [
{ "role": "user", "content": "refactor this module" }
]
}'Models
Selected open-weight models we host. Details on the hosted catalog.
| Access | Model | Best for |
|---|---|---|
| Confirm API ID during onboarding | GLM 5.3 | Hosted inference for your applications and coding workflows |
| Confirm API ID during onboarding | DeepSeek V4.1 Flash | Hosted inference for your applications and coding workflows |
| Confirm API ID during onboarding | Qwen 3.8 27B | Hosted inference for your applications and coding workflows |
Clients
Compatible with supported OpenAI-style clients. Tested or documented on the API page: CueCode, Cursor, scripts. Other OpenAI-compatible tools are expected, not claimed here.
- Cursor. Add an OpenAI-compatible provider. Base URL → Cue Cloud. API key → your cue_… key. Model → one of the flagship ids.
- CueCode. Cloud mode on Cue Cloud. Hosted model access; CueCode can send denser agent signals for this stack.
- Scripts & tools. Any OpenAI SDK or raw HTTP client works. Use it in CI or custom agents. We optimize serving for IDE coding workloads.
Errors
| Code | Meaning |
|---|---|
401 | Missing or invalid cue_… key |
400 | Bad request or unknown model id |
503 | Capacity path unavailable. Retry or check status. |