HOSTED CODING AGENTS
Keep the loop moving.
Coding agents make repeated model calls around local tools, tests, and review. Shared compute supplies the hosted inference behind that work.
Model calls are one part of the work.
The agent reads context, requests a model response, executes permitted tools, and checks the result. Every turn can add both input and output tokens to the monthly usage.
- Prepare approved context
- Request hosted inference
- Run tools in your client environment
- Review the output and continue
CueCode can use hosted inference. It is also the IDE component of Blitz, which is deployed privately with your organization’s tools and controls.
Explore private engineering with Blitz ↗Inference runs on infrastructure we operate.
CONNECT YOUR WORKFLOW
One endpoint. Your tools.
Make your first request.
- Get access.
Request hosted access. Once enabled, get your API key from the console.
- Configure your client.
Use the base URL and the model ID provided for your enabled workspace. Keep your key out of source control.
- Check the response.
Try a small, approved coding task before connecting a repository or agent workflow.
curl https://api.cuecloud.io/v1/chat/completions \
-H "Authorization: Bearer cue_••••••••" \
-H "Content-Type: application/json" \
-d '{
"model": "YOUR_ENABLED_MODEL_ID",
"messages": [
{ "role": "user", "content": "fix the failing test" }
]
}'Example only. Replace the key and YOUR_ENABLED_MODEL_ID with your enabled workspace values.
YOUR NEXT STEP
Ready to connect?
Request hosted access and tell us about your workload. We’ll follow up on availability and onboarding. Creating a console account is a separate step and does not itself grant inference access.
Create a console account
Already have an account? Sign in ↗