Open betaapply for access. We review each application and email you when it's approved, with your API key ready on the dashboard.

Examples

Short, copy-pasteable end-to-end snippets for each feature. The curl examples assume ACS_API_KEY + ACS_API_BASE are exported (see Quick start); the Python examples use the same openai SDK client.

Each page below shows one shared example response after the curl and Python snippets — both calls return the same JSON; the SDK just wraps it in typed objects. Always-null fields (service_tier, system_fingerprint, kv_transfer_params, etc.) are dropped from the shown responses for brevity.

Your own output won't match token-for-token: the displayed requests don't pin a seed, so under the default temperature=1.0 sampling varies. Add a seed to reproduce a specific run, or temperature=0 for greedy decoding.

Guided walkthrough

  • Hands-on Colab workshop — the from-zero path through most of what's below, run end-to-end in your browser: completions → sampling → logprobs → reading activations → building a steering vector, with exercises and an instructor solutions notebook — both in our tutorial repo. Start here if the per-feature snippets feel piecemeal.

Token-level inspection

  • logprobs — top-k logprobs per generated token.
  • prompt_logprobs — top-k logprobs at each prompt position (gotchas around rank-vs-actual-token).
  • echo — include the prompt in the response.

Streaming + concurrency

  • stream — SSE token-by-token.
  • Batch rollouts — 16 concurrent per key, with the canonical asyncio.gather pattern.

Recovery patterns

Evaluation

  • Evaluate with Inspect — run base-model evals through Inspect's openai-api-completions provider (raw prompts, no chat template).