---
title: Agents as models
description: Call any agent from the OpenAI or Anthropic SDKs.
---

Every agent can answer in the format the model providers use. Point an SDK at Keiki, name
the agent as the `model`, and one completion becomes one turn of that agent. The turn uses
the agent's own model, instructions, tools and memory, and it is traced like any other.

Keiki is a translation layer, not a proxy to OpenAI or Anthropic. Sampling settings your
client sends (`temperature`, `max_tokens`, `seed`) are ignored, because the agent's
configuration decides them. Settings that describe the response instead of the model are
honoured: `stream`, the tools your client runs itself (`tools`), how much the model thinks
(`reasoning_effort` or `thinking`), whether it must use a tool (`tool_choice`), and the shape
of the answer (`response_format`). Every streamed completion ends with a usage chunk. A
conversation that arrives this way is stored, searchable and traced just like one that
arrived over SMS.

## Turn it on

Go to **Agents**, open your agent, then **API gateway** and select **Turn on**. While it is
off, no client can name the agent as a model. `/models` leaves it out, and a call that names
it gets the provider's permission error.

The gateway holds no provider secret. Callers authenticate with your organization's own
developer API keys (`sk_live_…` or `sk_test_…`), so revoking a key revokes this surface too.

## OpenAI protocol

Base URL: `https://<host>/api/v1/openai`

| Endpoint | What it does |
| --- | --- |
| `POST /chat/completions` | One turn of the agent, streaming or not. |
| `GET /models` | The agents in your organization whose gateway is on. |

```ts
import OpenAI from 'openai'

const client = new OpenAI({
  baseURL: 'https://<host>/api/v1/openai',
  apiKey: process.env.KEIKI_API_KEY,
})

const stream = await client.chat.completions.create({
  model: 'support', // the agent's name, or its id
  messages: [{ role: 'user', content: 'where is my order?' }],
  stream: true,
})
```

Anything that speaks this protocol works the same way, including the Vercel AI SDK,
LangChain and other OpenAI-compatible clients:

```ts
import { createOpenAI } from '@ai-sdk/openai'
import { streamText } from 'ai'

const keiki = createOpenAI({
  baseURL: 'https://<host>/api/v1/openai',
  apiKey: process.env.KEIKI_API_KEY,
})

// keiki.chat(), not keiki(): the AI SDK's default factory speaks the
// Responses API, and this surface is Chat Completions.
const result = streamText({ model: keiki.chat('support'), prompt: 'hey' })
```

A non-streaming completion also carries a `keiki` object that names the conversation, the
agent and the turn's trace, so a caller can open what it just ran in the dashboard.

## Anthropic protocol

Base URL: `https://<host>/api/anthropic`

| Endpoint | What it does |
| --- | --- |
| `POST /v1/messages` | The same turn, in Anthropic's message format. |
| `POST /v1/messages/count_tokens` | An estimate of the text you are about to send. |

```ts
import Anthropic from '@anthropic-ai/sdk'

const client = new Anthropic({
  baseURL: 'https://<host>/api/anthropic',
  apiKey: process.env.KEIKI_API_KEY,
})

const message = await client.messages.create({
  model: 'support',
  max_tokens: 1024, // accepted and ignored: the agent's config decides
  messages: [{ role: 'user', content: 'where is my order?' }],
})
```

On either surface the key may travel as `x-api-key` or as `Authorization: Bearer …`.
`count_tokens` estimates the text in the request itself. What a turn actually spends depends
on the agent's prompt, its history window, and the tools it ends up calling, none of which
exist until it runs.

## Conversations

Send `x-conversation-id: <your id>` to keep several calls in one conversation. They share one
thread in the dashboard and one thread in the agent's memory. Without the header, every call
starts a new conversation.

The agent's own configuration travels with it: its prompt, tools, memory and history window
apply to an API turn as they do to a text or a Slack thread. If `messages` carries only the
message you want answered, the agent reads the conversation stored under that
`x-conversation-id` under its own history window, the same as on any other channel — there
is nothing to re-send and nothing to configure. If `messages` carries earlier turns, that
transcript is the context instead: an SDK that re-sends the whole thread on every call gets
the conversation it believes it is having. A client `system` message rides along as context
either way, but it does not replace the agent's own prompt.

A conversation runs one turn at a time. A second call while a turn is still running is
refused as busy, rather than queued behind an answer the caller will never read. Use a fresh
conversation id for calls you want to run at the same time.

## Streaming and tools

Streaming is real. Tokens are forwarded as the model produces them, from the first one, and
they continue across the agent's tool calls. The answer on this surface is the completion's
own text, because the response you are reading is the delivery. A message another tool sends,
such as an email or a text to someone else, arrives in the stream too.

Streaming from the first token has one consequence worth knowing. The agent's reply guards
run on the finished completion, so a completion the agent then rejects has already been read
by the client. That can happen if it is blocked, degenerate, or an echo of a tool's output.
The stream then ends with the protocol's error event instead of a normal stop, because no
provider protocol can unsend bytes. The stored turn holds the answer the agent settled on. A
non-streaming call never shows the rejected text at all. It returns the settled answer.

Keiki retries upstream only before the first byte reaches the client. After that, a failure
travels as the protocol's in-stream error, an `error` frame on the OpenAI surface or an
`error` event on the Anthropic one, which is what both SDKs raise on. Hanging up mid-answer
stops the turn.

## Reasoning

Each agent has a reasoning level, and a request can replace it for that call. The hidden
thinking phase runs before the first token, so it is the difference between an answer that
starts in a couple of seconds and one that starts in ten. A chat UI usually wants `low` or
none. An offline batch can afford `high`.

```ts
await client.chat.completions.create({
  model: 'support',
  reasoning_effort: 'none', // none | minimal | low | medium | high
  messages: [{ role: 'user', content: 'where is my order?' }],
})

await anthropic.messages.create({
  model: 'support',
  max_tokens: 1024,
  thinking: { type: 'enabled', budget_tokens: 8000 }, // or { type: 'disabled' }
  messages: [{ role: 'user', content: 'where is my order?' }],
})
```

A client that sends `reasoning: { effort }` or `reasoning: { enabled: false }` is read the
same way as `reasoning_effort`. A turn takes a level rather than a budget, so a token budget
is mapped onto one: under 4k becomes `low`, under 16k becomes `medium`, above that `high`,
and zero turns it off. OpenAI's `minimal` is read as off. Send nothing and the agent's own
setting stands. A level that is not one of these is a 400, rather than a quietly different
answer.

## Your own tools

The tools you declare in `tools` join the agent's own, and the model is offered one list.
Who runs what is the difference:

- the agent's tools, which are its capabilities, MCP servers and authored tools, run on
  Keiki inside the turn, and you never see them;
- your tools run in your process, so a call to one ends the response with
  `finish_reason: "tool_calls"` (Anthropic: `stop_reason: "tool_use"`) and waits for you to
  send the results.

Send the results back the way you would to the provider, with a `tool` message per call on
the OpenAI surface or `tool_result` blocks on the Anthropic one, and keep the same
`x-conversation-id`:

```ts
const first = await client.chat.completions.create({
  model: 'support',
  messages: [{ role: 'user', content: 'is my order out for delivery?' }],
  tools: [{ type: 'function', function: { name: 'get_location', parameters: {} } }],
})

const call = first.choices[0].message.tool_calls![0]
await client.chat.completions.create({
  model: 'support',
  messages: [
    { role: 'user', content: 'is my order out for delivery?' },
    first.choices[0].message,
    { role: 'tool', tool_call_id: call.id, content: 'Paris' },
  ],
  tools: [{ type: 'function', function: { name: 'get_location', parameters: {} } }],
})
```

The second call continues the same turn rather than starting one. The run is parked with its
progress, and the tools it already called are not called again, so nothing it did is
repeated. A conversation holds one parked run, which can be resumed for 24 hours and answered
once. Sending the same results twice gets you a 400, not a replay.

A few details are worth knowing:

- a tool named the same as one of the agent's tools stays the agent's, and your declaration
  is dropped;
- names must match `[A-Za-z0-9_-]{1,64}`, a name is declared once, and a request may carry at
  most 64 tools;
- keep declaring your tools on the follow-up request, which an SDK does for you;
- a result you leave out is answered for you with an error, and a result over 32k characters
  is truncated before the model reads it;
- your tools cannot reach anything of Keiki's. They are names and schemas in a prompt, run
  entirely by you.

The agent's own authored tools do have storage of their own: inside one, `ctx.storage`
is a durable JSON row store shared by every conversation of the agent (a roster, a lookup
table), alongside `ctx.userState`, which is per end user. `ctx.storage` exposes
`query(entity, { filter, limit, orderBy })`, `get(entity, id)`, `count(entity)`,
`insert(entity, data)`, `update(entity, id, data, expectedVersion)` and
`delete(entity, id)`. The same rows are reachable over the developer API:
`PUT /api/v1/developer/agents/:agentId/storage/:entity` replaces an entity's rows
atomically (bulk-load a roster; `{rows: []}` clears it), and
`GET /api/v1/developer/agents/:agentId/storage/:entity?limit=&offset=&filter=` reads them back,
`filter` an optional JSON-object containment match that `total` respects, paged with `offset`.

## Forcing, forbidding and shaping the answer

`tool_choice` applies to this request rather than the agent, so it is honoured. `"none"`
makes the agent answer without calling anything, `"required"` (Anthropic `{type:'any'}`)
makes it open with a tool call, and naming a tool picks which one. A turn is a loop, so a
forced choice constrains the first completion only, because re-forcing a tool at every step
could never reach an answer. `"none"` holds throughout. A named tool must be one you declared
in `tools`. The agent's own tools are its configuration, not yours to select, and naming
something you did not declare is a 400.

`response_format` shapes the answer on the OpenAI surface. The Anthropic protocol has no
equivalent field.

```ts
await client.chat.completions.create({
  model: 'support',
  response_format: {
    type: 'json_schema',
    json_schema: { name: 'order', schema: { type: 'object', properties: { id: { type: 'string' } } } },
  },
  messages: [{ role: 'user', content: 'where is my order?' }],
})
```

`{type:'json_object'}` asks for JSON, `{type:'json_schema'}` holds the model to your schema,
and `{type:'text'}` is the default. It applies to every completion of the turn, so the answer
keeps its shape however many steps the agent took to get there.

## Passthrough mode

Go to **Agents**, open your agent, then **API gateway** and select **Passthrough mode**. With
it off, a turn is the agent as configured. Your `system` message rides along as context, and
the agent's own prompt, tools and memory decide the answer. With it on, the caller owns the
prompt. Your `system` or `developer` messages become the turn's system prompt, only the tools
you declared in the request are offered, and none of the agent's skills, memory or reply
guards enter the turn. The agent still picks the model and the reasoning level, and the turn
is stored and traced like any other. That is what makes it a way to run a model call you
already wrote the prompt for through Keiki's observability.

One thing of the agent's does enter the turn: its stored system prompt, appended after
yours. It stays empty until someone writes it, either in the dashboard or through an
eval-driven fix you accept. That is how a passthrough agent improves without you touching
your code. Your prompt stays the contract and keeps winning on every request, Keiki's
amendments ride behind it, and the evals measure the two together exactly as production runs
them. Your prompt itself is only recorded, as the spec the evals are generated from. It is
never written back to the agent.

Tools work the provider's way on a passthrough gateway. The request is the whole exchange.
Each call runs one completion over exactly the `messages` you sent, including any assistant
`tool_calls` and `tool` results in the order they happened, and a call to one of your tools
ends the response. Nothing is parked on the server. The next request carries the transcript,
and each request is its own traced turn under the same `x-conversation-id`. A loop you
already run against a provider can point here unchanged, and the results you send back are
the ones the model is shown, whichever ids they carry.

Passthrough belongs to the gateway, not to the agent. The same agent keeps answering as
itself over SMS, Slack or the dashboard. It applies to the managed harness. An agent running
on your own server (`executionMode: 'sdk'`) receives the turn as it always did. A turn parked
on one of your tools finishes in the mode it started in, even if you flip the toggle before
the results come back.

## Errors

Errors use each provider's envelope, so an SDK raises the exception type a client already
handles.

| Situation | Status |
| --- | --- |
| Missing or invalid API key | 401 |
| Key without the `agents:chat` scope | 403 |
| No such agent in the organization | 404 |
| The agent's gateway is off | 403 |
| A turn is already running for the conversation | 409 |
| The turn itself failed | 500 |

A `tool_choice` that names a tool you did not declare, or a `response_format` that is not one
of the three shapes, is a 400. The turn never runs, rather than quietly answering something
else.

To create agents for your own users from your backend, see
[Partner deployments](/docs/partner-deployments).
