File size: 15,056 Bytes
8f16a6b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 | # AgentBridge HTTP API β reference
OpenAI-compatible endpoints plus a small set of **documented proprietary extensions**
for the features that have no OpenAI equivalent (voice speech, LLM switching, platform
capabilities), plus a native MCP JSON-RPC connector. Any OpenAI SDK, script, standalone
client or MCP client can drive the AI agents without modification to the agent core.
## Endpoint summary
### Standard (OpenAI-compatible)
| Endpoint | Purpose |
|---|---|
| `POST /v1/chat/completions` | Chat with the agents (streaming SSE, sessions, LLM switching) |
| `POST /v1/files` | Multipart upload + server-side Markdown conversion |
| `GET /v1/files` Β· `GET /v1/files/{id}` | List / retrieve converted files |
| `GET /v1/files/{id}/content` | Raw uploaded bytes (OpenAI Files API) |
| `DELETE /v1/files/{id}` | Delete an uploaded file |
| `GET /v1/models` | Agent sets **and** LLM providers with their characteristics |
| `GET /v1/models/{id}` | Single model details |
| `POST /v1/audio/speech` | Text-to-speech β WAV bytes (Kokoro neural TTS) |
| `GET /health` | Liveness probe |
### Proprietary extensions (documented, additive β ignored by strict OpenAI clients)
| Endpoint | Purpose |
|---|---|
| `POST /v1/control` | Pilot/steering: switch the LLM in use, toggle features, reset history, create sessions |
| `GET /v1/control` | Session state + platform capabilities (what is available here and now) |
| `POST /v1/voice/listen` | One-shot speech recognition from the server microphone (Windows only) |
| `GET /v1/audio/voices` | TTS voices available on this platform |
| `POST /mcp` | Native MCP JSON-RPC endpoint (`initialize`, `tools/list`, `tools/call`) |
> **Telegram is an in-process medium and exposes no HTTP endpoints** β messages travel
> directly through the WTelegramClient library; configuration is done from the TUI
> (`/telegram`) or in `telegram.json` (see [Telegram chat](#telegram-chat-an-in-process-medium-no-http-endpoints)).
The rule for platform-dependent features: the server reports them **unavailable (501)** when
the platform or the assets are missing, and `GET /v1/control` / `GET /v1/audio/voices` always
tell the client what is actually available β a chat client activates voice/TTS only where they
really run.
---
## `POST /mcp` β native MCP JSON-RPC connector
AgentBridge exposes a native MCP connector in the same process as the agent runtime, so MCP,
OpenAI API and TUI all drive the same orchestrator state.
Current minimal profile (intentionally small for immediate interoperability):
- `initialize`
- `tools/list`
- `tools/call`
The initial tool catalog exposes one high-level tool:
- `agent_run` β runs an autonomous AgentBridge execution for the provided prompt.
Example request:
```json
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "agent_run",
"arguments": {
"prompt": "Create a concise weekly report from the latest sales data",
"model": "default-agent",
"llm_provider": "Zai",
"max_iterations": 120,
"session_id": "sess-..."
}
}
}
```
`agent_run` returns MCP content blocks plus a structured payload (`success`, `code`,
`iterations`, optional `session_id`, optional `attachments`).
---
## `POST /v1/chat/completions` β chat with the agents
OpenAI Chat Completions compatible. `model` selects which agent set is used
(see `GET /v1/models`); `stream: true` returns Server-Sent Events (SSE).
```bash
curl -N http://localhost:5290/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "web-agent",
"messages": [{"role": "user", "content": "What is the weather today?"}],
"file_ids": ["file-..."],
"stream": true
}'
```
| Field | Meaning |
|---|---|
| `model` | Agent set (see `GET /v1/models`): `default-agent`, `web-agent`, `search-agent`, `research-agent`, `document-files`, `spreadsheet-files`, `email-agent`, `office-files`, `multi-files`, `all-files`. |
| `tools` | **Extension** β explicit tool-name list (e.g. `["FileTool", "OfficeTool", "EMailTool"]`) that overrides the preset from `model`. Unknown names are skipped; an empty list falls back to the preset. The core tools (`FileTool`, `GitTool`) are always part of the presets and of the TUI custom combinations β the TUI cannot remove them; the only way to change their status is `tools.json` (see below). |
| `messages` | OpenAI messages; the last `user` message is the prompt. |
| `file_ids` | Optional ids from `POST /v1/files` β attached as context (Markdown, server-side). |
| `max_tokens` | Roughly maps to agent loop iterations (`max_tokens / 100`, clamped 1β50). |
| `stream` | `true` β SSE chunks; `false` (default) β single JSON response with `usage`. |
| `session_id` | **Extension** β multi-turn session id (see [Sessions](#sessions-multi-turn-memory)). |
| `llm_provider` | **Extension** β LLM provider for this request (see [LLM switching](#llm-switching-the-pilot-endpoint)). |
> **Per-tool configuration (`tools.json`).** A JSON file next to the executable overrides a
> tool's default status β `{"tools": {"OfficeTool": true}}` enables the class-B `OfficeTool`
> (default OFF), `{"tools": {"FileTool": false}}` disables a core tool. The rule is
> **"unspecified β ON"**: a tool with no explicit entry uses its default (class-A tools ON,
> class-B tools OFF). An absent file means all defaults. The file is never overwritten by
> updates (same pattern as `telegram.json`). The dynamic `all-files` preset resolves to
> every loaded tool the config leaves enabled.
Responses carry an additive `session_id` field when a session was used.
> **Streaming caveat**: LLM-native streaming (`SendQueryStream`) does not support
> anonymization and throws for Gemini β the `/v1/chat/completions` SSE endpoint here is
> response-side only (the agent result is computed with non-streaming `SendQuery`).
## Sessions (multi-turn memory)
By default every request is stateless (fresh orchestrator, fresh history). Passing a
`session_id` keeps the conversation history across requests:
1. Create a session: `POST /v1/control {"create": true}` β returns the `session_id`
(or omit `session_id` on the first chat request β the response returns the new id;
for `stream: true`, create the session via `/v1/control` first).
2. Send chat requests with `"session_id": "sess-..."` β the agent remembers previous turns.
3. Inspect/reset: `GET /v1/control?session_id=...` and `POST /v1/control` with
`reset_history: true`.
Sessions are in-memory, expire after 30 minutes of inactivity, and are serialized (one chat
at a time per session). Unknown `session_id` β `404`.
## LLM switching (the pilot endpoint)
The LLM provider is not a server-wide constant: it can be changed **on the fly**, like
switching models in a code editor β per request, or per session. There is no OpenAI-standard
way to do this, so the server exposes the **`POST /v1/control` pilot endpoint** (proprietary
but stable and extensible):
```json
// switch the LLM currently in use for a session
{ "session_id": "sess-...", "llm_provider": "Zai" }
// toggle feature flags (extensible for future features)
{ "session_id": "sess-...", "features": { "voice": true, "tts": true } }
// start a fresh conversation
{ "session_id": "sess-...", "reset_history": true }
// create a session
{ "create": true }
```
`GET /v1/control?session_id=...` returns the full session state: provider in use, model name,
**context window, history size and estimated history tokens**, feature flags and platform
capabilities.
**Context-window guard.** A switch is refused with **`409 context_window_exceeded`** when the
accumulated conversation overflows the target provider's context window β the exact
"on-the-fly switch conflicts with the context window of the model in use" case:
```json
{
"error": "context_window_exceeded",
"detail": "The conversation needs β44744 tokens but provider 'ExllamaV2' has a context window of 8192 tokens. Reset the conversation (POST /v1/control with reset_history: true) or switch to a provider with a larger context window.",
"estimated_tokens": 44744,
"context_window": 8192,
"provider": "ExllamaV2"
}
```
The same check applies to a per-request `llm_provider` on a session chat. The switch itself
preserves the conversation (history is moved to the new provider's utility). Note that some
providers block while being activated β e.g. `ExllamaV2` auto-starts the local
ExLlamaV2 server and waits for it to become ready (up to 3 minutes, then it fails).
Per-request switching without a session works too: `"llm_provider": "Zai"` on any
`/v1/chat/completions` body.
## TTS β `POST /v1/audio/speech` (standard OpenAI)
In-process **Kokoro neural TTS** (the same engine/voices as the Windows VoiceAgent, but
cross-platform β it runs on Windows **and** Linux). Request:
```bash
curl http://localhost:5290/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"input":"Ciao! Oggi Γ¨ una bella giornata.","voice":"alloy","speed":1.0}' \
-o speech.wav
```
- `input` (required), `voice` (OpenAI names `alloy`, `echo`, `fable`, `onyx`, `nova`,
`shimmer`, `coral`, `sage`, `ash`, `ballad`, `verse` **or** raw Kokoro ids like `if_sara`,
`af_heart` β see `GET /v1/audio/voices`), `speed` (0.25β4.0, default 1.0).
- **`lang`** (extension): two-letter ISO language. Kokoro voices are per-language
(`if_*` Italian, `af_*`/`am_*` English, `ef_*` Spanish, `ff_*` French, `jf_*` Japanese, ...).
When `lang` is omitted the **server's system language** selects the voice β an Italian
machine speaks Italian (`if_sara`), not accented English. A named `voice` of a different
language is overridden by `lang` (e.g. `alloy` + `lang: it` β an `if_*` voice).
- Response: `audio/wav` (24 kHz mono 16-bit PCM).
- `response_format` accepts `wav` (default); others β `400`. `model` is accepted for
compatibility and ignored.
- **501 `tts_unavailable`** when the model assets are missing (see
[Build / assets](../docs-dev/ARCHITECTURE.md#build--assets)).
## Voice speech β `POST /v1/voice/listen` (proprietary, Windows)
One-shot speech recognition from the **server microphone** through the
`AIOffice.VoiceAgent.Win.exe` subprocess β the same chain as the AIOffice Voice panel.
```bash
curl http://localhost:5290/v1/voice/listen \
-H "Content-Type: application/json" \
-d '{"lang":"it","timeout_seconds":15}'
# β {"text":"quanto fa sette per otto","lang":"it","provider":"voiceagent-win"}
```
- `lang`: two-letter ISO code (default `it`); `timeout_seconds`: 1β60 (default 15).
- **501 `voice_unavailable`** on non-Windows or when the executable is missing
(`Voice:ExePath`, default: next to the server). The microphone is exclusive β one listener
at a time. `408` on timeout.
Typical voice chat flow: `voice/listen` β transcript β `chat/completions` β `audio/speech` β
audio back to the client.
## Files β upload once, reference later (OpenAI Files API)
```bash
curl http://localhost:5290/v1/files -F "file=@report.csv" -F "purpose=assistants"
```
- Upload: original binary + server-side Markdown conversion (AllToMarkdown for documents,
Z.ai GLM-OCR for images). Response: OpenAI metadata + additive `extracted_content` /
`content_format`; `status` is `processed`/`unsupported`.
- `GET /v1/files/{id}/content` returns the original bytes; `DELETE /v1/files/{id}` removes the
file (`{"deleted": true}`, `404` when unknown). Chat references files via `file_ids`.
- Limits: 25 MB per upload; in-memory cache, lost on restart (volatile by design).
## `GET /v1/models` β agents **and** LLM providers
Two kinds of entries:
- **Agent sets** (`owned_by: "ai-orchestrator"`): select the agent tools via the chat `model` field.
- **LLM providers** (`owned_by: "llm-provider"`): the actual LLMs behind the agents, each with
its characteristics β `provider`, `model_name`, `protocol` (`OpenAI`/`Gemini`),
`context_window`, `base_address`, `interaction_mode` (`API` or `CLI` β the effective agent
interaction mode; see below). This is the "read the LLM characteristics" surface: a
client can pick a provider whose context window fits the task.
**Agent interaction mode.** Each provider drives the agent tools either through the JSON
tool-calling API (`interaction_mode: "API"` β one tool per method) or through the
application CLI (`interaction_mode: "CLI"` β the agent issues `ClassName subcommand args`
commands against the terminal). It is configured per provider in the Models & Providers UI
or in `providers.json` (`AgentInteractionMode`, options `API`/`CLI`/`Default`); `Default`
delegates to the model size β CLI for small models (context window < 128 000 tokens), API
for large ones. `interaction_mode` always reports the **effective** value (the explicit
setting or the size default). The same field appears on `GET /v1/control` session state.
`GET /v1/models/{id}` returns a single entry (`404` for unknown ids).
## Telegram chat β an in-process medium, **no HTTP endpoints**
The **Telegram chat medium** β a [WTelegramClient](https://github.com/wiz0u/WTelegramClient)
4.4.8 userbot that acts as a chat client (text + file attachments), **not** a voice medium
(the Telegram Client API has no audio-call support; see [docs/telegram.md](telegram.md)) β
is **purely in-process**: messages flow directly through the WTelegramClient library into
the agent harness, and configuration is driven from the TUI (`/telegram`), which calls
`TelegramBridge` directly. **There are no `/v1/telegram/*` HTTP endpoints** β Telegram is
not a web client, so nothing about it is exposed over HTTP.
Configuration lives in `telegram.json` next to the executable (excluded from updates) β set
it by hand, with the setup scripts (`scripts/setup-telegram.bat` on Windows,
`scripts/setup-telegram.sh` on Linux/macOS), or from the TUI `/telegram` command. Config
keys (case-insensitive): `Enabled`, `ApiId`, `ApiHash`, `PhoneNumber`, `SessionPath`,
`AllowedUsers` (comma-separated list of ids / `@usernames`), `Agent`. Changing a
**connection-affecting key** (`Enabled`, `ApiId`, `ApiHash`, `PhoneNumber`, `SessionPath`)
restarts the bridge from the TUI.
## `GET /v1/control` β capabilities
Without a session id it returns what this platform can do right now:
```json
{
"capabilities": {
"platform": "windows",
"default_provider": "DeepSeekBridge",
"providers": [ { "name": "Zai", "model_name": "glm-4.7-flash", "protocol": "OpenAI", "context_window": 128000, "base_address": "https://api.z.ai/", "interaction_mode": "API" }, ... ],
"tts": { "available": true, "engine": "kokoro", "voices": [ ... ], "detail": "" },
"voice": { "available": true, "engine": "voiceagent-win", "detail": "" },
"telegram": { "available": true, "connected": true, "status": { "phase": "connected", ... } },
"sessions": 3
}
}
```
---
See also: [README](../README.md) Β· [Terminal UI](TUI.md) Β· [Architecture](../docs-dev/ARCHITECTURE.md) (developers, not shipped)
|