23 KiB
@elizaos/plugin-elizacloud
Eliza Cloud integration — multi-model inference, container provisioning, agent bridge, and billing for elizaOS agents.
Purpose / role
Connects an Eliza agent to Eliza Cloud for hosted AI inference (text, embeddings, TTS, STT, image), container lifecycle management, real-time agent bridging via WebSocket, and billing/credit flows. Auto-enables when ELIZAOS_CLOUD_API_KEY or ELIZAOS_CLOUD_ENABLED=true is present (see auto-enable.ts). This plugin has priority 50, which means it wins the default text-generation slot over other direct provider plugins (priority 0) when no explicit routing preference is configured — unless the host writes ELIZAOS_CLOUD_USE_INFERENCE=false (applyCloudConfigToEnv), in which case the chat-brain handlers (TEXT_*, RESPONSE_HANDLER, ACTION_PLANNER) are not registered at all and only the capability handlers (IMAGE, IMAGE_DESCRIPTION, TEXT_TO_SPEECH, TRANSCRIPTION, embeddings, RESEARCH) stay active. This capability-only mode is how an agent keeps Cloud image/media/TTS while an external provider (a CLI/SDK subscription brain, a local model) owns the text brain (elizaOS/eliza#10819).
The plugin has two distinct export surfaces:
elizaOSCloudPlugin(src/index.ts) — inference model handlers, cloud providers, and cloud services. Safe in both browser and Node.elizaCloudRoutePlugin(src/plugin.ts) — registers/api/cloud/*HTTP routes. Node-only; loaded lazily viasrc/register-routes.ts.
Plugin surface
Model handlers
Two registration groups (elizaOS/eliza#10819):
Capability handlers — always registered (static models map). These don't
compete with the chat brain and must survive an external text provider:
| Slot | Handler | File |
|---|---|---|
TEXT_EMBEDDING |
handleTextEmbedding |
src/models/embeddings.ts |
RESEARCH |
handleResearch |
src/models/research.ts |
IMAGE |
handleImageGeneration |
src/models/image.ts |
IMAGE_DESCRIPTION |
handleImageDescription |
src/models/image.ts |
TEXT_TO_SPEECH |
handleTextToSpeech |
src/models/speech.ts |
TRANSCRIPTION |
handleTranscription |
src/models/transcription.ts |
Chat-brain handlers — registered from init() (registerTextInferenceModels,
src/index.ts), skipped when the host writes ELIZAOS_CLOUD_USE_INFERENCE=false
(registered when it is true or unset — unset preserves standalone plugin use):
| Slot | Handler | File |
|---|---|---|
TEXT_NANO |
handleTextNano |
src/models/text.ts |
TEXT_SMALL |
handleTextSmall |
src/models/text.ts |
TEXT_MEDIUM |
handleTextMedium |
src/models/text.ts |
TEXT_LARGE |
handleTextLarge |
src/models/text.ts |
TEXT_MEGA |
handleTextMega |
src/models/text.ts |
RESPONSE_HANDLER |
handleResponseHandler |
src/models/text.ts |
ACTION_PLANNER |
handleActionPlanner |
src/models/text.ts |
Providers
| Name | File | Description |
|---|---|---|
elizacloud_status |
src/cloud-providers/cloud-status.ts |
Container and connection status (position 90, contexts: settings/finance) |
elizacloud_credits |
src/cloud-providers/credit-balance.ts |
Credit balance with 60 s cache, low/critical alerts (position 91) |
elizacloud_health |
src/cloud-providers/container-health.ts |
Container health approximation from cached state; private (position 92) |
elizacloud_models |
src/cloud-providers/model-registry.ts |
Available models grouped by provider, 5 min cache (position 92) |
Services (started in dependency order)
| Service type | Class | File | Description |
|---|---|---|---|
CLOUD_AUTH |
CloudAuthService |
src/services/cloud-auth.ts |
Auth entry points — device auto-signup and Cloud SSO OAuth flow |
CLOUD_BOOTSTRAP |
CloudBootstrapServiceImpl |
src/services/cloud-bootstrap.ts |
Exposes Cloud trust-anchor (JWKS URL, issuer, container id) without importing app-core |
CLOUD_MANAGED_GATEWAY_RELAY |
CloudManagedGatewayRelayService |
src/services/cloud-managed-gateway-relay.ts |
Long-poll relay enabling Cloud to push requests to a local agent |
CLOUD_MODEL_REGISTRY |
CloudModelRegistryService |
src/services/cloud-model-registry.ts |
Fetches and caches available models from Cloud (30 min TTL) |
CLOUD_CONTAINER |
CloudContainerService |
src/services/cloud-container.ts |
ECS container lifecycle: create, list, poll status, delete |
CLOUD_BRIDGE |
CloudBridgeService |
src/services/cloud-bridge.ts |
JSON-RPC 2.0 WebSocket bridge to cloud-hosted agents with exponential-backoff reconnect |
CLOUD_BACKUP |
CloudBackupService |
src/services/cloud-backup.ts |
Agent state snapshots/restore; periodic auto-backup and pre-eviction snapshots |
workflow_credential_provider |
CloudCredentialProvider |
src/services/cloud-credential-provider.ts |
Bridges plugin-workflow's credential slot to Cloud OAuth connector surface |
Events
| Event | Handler | File |
|---|---|---|
MODEL_USED |
createWaifuMeteringHandler() |
src/utils/waifu-metering.ts |
Forwards per-inference token and USD spend to the Cloud metering endpoint when the container is a hosted agent. Inactive otherwise.
Routes (via elizaCloudRoutePlugin)
All paths use rawPath: true. Handled by three route groups:
- Status (
handleCloudStatusRoutes):GET /api/cloud/status,GET /api/cloud/credits - Cloud routes (
handleCloudRoute): login, disconnect, relay-status, agents provisioning/connect/shutdown, coding-container create/sync/promotions - Billing proxy (
handleCloudBillingRoute):GET|POST|PUT|PATCH|DELETE /api/cloud/billing/:path*— forwards to authenticated Cloud API
Layout
plugins/plugin-elizacloud/
src/
index.ts Main plugin object (elizaOSCloudPlugin)
index.node.ts Node-specific re-exports
index.browser.ts Browser-compatible build entry
plugin.ts Route-only plugin (elizaCloudRoutePlugin)
register-routes.ts Lazy-loads route plugin via registerAppRoutePluginLoader
init.ts OpenAI-compatible client initialization
auto-enable.ts Auto-enable check (reads ELIZAOS_CLOUD_API_KEY / ELIZAOS_CLOUD_ENABLED)
cloud-setup.ts Interactive Cloud setup flow (runCloudSetup)
cloud-voice-catalog.ts Fetches available TTS voice catalog from Cloud
models/
text.ts Text generation handlers for all model tiers
embeddings.ts TEXT_EMBEDDING handler
image.ts IMAGE and IMAGE_DESCRIPTION handlers
speech.ts TEXT_TO_SPEECH handler + CloudTtsUnavailableError
research.ts RESEARCH handler
transcription.ts TRANSCRIPTION handler
tokenization.ts TEXT_TOKENIZER_ENCODE/DECODE handlers
index.ts Re-exports all model handlers
services/
cloud-auth.ts CloudAuthService (CLOUD_AUTH)
cloud-bootstrap.ts CloudBootstrapServiceImpl (CLOUD_BOOTSTRAP)
cloud-managed-gateway-relay.ts CloudManagedGatewayRelayService (CLOUD_MANAGED_GATEWAY_RELAY)
cloud-model-registry.ts CloudModelRegistryService (CLOUD_MODEL_REGISTRY)
cloud-container.ts CloudContainerService (CLOUD_CONTAINER)
cloud-bridge.ts CloudBridgeService (CLOUD_BRIDGE)
cloud-backup.ts CloudBackupService (CLOUD_BACKUP)
cloud-credential-provider.ts CloudCredentialProvider (workflow_credential_provider)
cloud-providers/
cloud-status.ts elizacloud_status provider
credit-balance.ts elizacloud_credits provider
container-health.ts elizacloud_health provider
model-registry.ts elizacloud_models provider
routes/
cloud-routes.ts Core cloud login/disconnect/agent routes
cloud-routes-autonomous.ts Autonomous-mode cloud route handler
cloud-status-routes.ts /api/cloud/status and /api/cloud/credits
cloud-status-routes-autonomous.ts Autonomous-mode status route handler
cloud-billing-routes.ts /api/cloud/billing/* proxy
cloud-relay-routes.ts Relay-status route
cloud-provisioning.ts isCloudProvisionedContainer helper
cloud-coding-container-routes.ts Coding-container management
cloud-compat-routes.ts Compat route shims
cloud-features-routes.ts Feature-flag routes
travel-provider-relay-routes.ts Travel provider relay routes
home-remote-runner-access-url.ts Remote runner access URL helper
cloud/
auth.ts Auth helpers
auth-service-types.ts CloudAuthApiKeyService interface, normalizeCloudApiKey, isCloudAuthApiKeyService
backup.ts Backup helpers
base-url.ts resolveCloudApiBaseUrl, normalizeCloudSiteUrl
bridge-client.ts ElizaCloudClient, CloudWalletDescriptor
cloud-api-key.ts resolveCloudApiKey, resolveCloudApiBaseUrl, normalizeCloudSecret
cloud-manager.ts CloudManager orchestrator
cloud-proxy.ts Proxy utilities
cloud-wallet.ts Wallet descriptor types
clack-observer.ts ClackObserver (interactive CLI setup feedback)
null-observer.ts NullCloudSetupObserver
setup-observer.ts CloudSetupObserver interface
reconnect.ts Reconnect helpers
validate-url.ts validateCloudBaseUrl
duffel-client.ts Duffel travel/flight booking client (searchFlights, createOrder, readDuffelConfigFromEnv, DuffelConfigError)
lifeops-schedule-sync-client.ts LifeOps schedule sync client (resolveLifeOpsScheduleSyncConfig)
lifeops-schedule-sync-contracts.ts LifeOps schedule sync contract types
managed-payment-clients.ts Managed payment client helpers
x402-payment-handler.ts x402 payment protocol handler (parseX402Response, requestPayment, PaymentRequiredError)
index.ts Barrel
lib/
cloud-connection.ts CloudAuthLike interface
cloud-secrets.ts getCloudSecret, clearCloudSecrets, scrubCloudSecretsFromEnv
config-env.ts Env-to-config mapping
config-like.ts ElizaConfig type
credential-type-map.ts credTypeToConnector mapping
feature-flags.ts Feature flag helpers
http.ts sendJson HTTP helper
server-cloud-tts.ts TTS compat layer, resolveCloudTtsBaseUrl
state-paths.ts State directory path helpers
tts-debug.ts TTS debug utilities
providers/
openai.ts createOpenAIClient (Vercel AI SDK OpenAI-compatible adapter)
utils/
cloud-api.ts CloudApiClient — base HTTP client for Cloud API
sdk-client.ts createCloudApiClient, createElizaCloudClient
waifu-metering.ts createWaifuMeteringHandler (MODEL_USED event bridge)
config.ts Model string resolution helpers (getNanoModel, getLargeModel, …)
events.ts emitModelUsageEvent, ModelUsageEventMeta
helpers.ts Misc internal helpers
responses-output.ts extractResponsesOutputText (Responses API output parser)
cloud/sdk/ Internal SDK surface wrappers
types/
cloud.ts CloudContainer, DevicePlatform, DEFAULT_CLOUD_CONFIG, and all Cloud API types
index.ts Type barrel
__tests__/
unit/ Unit tests (no live API)
integration/ Integration tests
*.test.ts Feature-level test suites
auto-enable.ts Auto-enable entry point (package.json elizaos.plugin.autoEnableModule)
build.ts Dual-target (node + browser) build script
package.json
Commands
bun run --cwd plugins/plugin-elizacloud build # compile node + browser bundles
bun run --cwd plugins/plugin-elizacloud typecheck # type check only (tsgo --noEmit)
bun run --cwd plugins/plugin-elizacloud test # run all tests via vitest
bun run --cwd plugins/plugin-elizacloud test:unit # unit tests only
bun run --cwd plugins/plugin-elizacloud test:integration # integration tests only
bun run --cwd plugins/plugin-elizacloud test:e2e # live smoke test via app-core script
bun run --cwd plugins/plugin-elizacloud lint # biome check --write --unsafe
bun run --cwd plugins/plugin-elizacloud clean # rm -rf dist .turbo .turbo-tsconfig.json tsconfig.tsbuildinfo
Config / env vars
All settings are optional except ELIZAOS_CLOUD_API_KEY (required for any authenticated call).
Required
| Var | Description |
|---|---|
ELIZAOS_CLOUD_API_KEY |
API key (eliza_xxxxx). Get from https://www.elizacloud.ai/dashboard/api-keys |
Optional — core
| Var | Default |
|---|---|
ELIZAOS_CLOUD_BASE_URL |
https://elizacloud.ai/api/v1 |
ELIZAOS_CLOUD_ENABLED |
false — when true, enables container provisioning, device auth, bridge, and backup services |
ELIZAOS_CLOUD_EXPERIMENTAL_TELEMETRY |
false |
ELIZAOS_CLOUD_APP_VERSION |
2.0.0-beta.0 |
ELIZAOS_CLOUD_NATIVE_CONCURRENCY |
8 — per-process cap on CONCURRENT native cloud text calls. Covers BOTH native text routes sharing the one cerebras key: /chat/completions (native-transport callers) AND /responses (bare-{ prompt } callers, incl. the primary reply action). The per-turn burst comes from the prompt batcher (dynamicPromptExecFromState, always sets providerOptions -> /chat/completions) and the merged evaluator call — NOT from composeState providers (no provider calls useModel during composeState); firing them at once can overrun the shared cerebras key's concurrent limit -> 429 -> retries. The default 8 is a SAFETY CEILING, not full serialization: with the paid cerebras key (1000 req/min) and leaner per-turn call counts the typical 1-3 concurrent calls/turn run unguarded while a pathological burst is still bounded. The limiter is process-global and keys on native transport (not the model), so it also bounds non-cerebras native calls (e.g. zai-glm-4.7) — hence the high default. Set to 1 to fully serialize on a cerebras-bottlenecked single-key deployment, or raise for more parallelism. Embeddings use a separate /embeddings route and are NOT gated. |
Optional — model tiers (each has a fallback bare env alias)
| Cloud var | Bare fallback | Default |
|---|---|---|
ELIZAOS_CLOUD_NANO_MODEL |
NANO_MODEL |
falls back to small model |
ELIZAOS_CLOUD_SMALL_MODEL |
SMALL_MODEL |
gemma-4-31b |
ELIZAOS_CLOUD_MEDIUM_MODEL |
MEDIUM_MODEL |
falls back to small model |
ELIZAOS_CLOUD_LARGE_MODEL |
LARGE_MODEL |
gemma-4-31b |
ELIZAOS_CLOUD_MEGA_MODEL |
MEGA_MODEL |
falls back to large |
ELIZAOS_CLOUD_RESPONSE_HANDLER_MODEL |
RESPONSE_HANDLER_MODEL |
falls back to small model |
ELIZAOS_CLOUD_ACTION_PLANNER_MODEL |
ACTION_PLANNER_MODEL |
falls back to large model |
ELIZAOS_CLOUD_RESEARCH_MODEL |
RESEARCH_MODEL |
o3-deep-research |
Optional — embeddings
| Var | Default |
|---|---|
ELIZAOS_CLOUD_EMBEDDING_MODEL |
text-embedding-3-small |
ELIZAOS_CLOUD_EMBEDDING_DIMENSIONS |
1536 |
ELIZAOS_CLOUD_EMBEDDING_URL |
unset (uses base URL) |
ELIZAOS_CLOUD_EMBEDDING_API_KEY |
falls back to ELIZAOS_CLOUD_API_KEY |
Optional — image / audio
| Var | Default |
|---|---|
ELIZAOS_CLOUD_IMAGE_DESCRIPTION_MODEL |
gpt-5.4-mini |
ELIZAOS_CLOUD_IMAGE_DESCRIPTION_MAX_TOKENS |
8192 |
ELIZAOS_CLOUD_IMAGE_GENERATION_MODEL |
google/nano-banana-2/text-to-image |
ELIZAOS_CLOUD_TTS_MODEL |
gpt-5-mini-tts |
ELIZAOS_CLOUD_USE_STT |
unset — per-service opt-in for Cloud STT in capability-only mode (ELIZAOS_CLOUD_ENABLED unset) |
ELIZAOS_CLOUD_STT_TIMEOUT_MS |
60000 |
Browser-only proxy vars (no secrets in client bundles)
| Var |
|---|
ELIZAOS_CLOUD_BROWSER_BASE_URL |
ELIZAOS_CLOUD_BROWSER_EMBEDDING_URL |
How to extend
Add a model handler
- Add a handler function in the appropriate file under
src/models/. - Export it from
src/models/index.ts. - Register it in the
modelsmap insrc/index.tskeyed by theModelTypeconstant. - Use
createCloudApiClient(runtime)for raw API-base calls (seesrc/utils/sdk-client.ts). - Call
emitModelUsageEvent(runtime, type, prompt, usage, meta)after each inference call.
Add a provider
- Create a new file in
src/cloud-providers/exporting aProviderobject. - Export it from
src/cloud-providers/index.ts. - Add it to the
providersarray insrc/index.ts. - Gate it with
contextGateso it only fires in relevant context windows.
Add a service
- Create a new file in
src/services/with a class extendingServicefrom@elizaos/core. - Set a static
serviceTypestring (used forruntime.getService(...)lookups). - Add it to the
servicesarray insrc/index.tsin dependency order. - Add a matching
await runtime.getService(YourService.serviceType)?.stop()line indispose().
Conventions / gotchas
- No direct
fetch()for Cloud API calls. UsecreateCloudApiClient(runtime)orcreateElizaCloudClient(runtime)fromsrc/utils/sdk-client.ts. The one exception is the plugin test suite downloading a public audio fixture. ELIZAOS_CLOUD_ENABLEDgates infrastructure services. When false, only inference model handlers are active. Container, bridge, backup, and relay services start only when this flag is true.- Browser build is separate.
src/index.browser.tsis the entry fordist/browser/. It must not import Node-only modules. The route plugin (src/plugin.ts) is Node-only and is excluded from the browser bundle. - Routes use
rawPath: true. All/api/cloud/*routes bypass the plugin-name prefix so paths stay stable. - TTS routing precedence. This plugin's priority (50) does not govern TTS routing. The router-handler in
plugin-local-inferenceruns atMAX_SAFE_INTEGERpriority and enforces theprefer-localpolicy. Cloud TTS is a fallback;CloudTtsUnavailableError(fromsrc/models/speech.ts) signals the router to try the next provider. - Cloud STT gate mirrors the TTS gate.
handleTranscriptionserves when a Cloud API key is present AND (ELIZAOS_CLOUD_ENABLEDORELIZAOS_CLOUD_USE_STT) is truthy —isCloudSttAvailableinsrc/utils/config.ts. Otherwise it throwsCloudSttUnavailableErrorso the local-inference router falls through to the next TRANSCRIPTION provider.audioUrl/string inputs are fetched through core'sfetchWithSsrfGuard. - Cloud TTS availability gate ≠ core
isCloudConnected.handleTextToSpeechandfetchCloudVoiceCatalogserve when a Cloud API key is present AND (ELIZAOS_CLOUD_ENABLEDORELIZAOS_CLOUD_USE_TTS) is truthy —isCloudTtsAvailableinsrc/utils/config.ts. TheUSE_TTSleg is what keeps Cloud TTS alive in capability-only mode, whereapplyCloudConfigToEnvdeliberately leavesELIZAOS_CLOUD_ENABLEDunset (many consumers read ENABLED as "cloud is the text brain"). Do not "simplify" this back to coreisCloudConnected— that regates TTS on inference and reopens the capability-only gap (elizaOS/eliza#10961 follow-up). - Services start in dependency order.
CloudAuthServicemust be first; every other service callsruntime.getService("CLOUD_AUTH").dispose()stops them in reverse order. CloudBootstrapServicefails closed.getExpectedIssuer()throws whenELIZA_CLOUD_ISSUERis unset. Never add a silent default.- No
ascasts or?? 0fallbacks for missing pipeline data. Follow the architecture rules inAGENTS.mdat the repo root.
⛔ NON-NEGOTIABLE — evidence, trajectories & real end-to-end tests
The binding, repo-wide standard is AGENTS.md. Read it. Nothing in this package is done until it is proven done — a reviewer must confirm it works without reading the code, from the artifacts you attach. This applies to every feature, fix, refactor, and chore here. "Tests pass" is not proof; "CI is green" is not proof.
- Record AND read model trajectories. Capture the actual inputs and outputs of the model
from a live LLM — not the deterministic proxy, not a mock: the prompt, the
providers/context, the raw model output, every tool/action call, and the result. Then open
the trajectory and review it by hand. A captured-but-unread trajectory is not evidence
(
packages/scenario-runner/bin/eliza-scenarios run <scenario> --report <out>). - Real, full-featured E2E — no larp. Every feature ships detailed end-to-end tests that drive the real path end to end. Not the happy "front door" only: cover error paths, edge/empty/invalid input, concurrency, roles/permissions, and adversarial input. A test that asserts against a mock/stub/fixture standing in for the thing under test does not count. If the real model/device/chain/connector/account is hard to reach, make it reachable — that is the work, not an excuse to mock. If the existing tests here are shallow or mocked, fixing them is part of your change.
- Screenshots + logs at every phase, plus a complete walkthrough video/run-through of
the entire feature or view, start to finish (
bun run test:e2e:record). - Manually review every artifact the change touches — never just the green check: client
logs (console + network), server logs (
[ClassName] …), the model trajectories in and out, before/after full-page screenshots, and the domain artifacts listed below for this package. - No residuals. No shortcuts. The goal is not "done" — it is everything done. Clear every blocker by the hard path: build the real architecture, stand up the real model/device/service, actually test it. Never leave a TODO, a stub, a stepping-stone, or a "follow-up." When unsure, research thoroughly, weigh the options, and ship the best, highest-effort, production-ready version. Keep going until every possibility is exhausted.
Artifacts → attached inline in the PR (MP4 video, JPG screenshots, logs in <details>); attach each evidence type or
explicitly mark it N/A with a reason — never leave it blank. If develop moved and changed
behavior, re-capture evidence; stale proof is worse than none.
Capture & manually review for this package — model provider:
- A trajectory from a live call to this provider (not the proxy, not a mock): full request, raw response, token usage, finish reason, and streamed chunks.
- Proof of tool/function-calling and structured-output parsing against the real model.
- The error paths exercised: bad key, model-not-found, oversized context, timeout, rate-limit, mid-stream disconnect — plus latency and cost from the real call.
- If no key is available in CI, attach the documented live-run transcript as evidence — never a mocked client passed off as a pass.