* chore: ignore local agency and team snapshots
* feat: add team agent profile templates
* KANBAN
* template
* prototype
* Fix tenant runtime pool deployment
* Fix local team profiles runtime override
* Update local team profiles testing guide
* Adapt local team profiles test deployment
* Patch tenant NFS server to ClusterIP
* Use NFS export root for workspace mounts
* Tighten team task completion projection
* Fix team event body map comparison
* before modification
* temp_peer1
* leadermodel
* team_fix
* fix(team): harden workflow completion and workspace handling
* fix(team): harden redis workflow protocol
* discord link updated
* chore: checkpoint team workflow fixes
* fix: stabilize leader kanban ledger
* fix: coalesce leader mediated work items
* fix: confirm leader-mediated member results
* fix: harden leader mediated team state
* fix: stabilize team shared runtime state
* Fix team task reconciliation and create flow
* fix: stabilize leader mediated team orchestration
* fix: refine team workspace presentation
* checkpoint: persist team workflow ledger fixes
* fix: reconcile team workflow delivery state
* fix: preserve meaningful team chat events
* docs: add team workspace quick guides
* chore: ignore local team deployment files
* fix: preserve setgid on team shared directories
* fix(team): align workspace overlays and completion projection
* fix(team): preserve workflow identity recovery
* fix(team): protect terminal workflow projection
* fix(team): preserve collaboration flow after terminal results
* feat(skill-hub): add Skill Hub catalog, publish flow, and lite materialize pipeline
Introduce Skill Hub for browsing, importing, publishing, and installing skills, with lite instance package materialization and runtime sync support.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(skill-hub): remove token-governance hooks from Skill Hub PR
Strip validateManagedRuntimeEnvironmentOverrides, network lock policy sync,
and egress proxy audit wiring that belong to the upcoming token-usage work,
so the Skill Hub branch compiles independently.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(skill-hub): update migration number in materialize docs
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(skill-hub): make RuntimeAgentClient test stub and hub tests compile-safe
Add ResyncInstanceSkills to the runtime pool handler fake client, and harden
skill hub payload helpers/tests against nil storage/instance repos so go test passes.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat: add Skill Hub hardening with session usage tracking and egress governance
Unify Skill Hub runtime sync improvements with session-token observability,
egress network policy, and admin/instance usage reporting for reopenable PR.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(skill-hub): repair CI tests and nested skill install
* fix(ci): restore release deployment configuration
---------
Co-authored-by: heshengran <heshengran@ieisystem.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(runtime): improve Lite gateway scheduling and proxy stability
- Expand the runtime gateway port range to match 100 instances per pod
- Scale runtime deployments based on pending backlog
- Serialize gateway creation per pod and handle starting/creating states
- Clean up missing gateway bindings from runtime agent reports
- Improve Lite gateway proxy token stripping and auth header injection
- Update deployment manifests, docs, and related tests
* fix(runtime): harden Hermes Lite gateway recovery
- Add hashed gateway token aliases for short-term token compatibility
- Fix /resume conversations failing when old sessions use mismatched gateway tokens
- Recover ports only from stale error bindings without affecting active gateways
* feat(runtime): enable bidirectional desktop clipboard and direct proxy
- Enable WebTop/Selkies/Kasm clipboard by default with per-instance overrides
- Document clipboard verification, configuration pitfalls, and IME troubleshooting
- Enable direct desktop proxy in the K8s and K3s manifests
---------
Co-authored-by: litiantian03 <litiantian03@ieisystem.com>
* chore: ignore local agency and team snapshots
* feat: add team agent profile templates
* KANBAN
* template
* prototype
* Fix tenant runtime pool deployment
* Fix local team profiles runtime override
* Update local team profiles testing guide
* Adapt local team profiles test deployment
* Patch tenant NFS server to ClusterIP
* Use NFS export root for workspace mounts
* Tighten team task completion projection
* Fix team event body map comparison
* before modification
* temp_peer1
* leadermodel
* team_fix
* fix(team): harden workflow completion and workspace handling
* fix(team): harden redis workflow protocol
* discord link updated
* chore: checkpoint team workflow fixes
* fix: stabilize leader kanban ledger
* fix: coalesce leader mediated work items
* fix: confirm leader-mediated member results
* fix: harden leader mediated team state
* fix: stabilize team shared runtime state
* Fix team task reconciliation and create flow
* fix: stabilize leader mediated team orchestration
* fix: refine team workspace presentation
* checkpoint: persist team workflow ledger fixes
* fix: reconcile team workflow delivery state
* fix: preserve meaningful team chat events
* docs: add team workspace quick guides
* chore: ignore local team deployment files
* fix: preserve setgid on team shared directories
* chore: ignore local agency and team snapshots
* feat: add team agent profile templates
* KANBAN
* template
* prototype
* Fix tenant runtime pool deployment
* Fix local team profiles runtime override
* Update local team profiles testing guide
* Adapt local team profiles test deployment
* Patch tenant NFS server to ClusterIP
* Use NFS export root for workspace mounts
* Tighten team task completion projection
* Fix team event body map comparison
* before modification
* temp_peer1
* leadermodel
* team_fix
* fix(team): harden workflow completion and workspace handling
* fix(team): harden redis workflow protocol
- Add APIs for batch creating and deleting Lite instances
- Support Lite batch create, selection, and delete flows in the instance list
- Fix Lite OpenClaw/Hermes workspace initial paths and skill directory paths
- Hide transient skill temp directories from workspace browsing
- Preserve runtime query tokens when stripping ClawManager desktop access tokens
- Add an available-skills endpoint for instances with actor-aware filtering
- Persist gateway token, agent bootstrap token, and OpenClaw config snapshots for Lite instances
- Inject agent and OpenClaw bootstrap env into gateway runtimes
- Materialize and remove Lite instance skills in the workspace filesystem
- Improve instance detail skill upload, bulk attachment, workspace refresh, and layout syncing
- Support multi-select skill attachment and show risk-policy filtering hints
- Mark instance skills as removed after successful uninstall commands
- Sync skill removal state when related workspace files are deleted
- Hide removed skills and avoid reactivating them from agent inventory reports
- Clear proxy-scoped OpenClaw control storage to prevent cross-instance state leakage
- Reuse desktop direct upstream resolution and include upstream in short external access tokens
- Default and normalize Selkies stream settings for desktop runtimes
- Remove legacy deployments with the same instance-id before ensuring the current deployment
- Preserve the unsaved desktop stream profile state in the frontend
- Add tests for stream profiles, pod envs, and deployment cleanup
- Add leader election to avoid duplicate sync and Team event consumers
- Proxy desktop traffic directly from nginx to instance Services to reduce load during concurrent viewers
- Add desktop stream profiles with low / standard / high options
- Limit hostPath PV fallback to manual storageClass only
- Disable Webtop/KasmVNC bidirectional clipboard sync by default to reduce accidental paste during remote desktop input
- Derive desktop runtime /dev/shm defaults from instance memory: 1G for 4G, 2G for 8G, and 4G for 12G+
- Preserve explicit SHM_SIZE_GB overrides and keep shell runtime behavior unchanged
- Add tests for openclaw/hermes clipboard policy and /dev/shm default selection
- Support multiple OpenClaw channel resources and fix alias injection
Add cursor-based GET /teams/:id/tasks and GET /teams/:id/events APIs.
Team detail page loads older messages on scroll/click and deduplicates
collaboration chat messages with improved thread ordering.
- Add migrations plus Team/Member/Task/Event APIs; Redis Streams consumer projects inbox/events into DB as source of truth
- Inject Team Secret (Redis URL, team token) via envFrom; shared RWX PVC at /team; sync ConfigMap roster to /team/team.json
- Create member Pods through InstanceService; extend K8s (PVC/Secret/Pod/ConfigMap); stale-task sweep and Team/member deletion with cleanup
- Add /teams, /teams/new, /teams/🆔 creation wizard (roster, presets, shared env/OpenClaw plan), per-member desktops, collaboration timeline, debug dispatch (defaults to Leader when target omitted)
Every channel form editor (telegram, dingtalk-connector, slack, feishu)
silently drops fields on Save. The user flow that loses data:
1. Open a channel resource in the JSON tab, add any field not owned
by the form (e.g. webhook, custom capabilities, additional feishu
accounts).
2. Switch to the Form tab and change anything (e.g. rotate a secret).
3. Click Save.
4. The field from step 1 is gone from the saved config — and stays
gone after a reload.
The data loss occurs in two places; both must be fixed:
Frontend (OpenClawConfigCenterPage.tsx):
Each update*ChannelContentText rebuilt the output from a hard-coded
allowlist instead of merging form fields into the parsed config. Any
field not in the allowlist was dropped on Save. For feishu this
collapsed a multi-account config to a single `main` account.
Backend (openclaw_config_service.go):
normalizeOpenClawResourceContent calls normalize*ChannelConfigForEnv
on every create/update to validate and rebuild the stored config.
Those helpers use the same allowlist pattern, so even if the
frontend sends a complete config, the backend strips tenant-authored
keys before persisting.
Fix:
- Frontend: introduce a single mergeChannelConfig(existing, allowlisted)
helper and route every editor's output through it. Each handler
builds an allowlisted object containing only its known keys; the
helper overlays it onto the parsed existing config so unknown keys
survive. The merge point is now structural rather than a per-handler
convention, so a new editor (e.g. an MS Teams handler added in a
downstream patch) can't regress by forgetting the spread. Feishu's
accounts deep-merge is preserved by building the merged accounts
inside the allowlisted object before handing off to the helper.
- Backend: introduce mergeOpenClawChannelConfigForStorage, which
overlays the allowlist-normalized output onto the original parsed
payload at storage time. Normalized keys still win; unknown keys
pass through. For feishu the accounts map is merged member-wise.
Env-rendering semantics are unchanged. appendCompiledOpenClawEnvPayload
continues to call the *ForEnv helpers at render time against the
stored content, so the runtime env payload sees the same allowlist as
before.
Tests:
- Two existing ResourcePayloadFromModelNormalizesStored*JSON tests
asserted the pre-fix behavior (dropping `legacyField`) and are
updated to assert it survives.
- New TestNormalizeOpenClawResourceContentPreservesUnknownChannelFields
(table-driven, three editors) guards the preservation invariant.
- New TestNormalizeOpenClawResourceContentPreservesFeishuSiblingAccounts
guards the feishu deep-merge.
Manual verification against a live instance: add `webhook` via the
JSON tab, switch to Form tab, rotate a secret, Save, reload — the
webhook field is retained.
Summary
-------
Logging in as an admin and navigating to Workspace → My Instances
returned every instance in the cluster instead of the admin's own.
Role was overloaded to both widen the admin console surface AND
widen the self-scoped list endpoint, so the workspace view broke
the owner-isolation contract users expect.
Root cause
----------
`InstanceService.GetVisibleInstances(userID, role, ...)` branched
on role: admin callers fell through to `instanceRepo.GetAll`,
non-admin callers fell through to `GetByUserID`. The single
`GET /instances` handler passed the caller's role into that
function, so the same URL meant "my instances" for users and
"every instance in the system" for admins.
Fix
---
Split the two views at the API surface, not inside the service:
- `GET /instances` is now always caller-scoped. The handler calls
`GetByUserID` unconditionally and never reads the caller's role.
- `GET /admin/instances` is a new route, gated by the existing
admin middleware trio (`Auth` + `SetUserInfo` + `NewAdminAuth`).
It calls a new `InstanceService.GetAllInstances` that does not
look at any userID.
- `GetVisibleInstances` is removed; nothing else in the codebase
called it.
Frontend follows: `AdminDashboard` and `InstanceManagementPage`
switch to a new `adminInstanceService.getInstances()` that hits
`/admin/instances`. Per-instance admin actions (start, stop,
delete, proxy, etc.) are intentionally out of scope for this PR
and continue to use the existing endpoints.
Tests
-----
Two new service tests in `instance_visibility_test.go`:
- `TestGetByUserIDFiltersByCaller` pins the workspace contract:
regardless of how many admins or other users exist, a caller
only ever sees their own instances through this path.
- `TestGetAllInstancesReturnsEveryUser` covers the admin-console
path, including pagination.
`go build ./...`, `go test ./...`, and `npm run build` all pass.
Scope note
----------
This PR fixes the listing leak only. Per-instance handlers in
`instance_handler.go` retain their existing inline role checks
and will be audited separately.
The previous error handler fell back to the button label when the server
did not return a JSON body (e.g. nginx 413 plain-HTML response), producing
a popup that said "Import .openclaw" on failure - indistinguishable from
a success message.
Add describeOpenClawError() that:
- surfaces server-supplied JSON { error: "..." } when present,
- strips HTML from plain-text bodies (nginx default error page),
- recognizes 413 as "archive too large",
- otherwise reports the HTTP status code.
Also emit distinct i18n strings for success (importOpenClawSuccess) and
failure (importOpenClawFailed / exportOpenClawFailed) across all 5
supported languages (en, zh, ja, ko, de).
- Change cpu_cores from int to float64 in backend models, handlers, services
- Update frontend CreateInstancePage to accept decimal CPU values (step 0.1)
- Add DB migration 009_cpu_cores_decimal.sql for DECIMAL(10,2) columns
- Update deployment manifests (k8s, k3s, incluster) schema definitions
- Add GOFLAGS/GOPROXY/GOSUMDB build args to Dockerfile so that
--build-arg values passed by build scripts take effect during go mod download
Add a protocol selector for local/internal model templates, route gateway requests through Anthropic native APIs when selected, and update the Anthropic template icon.
Background
This update focuses on three areas:
- Improve model onboarding capabilities by supporting unified configuration across multiple model service platforms
- Optimize the AI audit detail page to improve trace troubleshooting and readability
- Upgrade the OpenClaw runtime image and automatically enable the Gateway after instance startup
Changes
1. Support multiple model service platforms
- Introduce a vendor template mechanism for model onboarding, allowing quick selection of different model service platforms through a searchable dropdown
- Preconfigure fixed `base_url` values for common platforms to reduce manual input and configuration errors
- Support custom `base_url` for `Local / Internal`, compatible with local gateways, internal proxies, and self-hosted compatible services
- Add vendor icons to improve recognizability on the model configuration page
- Optimize the model creation experience so newly created cards appear at the top, making them easier to find when many models exist
- Optimize provider model discovery logic to support OpenAI-compatible services with non-standard version paths
2. Optimize the AI audit detail page
- Rework the trace detail layout to improve information hierarchy and readability
- Change audit timestamps to a combined relative + absolute format for more intuitive troubleshooting
- Add an execution flow view so the full process can be inspected by execution node
- Add a minimap to quickly locate execution nodes and keep navigation aligned with detail scrolling
- Remove duplicated or low-value information blocks to simplify the trace detail experience
- Optimize status rendering and failure reason unwrapping to prevent error payloads from polluting the status field
3. Upgrade the OpenClaw image and automatically enable the Gateway
- Upgrade the default OpenClaw runtime image to the new image address
- Automatically initialize and enable the Gateway when the new image starts, reducing manual instance-side operations
- Update the system default image configuration and add migration logic for existing default values