文件历史

78 次代码提交

作者 SHA1 备注 提交日期
Stark-Han 05480fdce9 fix(team): harden workflow recovery, terminal projection, and artifact visibility (#167)
Build / Backend Build (push) Failing after 1s
Build / Frontend Build (push) Failing after 1s
Build / Docker Build (push) Has been skipped
Build / Kubernetes Deploy Smoke Test (push) Has been skipped
Release / Prepare Scheduled Release (push) Has been skipped
Release / Publish Release (push) Failing after 0s
* chore: ignore local agency and team snapshots

* feat: add team agent profile templates

* KANBAN

* template

* prototype

* Fix tenant runtime pool deployment

* Fix local team profiles runtime override

* Update local team profiles testing guide

* Adapt local team profiles test deployment

* Patch tenant NFS server to ClusterIP

* Use NFS export root for workspace mounts

* Tighten team task completion projection

* Fix team event body map comparison

* before modification

* temp_peer1

* leadermodel

* team_fix

* fix(team): harden workflow completion and workspace handling

* fix(team): harden redis workflow protocol

* discord link updated

* chore: checkpoint team workflow fixes

* fix: stabilize leader kanban ledger

* fix: coalesce leader mediated work items

* fix: confirm leader-mediated member results

* fix: harden leader mediated team state

* fix: stabilize team shared runtime state

* Fix team task reconciliation and create flow

* fix: stabilize leader mediated team orchestration

* fix: refine team workspace presentation

* checkpoint: persist team workflow ledger fixes

* fix: reconcile team workflow delivery state

* fix: preserve meaningful team chat events

* docs: add team workspace quick guides

* chore: ignore local team deployment files

* fix: preserve setgid on team shared directories

* fix(team): align workspace overlays and completion projection

* fix(team): preserve workflow identity recovery

* fix(team): protect terminal workflow projection

* fix(team): preserve collaboration flow after terminal results
2026-07-17 17:48:58 +08:00
heranran 71c7672545 feat: Skill Hub hardening with session usage tracking and egress governance (#163)
* feat(skill-hub): add Skill Hub catalog, publish flow, and lite materialize pipeline

Introduce Skill Hub for browsing, importing, publishing, and installing skills, with lite instance package materialization and runtime sync support.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(skill-hub): remove token-governance hooks from Skill Hub PR

Strip validateManagedRuntimeEnvironmentOverrides, network lock policy sync,
and egress proxy audit wiring that belong to the upcoming token-usage work,
so the Skill Hub branch compiles independently.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(skill-hub): update migration number in materialize docs

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(skill-hub): make RuntimeAgentClient test stub and hub tests compile-safe

Add ResyncInstanceSkills to the runtime pool handler fake client, and harden
skill hub payload helpers/tests against nil storage/instance repos so go test passes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat: add Skill Hub hardening with session usage tracking and egress governance

Unify Skill Hub runtime sync improvements with session-token observability,
egress network policy, and admin/instance usage reporting for reopenable PR.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(skill-hub): repair CI tests and nested skill install

* fix(ci): restore release deployment configuration

---------

Co-authored-by: heshengran <heshengran@ieisystem.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-17 12:37:01 +08:00
Li-Day-Day 58c19856fd feat:Improve Lite gateway scheduling, proxy stability, and Hermes recovery (#160)
Release / Prepare Scheduled Release (push) Has been skipped
Release / Publish Release (push) Failing after 0s
* feat(runtime): improve Lite gateway scheduling and proxy stability

- Expand the runtime gateway port range to match 100 instances per pod
- Scale runtime deployments based on pending backlog
- Serialize gateway creation per pod and handle starting/creating states
- Clean up missing gateway bindings from runtime agent reports
- Improve Lite gateway proxy token stripping and auth header injection
- Update deployment manifests, docs, and related tests

* fix(runtime): harden Hermes Lite gateway recovery

- Add hashed gateway token aliases for short-term token compatibility
- Fix /resume conversations failing when old sessions use mismatched gateway tokens
- Recover ports only from stale error bindings without affecting active gateways

* feat(runtime): enable bidirectional desktop clipboard and direct proxy

- Enable WebTop/Selkies/Kasm clipboard by default with per-instance overrides
- Document clipboard verification, configuration pitfalls, and IME troubleshooting
- Enable direct desktop proxy in the K8s and K3s manifests

---------

Co-authored-by: litiantian03 <litiantian03@ieisystem.com>
2026-07-13 14:37:03 +08:00
Stark-Han 2dbe97d5b6 fix: stabilize OpenClaw Lite Team collaboration and documentation (#162)
Release / Prepare Scheduled Release (push) Has been skipped
Release / Publish Release (push) Failing after 1s
* chore: ignore local agency and team snapshots

* feat: add team agent profile templates

* KANBAN

* template

* prototype

* Fix tenant runtime pool deployment

* Fix local team profiles runtime override

* Update local team profiles testing guide

* Adapt local team profiles test deployment

* Patch tenant NFS server to ClusterIP

* Use NFS export root for workspace mounts

* Tighten team task completion projection

* Fix team event body map comparison

* before modification

* temp_peer1

* leadermodel

* team_fix

* fix(team): harden workflow completion and workspace handling

* fix(team): harden redis workflow protocol

* discord link updated

* chore: checkpoint team workflow fixes

* fix: stabilize leader kanban ledger

* fix: coalesce leader mediated work items

* fix: confirm leader-mediated member results

* fix: harden leader mediated team state

* fix: stabilize team shared runtime state

* Fix team task reconciliation and create flow

* fix: stabilize leader mediated team orchestration

* fix: refine team workspace presentation

* checkpoint: persist team workflow ledger fixes

* fix: reconcile team workflow delivery state

* fix: preserve meaningful team chat events

* docs: add team workspace quick guides

* chore: ignore local team deployment files

* fix: preserve setgid on team shared directories
2026-07-11 08:15:10 +08:00
Stark-Han 41d8d6a629 Codex/peer assisted team collab (#155)
Release / Prepare Scheduled Release (push) Has been skipped
Release / Publish Release (push) Failing after 1s
* chore: ignore local agency and team snapshots

* feat: add team agent profile templates

* KANBAN

* template

* prototype

* Fix tenant runtime pool deployment

* Fix local team profiles runtime override

* Update local team profiles testing guide

* Adapt local team profiles test deployment

* Patch tenant NFS server to ClusterIP

* Use NFS export root for workspace mounts

* Tighten team task completion projection

* Fix team event body map comparison

* before modification

* temp_peer1

* leadermodel

* team_fix

* fix(team): harden workflow completion and workspace handling

* fix(team): harden redis workflow protocol
2026-07-03 09:58:47 +08:00
litiantian03 2f2ad50793 fix:Backend Build Error 2026-07-01 17:23:20 +08:00
litiantian03 58f7b8aea3 feat(instance): support Lite instance batch management
- Add APIs for batch creating and deleting Lite instances
- Support Lite batch create, selection, and delete flows in the instance list
- Fix Lite OpenClaw/Hermes workspace initial paths and skill directory paths
- Hide transient skill temp directories from workspace browsing
- Preserve runtime query tokens when stripping ClawManager desktop access tokens
2026-07-01 17:04:14 +08:00
litiantian03 609996b65f feat(instance): improve skill injection and OpenClaw bootstrap config
- Add an available-skills endpoint for instances with actor-aware filtering
- Persist gateway token, agent bootstrap token, and OpenClaw config snapshots for Lite instances
- Inject agent and OpenClaw bootstrap env into gateway runtimes
- Materialize and remove Lite instance skills in the workspace filesystem
- Improve instance detail skill upload, bulk attachment, workspace refresh, and layout syncing
2026-07-01 17:04:14 +08:00
litiantian03 69d30df667 feat(skill): improve instance skill management and workspace sync
- Support multi-select skill attachment and show risk-policy filtering hints
- Mark instance skills as removed after successful uninstall commands
- Sync skill removal state when related workspace files are deleted
- Hide removed skills and avoid reactivating them from agent inventory reports
- Clear proxy-scoped OpenClaw control storage to prevent cross-instance state leakage
2026-07-01 17:04:14 +08:00
litiantian03 eca9ad5754 feat(instance): improve desktop stream apply flow and Team member config 2026-07-01 17:04:14 +08:00
Qingshan Chen ee4d9c4cc7 Fix runtime manifest test paths 2026-06-25 14:41:21 +08:00
Qingshan Chen 8c922189d2 Merge upstream main into sc 2026-06-25 14:29:49 +08:00
Qingshan Chen 49da74e87f Add single-node deployment manifests 2026-06-25 14:16:28 +08:00
litiantian03 ccb513eaf3 fix(instance): skip desktop direct proxy for runtime gateways 2026-06-18 15:41:58 +08:00
litiantian03 3c35b26b74 fix(instance): stabilize desktop access and stream defaults
- Reuse desktop direct upstream resolution and include upstream in short external access tokens
- Default and normalize Selkies stream settings for desktop runtimes
- Remove legacy deployments with the same instance-id before ensuring the current deployment
- Preserve the unsaved desktop stream profile state in the frontend
- Add tests for stream profiles, pod envs, and deployment cleanup
2026-06-18 11:28:40 +08:00
litiantian03 b8a7e8c67a feat: optimize shared desktop access and multi-replica control plane
- Add leader election to avoid duplicate sync and Team event consumers
- Proxy desktop traffic directly from nginx to instance Services to reduce load during concurrent viewers
- Add desktop stream profiles with low / standard / high options
- Limit hostPath PV fallback to manual storageClass only
2026-06-18 11:28:40 +08:00
Qingshan Chen abfd47c1d0 fix: align instance pods with pvc node affinity 2026-06-16 18:25:44 +08:00
Qingshan Chen f415f09716 fix openclaw lite gateway bind mode 2026-06-15 17:25:14 +08:00
Qingshan Chen aea224426f feat: support lite/pro runtime modes and runtime rollout 2026-06-14 14:40:49 +08:00
litiantian03 95d6faad73 fix: resource skill upload and team member numbering 2026-06-02 16:22:11 +08:00
litiantian03 7fa5c94093 feat(k8s): manage instance workloads with Deployments
Create instance workloads as single-replica Deployments instead of bare Pods so Kubernetes can recreate Pods after deletion, eviction, or node/runtime failures.
Update stop/delete and orphan cleanup to remove Deployments before residual Pods, and improve status sync to keep instances creating while a Deployment exists without a Pod and avoid stale Failed/Evicted Pods affecting instance state.
2026-06-02 10:14:23 +08:00
litiantian03 29375f04f2 fix(instance): mount Hermes/desktop PVCs at /config with legacy layout migration 2026-06-02 09:15:59 +08:00
litiantian03 6f02978997 Tune desktop clipboard and shared memory defaults
- Disable Webtop/KasmVNC bidirectional clipboard sync by default to reduce accidental paste during remote desktop input
- Derive desktop runtime /dev/shm defaults from instance memory: 1G for 4G, 2G for 8G, and 4G for 12G+
- Preserve explicit SHM_SIZE_GB overrides and keep shell runtime behavior unchanged
- Add tests for openclaw/hermes clipboard policy and /dev/shm default selection
- Support multiple OpenClaw channel resources and fix alias injection
2026-05-28 16:25:24 +08:00
litiantian03 561b263c8b feat: add Bundle Skills, WeCom channel, configurable archive limit, fix #123
- Bundle Skills: attach skills to config bundles, auto-resolve on creation. Migration 022 adds openclaw_config_bundle_skills table. Fix #123
- WeCom channel: wecom connector with botId/secret/dmPolicy/allowFrom
- CLAWMANAGER_WORKSPACE_ARCHIVE_MAX_MIB controls archive size limit, synced to nginx client_max_body_size via start.sh
- Update deployment manifests, i18n, and frontend UI
2026-05-26 17:17:07 +08:00
litiantian03 a64fa20997 feat(team): add paginated task/event history and improve chat deduplication
Add cursor-based GET /teams/:id/tasks and GET /teams/:id/events APIs.
Team detail page loads older messages on scroll/click and deduplicates
collaboration chat messages with improved thread ordering.
2026-05-25 10:29:25 +08:00
Qingshan Chen d47c5f0b8c feat: add openclaw shell runtime preset 2026-05-22 19:21:11 +08:00
Qingshan Chen c192908d95 feat: add shell runtime and harden team shared storage 2026-05-22 15:06:59 +08:00
Qingshan Chen 8bfd6340e9 Merge remote-tracking branch 'upstream/main' into codex/shell
# Conflicts:
#	backend/internal/services/instance_service.go
#	backend/internal/services/k8s/pod_service.go
#	backend/internal/services/k8s/pod_service_test.go
#	frontend/src/components/InstanceAccess.tsx
2026-05-18 22:12:46 +08:00
litiantian03 3fe1418d5e feat(team): Multi-agent Team control plane across runtimes (Leader-mediated collaboration)
- Add migrations plus Team/Member/Task/Event APIs; Redis Streams consumer projects inbox/events into DB as source of truth
- Inject Team Secret (Redis URL, team token) via envFrom; shared RWX PVC at /team; sync ConfigMap roster to /team/team.json
- Create member Pods through InstanceService; extend K8s (PVC/Secret/Pod/ConfigMap); stale-task sweep and Team/member deletion with cleanup
- Add /teams, /teams/new, /teams/🆔 creation wizard (roster, presets, shared env/OpenClaw plan), per-member desktops, collaboration timeline, debug dispatch (defaults to Leader when target omitted)
2026-05-18 16:30:26 +08:00
Qingshan Chen 7ceefbf6a8 shell init 2026-05-13 20:19:44 +08:00
Qingshan Chen a3bbedbbdc fix: paginate user management list 2026-05-12 20:35:37 +08:00
litiantian03 c0c6b763af fix build issue 2026-05-07 11:00:33 +08:00
litiantian03 054a2b9b44 fix: resolve OpenClaw upload crash with least-privilege pod settings
- add configurable /dev/shm mounts for instance pods with a bounded SHM_SIZE_GB override
- introduce pod security modes and use chromium-compat for OpenClaw instead of privileged
- keep privileged mode only as an explicit admin fallback
- remove the node-level clawmanager-node-tuner manifest to avoid changing host security defaults
- fix k3s HTTPS and API/proxy service port mappings
- add tests for SHM parsing and pod security mode behavior
2026-05-07 10:34:37 +08:00
hippoley 963a1c378e fix: use persistent path for PV hostPath instead of /tmp/
BREAKING: PV hostPath prefix changed from /tmp/clawreef/ to /data/clawreef/

Problem:
- PV hostPath was hardcoded to /tmp/clawreef/user-{id}/instance-{id}
- /tmp/ is a volatile directory that may be cleaned by systemd-tmpfiles
  or cleared on node reboot (depending on OS configuration)
- This caused data loss when worker-02 was rebooted by the hypervisor

Fix:
- Add configurable hostPathPrefix to RuntimePVCConfig (default: /data/clawreef)
- Support K8S_PV_HOST_PATH_PREFIX environment variable override
- Update deployment manifests to use /data/clawmanager/ for MySQL and MinIO

Migration:
- Existing deployments should move data from /tmp/clawreef/ to /data/clawreef/
  and create a symlink for backward compatibility:
    mkdir -p /data/clawreef
    mv /tmp/clawreef/* /data/clawreef/
    rm -rf /tmp/clawreef && ln -s /data/clawreef /tmp/clawreef
2026-05-07 07:57:44 +08:00
Qingshan Chen 55e9edfb83 fix: make instance deletion asynchronous 2026-04-30 19:24:17 +08:00
Qingshan Chen 6c8bf5180c fix: normalize agents runtime image references 2026-04-29 17:24:49 +08:00
Qingshan Chen 58c35f6ecd Merge pull request #88 from naiqus/fix/openclaw-channel-editor-field-dropping
fix: channel editors silently dropping unknown fields on save
2026-04-29 15:14:37 +08:00
Qingshan Chen e2613b4c45 feat: add hermes runtime integration 2026-04-29 15:00:57 +08:00
Suqian Zhang 08dce2a29d fix: remove duplicate test function declarations
TestNormalizeOpenClawResourceContentPreservesUnknownChannelFields and
TestNormalizeOpenClawResourceContentPreservesFeishuSiblingAccounts were
declared twice causing build failure.
2026-04-27 23:48:07 +02:00
Suqian Zhang 032c038275 Merge branch 'fix/openclaw-channel-editor-field-dropping' of github.com:naiqus/ClawManager into fix/openclaw-channel-editor-field-dropping 2026-04-27 23:08:58 +02:00
Suqian Zhang c39df72d51 fix(openclaw): preserve unknown fields in channel config save path
Every channel form editor (telegram, dingtalk-connector, slack, feishu)
silently drops fields on Save. The user flow that loses data:

  1. Open a channel resource in the JSON tab, add any field not owned
     by the form (e.g. webhook, custom capabilities, additional feishu
     accounts).
  2. Switch to the Form tab and change anything (e.g. rotate a secret).
  3. Click Save.
  4. The field from step 1 is gone from the saved config — and stays
     gone after a reload.

The data loss occurs in two places; both must be fixed:

Frontend (OpenClawConfigCenterPage.tsx):
  Each update*ChannelContentText rebuilt the output from a hard-coded
  allowlist instead of merging form fields into the parsed config. Any
  field not in the allowlist was dropped on Save. For feishu this
  collapsed a multi-account config to a single `main` account.

Backend (openclaw_config_service.go):
  normalizeOpenClawResourceContent calls normalize*ChannelConfigForEnv
  on every create/update to validate and rebuild the stored config.
  Those helpers use the same allowlist pattern, so even if the
  frontend sends a complete config, the backend strips tenant-authored
  keys before persisting.

Fix:

  - Frontend: introduce a single mergeChannelConfig(existing, allowlisted)
    helper and route every editor's output through it. Each handler
    builds an allowlisted object containing only its known keys; the
    helper overlays it onto the parsed existing config so unknown keys
    survive. The merge point is now structural rather than a per-handler
    convention, so a new editor (e.g. an MS Teams handler added in a
    downstream patch) can't regress by forgetting the spread. Feishu's
    accounts deep-merge is preserved by building the merged accounts
    inside the allowlisted object before handing off to the helper.

  - Backend: introduce mergeOpenClawChannelConfigForStorage, which
    overlays the allowlist-normalized output onto the original parsed
    payload at storage time. Normalized keys still win; unknown keys
    pass through. For feishu the accounts map is merged member-wise.

Env-rendering semantics are unchanged. appendCompiledOpenClawEnvPayload
continues to call the *ForEnv helpers at render time against the
stored content, so the runtime env payload sees the same allowlist as
before.

Tests:

  - Two existing ResourcePayloadFromModelNormalizesStored*JSON tests
    asserted the pre-fix behavior (dropping `legacyField`) and are
    updated to assert it survives.
  - New TestNormalizeOpenClawResourceContentPreservesUnknownChannelFields
    (table-driven, three editors) guards the preservation invariant.
  - New TestNormalizeOpenClawResourceContentPreservesFeishuSiblingAccounts
    guards the feishu deep-merge.

Manual verification against a live instance: add `webhook` via the
JSON tab, switch to Form tab, rotate a secret, Save, reload — the
webhook field is retained.
2026-04-27 23:07:24 +02:00
Qingshan Chen 30abf94c45 Merge pull request #95 from hippoley/fix/image-pull-policy-configurable
Release / Prepare Scheduled Release (push) Has been skipped
Release / Publish Release (push) Failing after 0s
fix: default imagePullPolicy to IfNotPresent for air-gapped environments
2026-04-27 21:46:35 +08:00
hippoley 584b071e35 fix: default imagePullPolicy to IfNotPresent for air-gapped environments
Kubernetes defaults imagePullPolicy to Always when the image tag is
:latest. This causes pod creation failures in air-gapped and enterprise
environments where nodes cannot reach external registries.

Changes:
- Add ImagePullPolicy field to PodConfig struct
- Default to IfNotPresent in CreatePod when not explicitly set
- Support IMAGE_PULL_POLICY env var for operator override
- Apply policy to both CreateInstance and StartInstance paths

Closes #94
2026-04-27 16:32:30 +08:00
litiantian03 bf0125ab23 fix(instances): delete legacy network policy before create/start instead of ensuring default 2026-04-27 09:43:52 +08:00
Suqian Zhang 87fe5a6bdf fix(openclaw): preserve unknown fields in channel config save path
Every channel form editor (telegram, dingtalk-connector, slack, feishu)
silently drops fields on Save. The user flow that loses data:

  1. Open a channel resource in the JSON tab, add any field not owned
     by the form (e.g. webhook, custom capabilities, additional feishu
     accounts).
  2. Switch to the Form tab and change anything (e.g. rotate a secret).
  3. Click Save.
  4. The field from step 1 is gone from the saved config — and stays
     gone after a reload.

The data loss occurs in two places; both must be fixed:

Frontend (OpenClawConfigCenterPage.tsx):
  Each update*ChannelContentText rebuilt the output from a hard-coded
  allowlist instead of merging form fields into the parsed config. Any
  field not in the allowlist was dropped on Save. For feishu this
  collapsed a multi-account config to a single `main` account.

Backend (openclaw_config_service.go):
  normalizeOpenClawResourceContent calls normalize*ChannelConfigForEnv
  on every create/update to validate and rebuild the stored config.
  Those helpers use the same allowlist pattern, so even if the
  frontend sends a complete config, the backend strips tenant-authored
  keys before persisting.

Fix:

  - Frontend: introduce a single mergeChannelConfig(existing, allowlisted)
    helper and route every editor's output through it. Each handler
    builds an allowlisted object containing only its known keys; the
    helper overlays it onto the parsed existing config so unknown keys
    survive. The merge point is now structural rather than a per-handler
    convention, so a new editor (e.g. an MS Teams handler added in a
    downstream patch) can't regress by forgetting the spread. Feishu's
    accounts deep-merge is preserved by building the merged accounts
    inside the allowlisted object before handing off to the helper.

  - Backend: introduce mergeOpenClawChannelConfigForStorage, which
    overlays the allowlist-normalized output onto the original parsed
    payload at storage time. Normalized keys still win; unknown keys
    pass through. For feishu the accounts map is merged member-wise.

Env-rendering semantics are unchanged. appendCompiledOpenClawEnvPayload
continues to call the *ForEnv helpers at render time against the
stored content, so the runtime env payload sees the same allowlist as
before.

Tests:

  - Two existing ResourcePayloadFromModelNormalizesStored*JSON tests
    asserted the pre-fix behavior (dropping `legacyField`) and are
    updated to assert it survives.
  - New TestNormalizeOpenClawResourceContentPreservesUnknownChannelFields
    (table-driven, three editors) guards the preservation invariant.
  - New TestNormalizeOpenClawResourceContentPreservesFeishuSiblingAccounts
    guards the feishu deep-merge.

Manual verification against a live instance: add `webhook` via the
JSON tab, switch to Form tab, rotate a secret, Save, reload — the
webhook field is retained.
2026-04-26 17:50:08 +02:00
Qingshan Chen 3a44141d97 Merge pull request #92 from hippoley/fix/graceful-shutdown-and-goroutine-leaks
Release / Prepare Scheduled Release (push) Has been skipped
Release / Publish Release (push) Failing after 1s
fix: add graceful shutdown, fix WebSocket Hub race, plug goroutine leaks
2026-04-26 00:42:45 +08:00
Qingshan Chen 49725a8e24 Merge pull request #87 from naiqus/fix/openclaw-transfer-path
fix: .openclaw workspace transfer path and upload handling
2026-04-26 00:27:20 +08:00
Qingshan Chen 313cc57ddc Merge pull request #89 from naiqus/fix/admin-instance-view-leakage
fix: admin sees all users' instances in workspace view
2026-04-26 00:04:30 +08:00
hippoley cfd9c46ca8 fix: add graceful shutdown, fix WebSocket Hub race, plug goroutine leaks
Address three HIGH-severity issues from the memory-leak and bloat
analysis (issue #56).

1. Graceful shutdown (main.go)
   - Replace gin's r.Run() with an explicit http.Server so the process
     can intercept SIGINT / SIGTERM.
   - On signal: drain active HTTP requests (10 s timeout), then stop
     SyncService, WebSocket Hub, and InstanceAccessService cleanup
     goroutine in order.
   - Ensures database connections, K8s watchers, and background loops
     are released cleanly on deploy or restart.

2. WebSocket Hub init race (websocket_service.go)
   - GetHub() used a bare nil-check with no synchronisation; two
     goroutines could each create a Hub and start a Run() loop.
   - Replaced with sync.Once to guarantee exactly one Hub instance.
   - Added a stop channel to Hub.Run() so the hub can be shut down
     gracefully, closing all connected clients.

3. InstanceAccessService goroutine leak (instance_access_service.go)
   - cleanupExpiredTokens() looped on ticker.C with no exit path,
     leaking the goroutine for the lifetime of the process.
   - Added a stopChan; cleanupExpiredTokens now selects on both the
     ticker and the stop signal.
   - Exposed Stop() on the service; InstanceHandler.Shutdown() calls
     it during graceful shutdown.

Tests:
- TestGetHubReturnsSameInstance: singleton guarantee
- TestGetHubConcurrentAccess: 50-goroutine race test
- TestHubStopClosesClients: verifies client cleanup on Stop()
- TestInstanceAccessServiceStopTerminatesCleanup: Stop() is safe and
  the service remains functional for token ops afterward

All existing tests continue to pass; full project build and regression
verified.

Ref: #56
2026-04-24 20:07:13 +08:00
hippoley 7bb6a19b85 fix: cascade recompile active snapshots when a resource is updated
When a resource is updated, all compiled/active snapshots that reference
it (directly or via dependsOn) are now automatically recompiled so their
rendered_env_json stays in sync.

Key changes:
- snapshotReferencesResource checks ResolvedResourcesJSON first (full
  resolved set including dependsOn), falling back to
  SelectedResourceIDsJSON for backward compatibility with older snapshots
- Repository-layer CAS via UPDATE ... WHERE updated_at = ? prevents
  concurrent cascades from clobbering each other
- ListActiveSnapshots queries both 'compiled' and 'active' statuses
- Cascade runs in a background goroutine; errors are logged but never
  fail the resource update itself
2026-04-24 18:08:03 +08:00