Troubleshooting & Diagnostics
Common Issues & Fixes
1. docker: No such image: cloud-harness-executor:local
Cause: Local executor image was pruned by a host Docker cleanup or never built. Fix: Rebuild the image from the project root:
docker compose --profile images build executor-image2. repository clone failed: unauthorized
Cause: Attempting to clone a private repository without a valid GitHub App installation. Fix:
- Open the Operator Dashboard → GitHub.
- Click Install GitHub App and authorize the target repository.
- Ensure the repository URL matches the format
https://github.com/owner/repo.git.
3. workspace expired: TTL exceeded
Cause: The workspace reached its 15-minute wall-clock limit or 5-minute idle limit. Fix: Workspaces are ephemeral by design. Re-open a workspace using workspace_open with a fresh idempotency key.
4. API Key Denied (401 Unauthorized)
Cause: Expired key, revoked key, or key used against the Managed OAuth URL instead of the gateway. Fix:
- Ensure the client URL is
https://api.harness.zuey.me/mcp(NOThttps://harness.zuey.me/mcp). - Verify key validity in the Dashboard under API Keys.
5. OAuth DCR Error (redirect_uri is not allowed by the account configuration)
Cause: Cloudflare Access Managed OAuth rejected Dynamic Client Registration because the client's callback URL was not allowlisted. Fix:
- Log into Cloudflare Zero Trust → Access controls → Applications.
- Edit the application for your MCP hostname → Advanced settings → Managed OAuth.
- Add the required callback URLs to Allowed redirect URIs:
- Claude Desktop:
https://claude.ai/api/mcp/auth_callbackandhttps://claude.com/api/mcp/auth_callback - Codex App / Native Clients: Pin
mcp_oauth_callback_port = 3118in~/.codex/config.tomland addhttp://127.0.0.1:3118/callback/*,http://127.0.0.1:3118/*,http://localhost:3118/callback/*, andhttp://localhost:3118/*. - ChatGPT Web:
https://chatgpt.com/connector/oauth/*,https://chatgpt.com/connector_platform_oauth_redirect, andhttps://chatgpt.com/api/aip/p/oauth/callback.
- Claude Desktop:
6. Frequent MCP Sign-out or Re-authentication Prompts in AI Tools
Cause: When connecting via Managed OAuth (https://harness.zuey.me/mcp), client continuity depends on Cloudflare Access's Grant session duration (refresh token lifetime). When the grant expires, the client prompts for interactive browser re-authentication.
Fix:
- Adjust Managed OAuth Grant Session Duration (OAuth Clients):
- Log into Cloudflare Zero Trust → Access controls → Applications.
- Edit the MCP application → Advanced settings → Managed OAuth.
- Set Grant session duration to your preferred continuity interval (Cloudflare recommends 1–2 weeks for CLI/agent clients, or longer up to 1 month where supported by the tenant).
- Keep the Access token lifetime short (5–15 minutes, default 15 minutes) so silent refresh and policy re-evaluation continue normally.
- Switch to Static API Key Gateway (Zero-Reauth for Coding Tools):
- For IDE/CLI coding agents (Claude Code, Cursor, Codex, etc.) that support static headers, generate an API key from the Dashboard at
https://harness.zuey.me/dashboard/api-keys(configurable for 1 to 3,650 days, approximately 10 years). - Configure the tool to connect directly to
https://api.harness.zuey.me/mcpwithAuthorization: Bearer <api-key>to eliminate interactive OAuth prompts entirely.
- For IDE/CLI coding agents (Claude Code, Cursor, Codex, etc.) that support static headers, generate an API key from the Dashboard at
7. Local Stdio: --workspace path must be absolute or Directory Error
Cause: The --workspace argument provided to cloud-harness-mcp --transport stdio is relative, does not exist, or points to a regular file instead of a directory. Fix: Provide a valid, existing absolute directory path (e.g. /home/user/project or /mnt/c/Users/user/project in WSL). Native Windows path formats (like C:\...) are unsupported in v1 local stdio mode; run the process inside WSL instead.
8. ChatGPT: FORBIDDEN: This conversation does not support developer MCPs
Cause: ChatGPT allows tool discovery, but blocks invocation because the active conversation surface (e.g. Custom GPT, Project chat, Canvas, Mobile app, or temporary chat) or user account restricts draft developer MCPs, or the connector is in draft state without workspace publishing. Fix:
- Open Standard 1-on-1 Web Chat: Use ChatGPT Web in a standard chat thread and select or
@mentionCloudHarness. - Enable Developer Mode: Verify that Settings → Apps → Advanced Settings → Developer mode is enabled for your account.
- Publish Connector (Workspace Admins): In Workspace Settings → Apps → Drafts, select CloudHarness and click Publish to promote it from a draft Developer MCP to an approved workspace Custom Connector.
- Verify Plan Support: Full MCP write actions (such as
workspace_open) are in beta for ChatGPT Business, Enterprise, and Edu plans. - Start Fresh Thread: If the connector was recently created or authorized, open a new chat session to clear stale conversation state.
- See ChatGPT Configuration Guide for complete setup steps.
9. Dashboard Shows a Diagnostic Page or JSON authentication_failed
Cause: The request reached the Cloud Harness origin without a valid Cloudflare Access assertion, so the API could not identify the caller. Typical causes: the Access application does not cover the dashboard hostname and path, a bypass or service-auth policy matched the request, the browser resolved the origin address instead of the Cloudflare-proxied hostname, or the origin no longer agrees with the live Access application (for example after the application was recreated, or after the team's signing keys rotated).
Fix:
- Read the reason code shown on the page, or the
access assertion rejectedline in the API log. It names the failing check —missing_assertion,wrong_audience, andjwks_unavailablecover most incidents, and the diagnostic module owns the full set. - In Cloudflare Zero Trust → Access controls → Applications, confirm the application covers the dashboard hostname and path, and compare its Application Audience (AUD) tag with the origin's
CLOUDFLARE_ACCESS_AUDIENCE. - Open the dashboard on the Cloudflare-proxied public hostname (
https://harness.zuey.me/dashboard). A hosts-file or router override that resolves it to the origin address bypasses Access and produces this page. - For
jwks_unavailableorunknown_key, check that the API container can reach the team's/cdn-cgi/access/certsendpoint and that the host clock is correct.
Non-browser clients keep receiving the compact {"error":"authentication_failed"} JSON body; only browser navigations render the diagnostic page, and no token, assertion, or identity claim is ever shown.
10. DEPENDENCY_EGRESS_UNAVAILABLE on workspace_open (HTTP 503)
Cause: The effective network profile is dependency-access — the shipped default — but the Linux host firewall is not provisioned, or its rules have drifted, so the runner fails the open closed instead of silently downgrading to network-none. Fix:
- Provision the host firewall on the Docker host:
bash deploy/scripts/setup-dependency-firewall.sh- Confirm Egress readiness reports
Readyon the dashboard Settings page. - Open the workspace again with a fresh idempotency key. The failed attempt kept its key with a
FAILEDstatus, and replaying that key returns the failed record without retrying the attestation. To work without egress meanwhile, reset the default tonetwork-noneon the Settings page or open the workspace withnetworkProfile: "network-none".
11. GITHUB_PERMISSION_MISSING / 403 Resource not accessible by integration
Cause: No configured credential can perform the requested github_action. The GitHub App installation did not grant the scope the action needs, and no fallback credential is available for the requesting principal. Fix:
- The error names the missing scope. Add that permission to the GitHub App and approve the pending installation change on GitHub, then retry. The GitHub App setup guide lists which operations need which permission. An action-scoped token also carries Contents: Read-only so that
ghcan resolve repository metadata, so that permission is required forgithub_actioneven when the named action scope is already granted. - Alternatively, configure the fallback credential for that principal: the runner-environment
GH_TOKEN/GITHUB_TOKENinowner-bearermode, or that principal's global runtime secret in Access mode. The runner-environment credential is harness-side only and never enters an executor; a principal's global runtime secret is injected into that principal's workspaces, so it also authenticates the workspaceghCLI. A403from the helper is never retried, because the operation may already have had side effects; inspect the issue or pull request before retrying.
12. LIMIT_EXCEEDED: active workspace limit reached on workspace_open
Cause: The principal already holds MAX_ACTIVE_WORKSPACES_PER_OWNER counted workspaces (CREATING, ACTIVE, NETWORK_QUARANTINED). A record in REAPING is in flight to teardown and holds no slot. Multiple concurrent workspaces are supported by design, so this is a quota rather than a harness limitation. Fix:
- Call
workspace_listto see the counted workspaces, thenworkspace_closeone you no longer need. Closing removes that workspace's files, so finalize or push unpushed work first. - An
ACTIVEworkspace also frees its slot when its idle or wall TTL expires. ANETWORK_QUARANTINEDrecord does not expire, so close it explicitly. - Raise
MAX_ACTIVE_WORKSPACES_PER_OWNERon the runner (default3, maximum64) after sizing host memory for the new limit, because each counted workspace may use up to 1 GiB of container memory, one CPU, and 256 pids. Lowering the limit never reaps an existing workspace; it only blocks new admission and recovery until the counted total drops.
13. A workspace is stuck in REAPING
Cause: A teardown that failed before its final CLOSED write leaves the record in REAPING. A REAPING record holds no capacity slot, so it never blocks workspace_open; the visible symptom is a workspace that will not disappear from the dashboard. Fix: Call workspace_close again on that workspace. The close path skips the claim for a record already in REAPING and retries container and path removal, so a repeat close is the supported remedy. Only the fenced dashboard close refuses a REAPING record with 409 CONFLICT. If removal keeps failing, fix the underlying Docker or filesystem fault rather than running broad Docker or database cleanup.
14. Dashboard skill ZIP upload returns 413 below 8 MiB
Cause: Older managed nginx dashboard routes inherit the server-level 1 MiB request cap even though Cloud Harness accepts skill archives up to 8 MiB. Fix: Deploy the current release or run deploy/scripts/upgrade-nginx-dashboard.sh on the host. The managed /dashboard/ route is upgraded to client_max_body_size 8m. Archives larger than 8 MiB are still rejected intentionally by the API and runner.
Also rejected as INVALID_INPUT: an archive with more than 200 SKILL.md documents, more than 20,000 files, or a SKILL.md longer than 64 Ki characters. Only SKILL.md documents are imported, so bundled assets and macOS __MACOSX entries are ignored rather than counted against the skill limit. Split a larger library into several archives.
15. agentkit toolkit fails during workspace_open
Cause: The licensed AgentKit kit kind fails closed by design, and each message names the missing prerequisite. Fix:
AgentKit kits are not configured on this instance— the operator must set bothAGENTKIT_REGISTRY_KEY_IDandAGENTKIT_REGISTRY_PUBLIC_KEY(the pinned Ed25519 registry signing key) and restart the runner.needs the <name> secret stored for this principal— store the AgentKit licence token (anak_dev_/ak_cli_credential) as a global secret with that name (AGENTKIT_REGISTRY_TOKENunlessAGENTKIT_REGISTRY_CREDENTIAL_SECRETchanges it). The secret must be created withpurpose: provisioning; the runner refuses aruntime-purpose token so it can never be injected into an executor. Credentials are never accepted as tool arguments.registry rejected the credential ... (not_licensed | not_authenticated | license_inactive)— the token has no entitlement for thatkitId, or is not a registry bearer. Re-issue it from the AgentKit account that owns the licence.manifest signature did not verify/signed by key ...; pinned key is ...— the registry signing key rotated or the pinned key is stale. UpdateAGENTKIT_REGISTRY_KEY_IDandAGENTKIT_REGISTRY_PUBLIC_KEYtogether; do not disable verification.package digest did not match the signed manifest,must contain exactly one <kitId> root directory, oroutside the <kitId> root— the downloaded artifact is not the published package. Retry once; if it persists, treat it as an upstream publish problem instead of mounting the content.no published release for kit <kitId>— that kit has no signed release on the requested channel yet. Trychannel: "beta", or pin a version that exists.
15. Uploaded operator skills do not appear in skills_list
Cause: The built-in tier is populated only when the runner is pointed at an operator-owned host directory, and only reads a strict layout. Fix:
- Set
BUILTIN_SKILLS_ROOTto an absolute host directory in the runner environment (/etc/cloud-harness-mcp/runtime.envor.env) and restart the stack; with the variable unset the executor mounts nothing and the tier stays empty by design. - Upload skills as
<root>/<skill-name>/SKILL.md; a directory withoutSKILL.mdis ignored. - Confirm the host path is readable by the runner and that the directory exists (
deploy/scripts/bootstrap-vps.shcreates/var/lib/cloud-harness/skillson first install). - Remember the tier outranks project skills: a same-named
.agents/skillsor.cloud-harness/skillsentry appears undershadowed, not as the selected skill. - Only changes to already-open workspaces need a reopen; a new workspace sees an updated upload immediately.
BUILTIN_SKILLS_ROOTis reserved and is now the single name the worker reads. If the runner rejects a stored secret or environment record with that name (INVALID_INPUT), remove it withsecret_delete(deletion of a reserved name stays allowed) and reopen. The removedCH_BUILTIN_SKILLS_ROOToverride is no longer read: rename it toBUILTIN_SKILLS_ROOT. A container created before this change keeps its old environment until it is closed or rebuilt.
16. git_log only shows one commit, or read-only GitHub calls prompt for approval every time
Cause: workspace_open clones a single commit by default, so a fresh workspace's history is one commit deep until deepened. Separately, MCP client approval prompts key off tool-level annotations: github_action is annotated destructive as a whole tool because it also performs mutations, so a client can prompt even for a read action like pr_list since there is no per-action server approval gate. Fix:
- For history depth, pass
fetchDepth(0 for full history, or a commit count) orshallowSince(an ISO date/datetime) toworkspace_open, or callgit_fetchafterward withdepth,unshallow, orshallowSince. - For GitHub reads, call the dedicated
github_readtool (pr_list,pr_view,issue_list,issue_view,commit_list,compare,release_list,tag_list) instead ofgithub_action. It is annotated read-only/idempotent/non-destructive, so compliant clients do not prompt per call.github_actionstill accepts the same read actions for compatibility.