When a container won't start, keeps restarting, or stops responding, you can ask your LLM (I'll use Claude as my reference for this guide) to investigate through the Cycle MCP server.
Claude reads your hub's events, logs, and telemetry, narrows down the cause, and suggests a fix.
This guide explains how that investigation works, what to ask for, and how to read the results. It walks through a real example: a Redis container that failed to start because of a missing capability.
If you'd like to try this out yourself use redis:8.10.2 , it currently requires CAP_SETPCAP, which means the container will fail on start.
Before you start
- Connect the Cycle MCP server to your AI client and sign in to the hub you want to investigate.
- Check your connector scopes. With read access, Claude can diagnose problems and recommend fixes, but it can't change anything. You apply the fix yourself in the Cycle portal. Actions that change state, such as starting a container or creating a DNS record, need a write-capable scope and your explicit confirmation.
- Name the environment if you can. Container names aren't unique across a hub. If more than one container matches, Claude will ask which one you mean or check each of them.
Ask in plain language
You don't need to know any tool names. Describe the symptom and, if you know them, the container and environment:
- "Diagnose the issue with container
container-nameand suggest how to fix it." - "My
apicontainer in the Demo environment keeps restarting. Check the logs and telemetry." - "Why can't I reach
app.example.com?" - "Give me an overview of my hub. Is anything unhealthy?"
How Claude investigates
Claude usually works through the steps below, and stops as soon as the evidence is clear.
1. Find the container
list_containers finds the container by name, and reports its environment, current state, and instance count. If the name matches more than one container, this is where that shows up.
2. Run a targeted diagnosis
diagnose runs several checks at once. It returns findings ranked critical, warning, or info, and each finding names the tool to run next. Claude picks a focus based on your symptom:
Symptom | Focus | Scope |
|---|---|---|
Won't start, crashing, restarting |
| Container |
URL or domain unreachable |
| Environment, optionally with a container |
502s or other load balancer errors |
| Environment |
Out of disk, or "no space left" errors |
| Server or container |
CPU or RAM running low |
| Server or container |
Containers can't reach each other |
| Environment |
Deploy or build failed |
| Stack |
By default, diagnose looks at the last 60 minutes, and it can look back as far as 24 hours. For a problem that comes and goes, ask Claude to check a longer window, for example "look at the last 24 hours." For anything older, Claude can search logs over a custom time range and pull up to 72 hours of metrics trends.
3. Read the logs
get_logs pulls log lines for a container or a single instance. It can filter by text or by a regular expression, such as (?i)error|panic|fatal, and can search any time range. The cause of a crash is usually in the last few lines before the instance exits.
4. Check resource telemetry
get_telemetry summarizes each instance's CPU usage and throttling, memory use against its limit, out-of-memory (OOM) kill count, and network errors. query_metrics with the container_instance_count preset shows the instance count hour by hour, which reveals flapping or unexpected scale-downs.
Telemetry can lag by up to about 10 minutes. Claude checks the latest sample time before drawing conclusions.
5. Check the server when a pattern spans containers
If several unrelated containers show the same symptom at the same moment, the cause is probably their host, not the containers. Claude can run diagnose against the server, or pull server telemetry with get_telemetry, to look for host-level events, load, and storage pressure.
6. Recommend a fix
Claude explains the root cause, shows the evidence behind it, and lists fixes from best to last resort. With read-only access, you make the change in the portal. Afterwards, you can ask Claude to confirm the container recovered.
Example: a Redis container that fails on start
Prompt: "Diagnose issue with container mcp-redis and suggest how to fix"
What Claude checked
Claude needed three tool calls.
list_containers, searching for "redis," found two containers in the same environment:redis, which was running, andmcp-redis, which was stopped.diagnoseonmcp-redis, with thecrashloopfocus over the last 24 hours, returned:- The only instance was in the
failedstate, with the error "Container exited with error code 127." - There was one error event in the window.
- The container was
stoppedbut configured to be running (state drift). get_logsshowed that the instance ran for about 680 ms and printed a single line:
setpriv: apply bounding set: Operation not permitted
Diagnosis
This was one failed start, not a crash loop, and Redis itself never ran.
The image's startup script runs as root. It uses setpriv to switch to the redis user and remove every Linux capability from the process. Removing capabilities requires the CAP_SETPCAP capability, and the container didn't have it. So setpriv failed and exited with code 127 before Redis could launch.
Suggested fixes, best first
- Add
CAP_SETPCAPto the container's capabilities in its runtime configuration, then start it. This keeps the image's intended behavior, including running Redis as a non-root user. This is the actual fix, returned as the highest liklihood.
Other possibilities if the first were to fail:
- Run the container as the
redisuser (UID 999 in the official image). The startup script only switches users when it runs as root, so it skips the failing step. Make sure the/datavolume is writable by that user. - Use an image version whose startup script uses
gosuinstead ofsetpriv.gosudoesn't remove capabilities. Check the script in the tag you choose before relying on this. - Override the entrypoint to
redis-serverto skip the startup script. This works, but Redis runs as root, so treat it as a last resort.
Confirming the cause and the fix
Verify recovery. After you apply a fix, ask Claude to "check whether the container started cleanly." Claude can wait on the logs for Redis's Ready to accept connections line instead of repeatedly checking.
Reading common signals
Signal | Usually means | Next step |
|---|---|---|
Exit code 127 | Command not found: the entrypoint or command is wrong, or the binary is missing from the image. Some tools, such as | Read the last log lines before the exit |
Exit code 137 | The process was killed with SIGKILL, often for running out of memory | Check the OOM kill count and the memory limit |
Exit code 1 | The application exited with an error | Read the logs |
Instance ran under a second and printed one line | It failed in the entrypoint, before the app started | Check the image, entrypoint, and runtime config |
| The last start failed, or something stopped the container | Run |
| Cycle attached to the instance's console. This appears on start and when the compute service reconnects. | On its own this doesn't mean a restart; look for container events and app startup lines |
Instances in different containers reattach in the same second | Something happened on their shared host | Run |
"No LINKED record points to Container X" | The container has no public domain | Expected for internal services such as databases and backends |
Network errors or drops on an instance | Packet loss on the instance's network interfaces | Compare other instances on the same server |
Tips for trustworthy results
- Ask for the evidence. Following up with "What makes you say this?" makes Claude separate what it observed from what it inferred. In testing, this caught cases where a correlation, such as timestamps lining up, had been stated more confidently than the evidence supported.
- Notice what's missing. A container that "keeps restarting" should show restart events, changes in instance count, or repeated startup lines. If none of these appear, the symptom may be coming from somewhere else, such as the host or the load balancer.
- Widen the window for intermittent problems. The default 60-minute lookback can miss a failure that happens a few times a day.
- Account for telemetry lag. Metrics can be up to about 10 minutes behind. For what's happening right now, ask Claude to capture live output with
capture_stream.
Limitations
- Diagnosis is read-only. Even with write access, anything that changes state waits for your confirmation.
- Brand-new containers may not have logs yet. If
get_logsreturns nothing, widen the time range or try again shortly. - Blocking waits are capped at 60 seconds. For longer operations, Claude repeats the call or tracks the work with
get_jobs. - Load balancer telemetry can be missing. It isn't reported when the load balancer hasn't served traffic recently, so a missing report doesn't mean the load balancer is down.
- Root causes can be inferences. Claude says when a conclusion is inferred from indirect evidence, such as an error message, and not confirmed from configuration. Check those points before acting on them.