Diagnosing Container Issues with the Cycle MCP Server.

When a container won't start, keeps restarting, or stops responding, you can ask your LLM (I'll use Claude as my reference for this guide) to investigate through the Cycle MCP server.

Claude reads your hub's events, logs, and telemetry, narrows down the cause, and suggests a fix.

This guide explains how that investigation works, what to ask for, and how to read the results. It walks through a real example: a Redis container that failed to start because of a missing capability.

If you'd like to try this out yourself use redis:8.10.2 , it currently requires CAP_SETPCAP, which means the container will fail on start.

Before you start

  • Connect the Cycle MCP server to your AI client and sign in to the hub you want to investigate.
  • Check your connector scopes. With read access, Claude can diagnose problems and recommend fixes, but it can't change anything. You apply the fix yourself in the Cycle portal. Actions that change state, such as starting a container or creating a DNS record, need a write-capable scope and your explicit confirmation.
  • Name the environment if you can. Container names aren't unique across a hub. If more than one container matches, Claude will ask which one you mean or check each of them.

Ask in plain language

You don't need to know any tool names. Describe the symptom and, if you know them, the container and environment:

  • "Diagnose the issue with container container-name and suggest how to fix it."
  • "My api container in the Demo environment keeps restarting. Check the logs and telemetry."
  • "Why can't I reach app.example.com?"
  • "Give me an overview of my hub. Is anything unhealthy?"

How Claude investigates

Claude usually works through the steps below, and stops as soon as the evidence is clear.

1. Find the container

list_containers finds the container by name, and reports its environment, current state, and instance count. If the name matches more than one container, this is where that shows up.

2. Run a targeted diagnosis

diagnose runs several checks at once. It returns findings ranked critical, warning, or info, and each finding names the tool to run next. Claude picks a focus based on your symptom:

Symptom

Focus

Scope

Won't start, crashing, restarting

crashloop

Container

URL or domain unreachable

unreachable

Environment, optionally with a container

502s or other load balancer errors

load_balancer

Environment

Out of disk, or "no space left" errors

storage

Server or container

CPU or RAM running low

resources

Server or container

Containers can't reach each other

networking

Environment

Deploy or build failed

build

Stack

By default, diagnose looks at the last 60 minutes, and it can look back as far as 24 hours. For a problem that comes and goes, ask Claude to check a longer window, for example "look at the last 24 hours." For anything older, Claude can search logs over a custom time range and pull up to 72 hours of metrics trends.

3. Read the logs

get_logs pulls log lines for a container or a single instance. It can filter by text or by a regular expression, such as (?i)error|panic|fatal, and can search any time range. The cause of a crash is usually in the last few lines before the instance exits.

4. Check resource telemetry

get_telemetry summarizes each instance's CPU usage and throttling, memory use against its limit, out-of-memory (OOM) kill count, and network errors. query_metrics with the container_instance_count preset shows the instance count hour by hour, which reveals flapping or unexpected scale-downs.

Telemetry can lag by up to about 10 minutes. Claude checks the latest sample time before drawing conclusions.

5. Check the server when a pattern spans containers

If several unrelated containers show the same symptom at the same moment, the cause is probably their host, not the containers. Claude can run diagnose against the server, or pull server telemetry with get_telemetry, to look for host-level events, load, and storage pressure.

6. Recommend a fix

Claude explains the root cause, shows the evidence behind it, and lists fixes from best to last resort. With read-only access, you make the change in the portal. Afterwards, you can ask Claude to confirm the container recovered.

Example: a Redis container that fails on start

Prompt: "Diagnose issue with container mcp-redis and suggest how to fix"

What Claude checked

Claude needed three tool calls.

  1. list_containers, searching for "redis," found two containers in the same environment: redis, which was running, and mcp-redis, which was stopped.
  2. diagnose on mcp-redis, with the crashloop focus over the last 24 hours, returned:
    • The only instance was in the failed state, with the error "Container exited with error code 127."
    • There was one error event in the window.
    • The container was stopped but configured to be running (state drift).
  3. get_logs showed that the instance ran for about 680 ms and printed a single line:
    setpriv: apply bounding set: Operation not permitted

Diagnosis

This was one failed start, not a crash loop, and Redis itself never ran.

The image's startup script runs as root. It uses setpriv to switch to the redis user and remove every Linux capability from the process. Removing capabilities requires the CAP_SETPCAP capability, and the container didn't have it. So setpriv failed and exited with code 127 before Redis could launch.

Suggested fixes, best first

  1. Add CAP_SETPCAP to the container's capabilities in its runtime configuration, then start it. This keeps the image's intended behavior, including running Redis as a non-root user. This is the actual fix, returned as the highest liklihood.

Other possibilities if the first were to fail:

  1. Run the container as the redis user (UID 999 in the official image). The startup script only switches users when it runs as root, so it skips the failing step. Make sure the /data volume is writable by that user.
  2. Use an image version whose startup script uses gosu instead of setpriv. gosu doesn't remove capabilities. Check the script in the tag you choose before relying on this.
  3. Override the entrypoint to redis-server to skip the startup script. This works, but Redis runs as root, so treat it as a last resort.

Confirming the cause and the fix

Verify recovery. After you apply a fix, ask Claude to "check whether the container started cleanly." Claude can wait on the logs for Redis's Ready to accept connections line instead of repeatedly checking.

Reading common signals

Signal

Usually means

Next step

Exit code 127

Command not found: the entrypoint or command is wrong, or the binary is missing from the image. Some tools, such as setpriv, also use 127 for their own errors.

Read the last log lines before the exit

Exit code 137

The process was killed with SIGKILL, often for running out of memory

Check the OOM kill count and the memory limit

Exit code 1

The application exited with an error

Read the logs

Instance ran under a second and printed one line

It failed in the entrypoint, before the app started

Check the image, entrypoint, and runtime config

stopped but should be running (state drift)

The last start failed, or something stopped the container

Run get_jobs to see what changed its state

[CYCLE COMPUTE] Console attached

Cycle attached to the instance's console. This appears on start and when the compute service reconnects.

On its own this doesn't mean a restart; look for container events and app startup lines

Instances in different containers reattach in the same second

Something happened on their shared host

Run diagnose on the server

"No LINKED record points to Container X"

The container has no public domain

Expected for internal services such as databases and backends

Network errors or drops on an instance

Packet loss on the instance's network interfaces

Compare other instances on the same server

Tips for trustworthy results

  • Ask for the evidence. Following up with "What makes you say this?" makes Claude separate what it observed from what it inferred. In testing, this caught cases where a correlation, such as timestamps lining up, had been stated more confidently than the evidence supported.
  • Notice what's missing. A container that "keeps restarting" should show restart events, changes in instance count, or repeated startup lines. If none of these appear, the symptom may be coming from somewhere else, such as the host or the load balancer.
  • Widen the window for intermittent problems. The default 60-minute lookback can miss a failure that happens a few times a day.
  • Account for telemetry lag. Metrics can be up to about 10 minutes behind. For what's happening right now, ask Claude to capture live output with capture_stream.

Limitations

  • Diagnosis is read-only. Even with write access, anything that changes state waits for your confirmation.
  • Brand-new containers may not have logs yet. If get_logs returns nothing, widen the time range or try again shortly.
  • Blocking waits are capped at 60 seconds. For longer operations, Claude repeats the call or tracks the work with get_jobs.
  • Load balancer telemetry can be missing. It isn't reported when the load balancer hasn't served traffic recently, so a missing report doesn't mean the load balancer is down.
  • Root causes can be inferences. Claude says when a conclusion is inferred from indirect evidence, such as an error message, and not confirmed from configuration. Check those points before acting on them.
Cookies

Cookies Preferences

We run basic, anonymous analytics by default to measure site traffic. By clicking "Accept," you allow additional cookies for advanced app improvements and tailored advertising. Choose what you share by clicking "Customize."