AppStatus Documentation Hub for Production Operations

Docs > Platform Observability > Getting Started with Container Monitoring

Getting Started with Container Monitoring

Overview

This guide introduces AppStatus Container Monitoring and covers:

  • Workload inventory across hosts and clusters
  • CPU and memory usage measured against configured limits
  • Restart and eviction tracking per workload
  • Image and version visibility per running container
  • Correlation with the request that the container served

What is Container Monitoring in AppStatus?

Container monitoring shows what is actually running, what each workload consumes against the limits it was given, and which containers are restarting or being evicted.

Because usage is compared to limits rather than to raw host capacity, you can see a workload heading for trouble while there is still time — a container sitting at 92% of its memory limit fails on the next traffic spike, not today.

Key capabilities:

  • Per-workload CPU and memory against configured limits
  • Restart count and reason, including out-of-memory kills
  • Image and tag visibility for every running container
  • Eviction and scheduling failure events
  • Link from a failing request to the container that served it

What you will see

What the workload list looks like

app.appstatus.ioSample data
WorkloadReplicasCPU vs limitMemory vs limitRestarts (24h)
checkout-api4/438%61%0
search-service6/681%92%7
payments-worker2/222%34%0
image-resizer1/296%99%23

Scroll the table sideways to see every column.

Memory limits cause hard kills; CPU limits cause slow requests. The symptom tells you which one you hit.

Troubleshooting

Usage shows but no percentage against limits

The workload has no resource limits configured. Without limits there is no ceiling to measure against — set requests and limits, and the pressure view starts working.

A container restarts repeatedly with low resource usage

A restart loop at low usage is almost never a resource problem. Check Error Tracking and the container logs — a failing dependency or bad configuration at startup is the usual cause.

Replica count is lower than expected

Look at the scheduling events. Containers that cannot be placed are usually blocked by resource requests that no node can satisfy, rather than by anything wrong with the image.

Operational Guidance

  • A restart loop is usually a config or dependency failure, not a resource one — check errors first.
  • Memory limits cause hard kills; CPU limits cause slow requests. The symptom tells you which.
  • Watch workloads sitting near their limit; they fail on the next traffic spike, not today.

Step-by-Step Setup

Container monitoring comes from the same agent as Infrastructure, so there is usually nothing new to install. The setup work is making sure resource limits exist, because without them there is no ceiling to measure pressure against.

Before you start

  • The agent installed on container hosts, or the in-cluster agent for Kubernetes
  1. 1

    Confirm workloads appear

    Open the workload list and confirm your containers are listed with their image and tag. Nothing extra is installed for this if the agent is already running on the host.

    WhereContainers → Workloads
  2. 2

    Set resource limits

    Make sure every workload has requests and limits configured. Pressure is measured against limits, so a workload without them shows usage but no percentage and cannot warn you before it fails.

    WhereYour deployment manifest
    Tip

    This is the step people skip, and it is the one that makes the whole page useful.

  3. 3

    Set restart thresholds

    Choose how many restarts count as unhealthy, per workload type. Stateful services should alert almost immediately; batch jobs can tolerate more.

    WhereContainers → Workload → Thresholds
  4. 4

    Name workloads to match services

    Where possible give the workload the same name as its APM service. That is what lets you move from a failing request to the container that served it without guessing.

    WhereDeployment labels

Configuration Options

Every option you can set, what each choice means, and what to pick. Use this as a reference while you fill in the form.

Workload settings

FieldOptionsWhat it doesRecommended
Resource limitsCPU / memoryCeiling that pressure is measured against.Always set. Without limits, pressure cannot be calculated.
Restart thresholdCount per windowWhen restarts count as unhealthy.Low for stateful workloads, higher for batch jobs.
Workload nameAny labelLinks the container to a service.Match the APM service name.

Feature Reference

Every feature, where to find it in the app, and what it does. Use this when you know what you want to do but not where it lives.

FeatureWhere in appDescription
Workload listContainers → WorkloadsWhat runs where, with usage against limits.
Restart trackingWorkload detailRestart count and reason, including out-of-memory kills.
Scheduling eventsWorkload detail → EventsWhy a container could not be placed.
Service correlationAPM → serviceTie a failing request to the container behind it.

Next Steps

Continue building your monitoring stack: