vibe-service-health-dashboard
Queries all configured services for health status — uptime, restarts, errors, resources. Use when checking the health of running services.
vibe-service-health-dashboard
Know what's running, what's crashed, and what's about to crash.
When to Use This Skill
- Checking service health after deployment
- Investigating performance or availability issues
- Morning standup health check
- After infrastructure changes
When NOT to Use This Skill
- No services running (pure library/CLI project)
- Already have a monitoring dashboard (Grafana, Datadog, etc.)
- Development-only local services that don't need monitoring
Steps
-
Discover services — Check for:
- systemd:
systemctl --user list-units --type=service --state=running - Docker:
docker ps - PM2:
pm2 list - Kubernetes:
kubectl get pods
- systemd:
-
For each service, collect:
- Status (running/stopped/crashed)
- Uptime
- Last restart time and reason
- Recent error count (from logs)
- Resource usage (CPU, memory) if available
-
Identify issues:
- Crashed services (not running when expected)
- High-restart services (restarted >3 times recently)
- Resource-hungry processes (>80% CPU or memory)
- Error spikes (more errors than usual)
-
Provide remediation commands for each issue
Output Format
Service Health Dashboard
Services: X running, Y stopped, Z error Overall: HEALTHY / DEGRADED / CRITICAL
| Service | Status | Uptime | Restarts | Errors (1h) | CPU | Memory |
|---|---|---|---|---|---|---|
| api | ✓ Running | 3d 4h | 0 | 2 | 15% | 256MB |
| worker | ✗ Crashed | - | 5 | 38 | - | - |
Issues
- worker crashed — Last error: "OOM killed"
- Fix:
systemctl --user restart workeror increase memory limit
- Fix:
Recommended Actions
- [Action with command]