Skip to content

Metrics and Observability

ShinyHub exposes Prometheus metrics for the server process itself (the control plane), emits a structured access log for every request, and - when tracing is enabled - records control-plane spans correlated with that access log. This is separate from the per-app CPU/RAM sampling shown in the dashboard and from the per-app proxy trace buffer documented in tracing.md. Durable, human-facing app-open reporting is also separate; see Usage analytics.

Prometheus is an optional, one-way export. ShinyHub never queries Prometheus to populate its API, dashboard, autoscaling decisions, or durable usage reports. Historical peak concurrency in the Usage tab is calculated and retained by ShinyHub itself.

Memory measurement

The live app metrics API always reports RSS for a PID-backed replica. Native Linux replicas also report three nullable attribution counters read from /proc/<pid>/smaps_rollup and summed across the replica's process group:

Field Meaning Operator use
rss_bytes Resident pages mapped by the process; shared pages appear in every sharer's RSS. Working-set signal and continuity with existing dashboards.
pss_bytes Resident shared pages divided proportionally among their current sharers. Additive physical-memory attribution and pre-fork evaluation.
uss_bytes Private clean, dirty, and private huge pages. Memory that should be reclaimed when the replica exits.
swap_pss_bytes Swapped pages divided proportionally among sharers. Detect attributed swap use without double-counting.

The fields are null on unsupported hosts and remote/PID-less backends. If a process exits or its rollup cannot be read while the process group is sampled, memory_attribution_partial is true and the non-null counters are lower bounds. shinyhub top --json sums the attribution fields across running replicas and preserves the same *_partial distinction. The dashboard keeps RSS labelled as Memory and shows PSS separately; the CLI replica inspector exposes all four.

PSS is for attribution, not enforcement or admission. Per-replica limits remain cgroup/container limits, and the elastic-worker safety floor remains host MemAvailable. A PSS value can move merely because another sharer starts or stops, so using it as a hard cap would make the cap non-local and unstable.

Native workload metrics over OTLP

ShinyHub can export resource usage for each native replica and scheduled command without instrumenting the application. This export is opt-in and independent of the Prometheus listener and dashboard history:

tracing:
  enabled: true
  otlp_endpoint: http://collector:4318
  otlp_protocol: http/protobuf  # or grpc
  resource_attributes:
    deployment.environment.name: production
metrics:
  process_interval: 30s         # default 0: disabled; allowed range 1s–10m

SHINYHUB_METRICS_PROCESS_INTERVAL overrides the interval. 0 disables it. Enabling export requires tracing and its OTLP endpoint; an endpoint alone does not enable workload metrics. The collector must have an OTLP metrics pipeline, even if it already accepts traces. Endpoint protocol and authentication headers come from the platform tracing configuration, including /v1/metrics appended to the base URL for HTTP. Per-app endpoint overrides do not redirect these platform observations.

Metric OTLP type Unit Meaning
shinyhub.process.memory.usage Gauge By Summed RSS of the native launch process group, including the launcher and workers that remain in that group.
shinyhub.process.cpu.time Monotonic cumulative sum s Observed user and system CPU seconds accumulated since registration.
shinyhub.replicas Gauge {replica} Native replica processes supervised by this ShinyHub instance, per app. Includes starting, draining and frozen replicas; excludes scheduled commands.

Resources carry shinyhub.app, shinyhub.app.slug, deployment identity when known, operator resource attributes, and either shinyhub.replica or shinyhub.schedule, shinyhub.schedule.name, and shinyhub.schedule.run_id. Schedule run IDs are strings, matching the injected OTEL_RESOURCE_ATTRIBUTES; the control-plane schedule.run span uses an integer run-ID attribute, so some backends require type normalization when joining it. Environment resource attributes are decoded with the same percent-encoding as application telemetry. Explicit app resource overrides are included, while workload identity and host identity remain authoritative. Unrelated environment variables and command lines are never exported.

Each observed launch has a unique service.instance.id and a host.name. Registration after server recovery starts a new metric stream and CPU accounting interval. This distinguishes restarts and overlapping deployment generations. Replica counts use the hostname and server listener address, so a server restart does not create another count series; sum across hosts/listeners for a fleet count. Resource attributes remain resources in OTLP; Prometheus exporters may need resource-to-label conversion or a resource-info join to expose them as query labels.

CPU data points include shinyhub.process.cpu.accounting. cgroup uses an existing dedicated cgroup's kernel counter, including children that exit between samples or leave the launch process group. Prior usage of a reused cgroup is excluded. sampled retains observed per-process CPU contributions after children exit and protects against PID reuse, but misses consumption from children that start and finish between observations. Accounting mode stays fixed for the observed launch. Workload metrics do not create cgroups or require root.

RSS counts shared pages in every process that maps them; it is not additive physical-memory attribution. A failed member read omits that interval's RSS observation rather than reporting a partial sum as complete. CPU sampling can still report its observed lower bound. Detached processes outside the original process group are excluded from RSS.

An initial observation is taken at registration, and available observations for a completed run are retained for export even if it ends before the first tick. The initial RSS may represent only the launcher. Periodic samples do not guarantee a run's peak memory or complete short-run CPU usage. Cgroup CPU is read again before teardown. Completed workloads are then exported once with their actual observation timestamps, subject to retries, and no longer sampled. An app's replica count receives a zero observation after its last replica exits. Backend retention and staleness determine when old series disappear from queries.

Sampling and export run independently, with five-second network timeouts. Live observations coalesce during an outage; up to 4,096 completed workload observations are retained in memory, with oldest-first eviction and a warning when full. Transient failures retry with backoff and collector throttling hints. Permanently rejected observations are dropped; partial success is not retried. There is no durable telemetry spool. A bounded flush runs on shutdown.

Run IDs create new historical series: a 15-minute schedule creates 192 series per day for these two resource metrics, before additional dimensions. Choose backend retention accordingly. This version covers native workloads on the ShinyHub host; Docker and remote runtimes are not sampled by this exporter.

The /metrics endpoint

Metrics are opt-in and served on their own listener, separate from the main application port so server internals are never exposed on a routable interface by accident:

metrics:
  enabled: true
  addr: "127.0.0.1:9090"   # default when enabled and unset

Environment overrides (last-wins over YAML):

YAML field Environment variable
enabled SHINYHUB_METRICS_ENABLED
addr SHINYHUB_METRICS_ADDR

The endpoint defaults to loopback. Operators scraping from another host set addr to a private interface behind their own network controls (the conventional Prometheus pattern). When enabled: false no handler and no listener are created.

Scrape it like any Prometheus target:

scrape_configs:
  - job_name: shinyhub
    static_configs:
      - targets: ["shinyhub-host:9090"]

Exposed series

Process and build

Metric Type Description
shinyhub_build_info{version} gauge Always 1; the build version is a label.
shinyhub_uptime_seconds gauge Seconds since the server started serving.
go_*, process_* various Standard Go runtime + process collectors (heap, goroutines, server RSS/CPU/FDs).

Control-plane HTTP

Labeled by the matched chi route pattern (not the raw path), so high-cardinality path parameters and unmatched 404 scans cannot explode the series count.

Metric Type Labels Description
shinyhub_http_requests_total counter method, route, status Control-plane HTTP requests.
shinyhub_http_request_duration_seconds histogram method, route, status Control-plane request latency.

Data-plane admission

Metric Type Labels Description
shinyhub_admission_rejects_total counter slug, reason Proxy admission rejections. slug is __unknown__ for requests to slugs that are not registered apps.
shinyhub_app_sessions gauge slug Active proxied sessions for an app, summed across live replicas (evaluated at scrape time).
shinyhub_app_sessions_limit gauge slug Admission ceiling for an app: the number of replicas that admit new sessions (live, not draining) times the per-replica session cap. Absent for uncapped apps, so shinyhub_app_sessions / shinyhub_app_sessions_limit is the saturation fraction wherever a cap applies.
shinyhub_ws_session_ends_total counter slug, closed_by, transport_end_side, abnormal Successfully hijacked WebSocket tunnels that ended. closed_by identifies the first observed close-frame sender or a known proxy action; it is unknown when only a transport end was observed. transport_end_side records the first side whose read ended. abnormal=true means an upstream close frame or upstream transport end without a 1000/1001 close code.
shinyhub_ws_abnormal_bursts_total counter slug Worker-local groups of at least three abnormal WebSocket endings within ten seconds, with a 30-second warning cooldown. The counter is local to each ShinyHub instance.

Each completed tunnel also emits one structured ws_session_end log with its connection_id, replica, deployment, duration, close code and reason when observed, and bytes written in each tunnel direction. end_signal distinguishes a close frame, transport end, proxy action, or unknown ending. Upstream abnormal endings are WARN; other endings are INFO. A transport EOF does not establish why the worker stopped responding. This counter measures the same local end events as the logs; it cannot prove that a specific log record reached an external log store. When a known lifetime stop or generation cleanup ends a connection, the event uses closed_by=lifetime or closed_by=drain and stays at INFO. A ws_abnormal_burst WARN names the app, deployment, worker slot, and number of abnormal endings in its window, plus their observed time span. It includes a close code and reason only when all endings in that window agree. It reports correlation, not a CPU-stall diagnosis. For Shiny's /websocket/ connection, the injected interruption overlay sends a random connection ID on the upgrade and shows that same ID after a connected session drops. The proxy removes the tag before forwarding the request to the app. If the script loads after the socket opened or browser crypto is unavailable, the overlay omits the ID rather than risk showing a different session's ID; the server still generates and logs one.

The reason label is a closed vocabulary. The same value is returned on the X-Shinyhub-Reject response header, so a rejected request can be traced from the client back to this counter. Reasons differ in what they mean you should do, which is why they are not collapsed into one:

reason What happened Remedy
unknown-slug No app with this slug is registered (404). Nothing, unless you expected the app to exist. A rising rate is usually scanning.
pool-saturated Every replica is live and at its per-replica session cap. Raise --max-sessions-per-replica and/or --replicas. This is the only scale-up signal here.
pool-degraded Fewer replicas are registered than configured, and the survivors are at cap. Check replica health first. Adding capacity on top of a crash loop hides it.
app-not-ready The app has no replica that has completed a WebSocket handshake yet. Nothing during a normal cold start. Sustained means the app is failing to come up.
memory-pressure The host is below server.min_available_memory_mb, so no new elastic worker may start. Free host memory, lower per-app ceilings, or add hardware.
render-paced A new session was deferred because the app's render-admission bucket was empty, then shed after the park window. Raise the app's render_seconds accuracy or add cores. More replicas do not help: they do not add CPU.
cpu-saturation The host CPU watermark is breached, so a new session was shed to protect connected ones. Add cores or move apps off this host.
render-deferred A page load was shown the "Waiting for capacity" page because the app had no render capacity at that instant. Same as render-paced. See the caveat below before alerting on it.
replica-starting A request could not be forwarded because replica readiness has not completed. Non-document clients receive 503 with Retry-After. Allow startup to finish; inspect startup logs if it persists.

Readiness observations (app-not-ready) and expected startup waits (replica-starting) retain their labelled counters but are excluded from the dashboard's ten-minute admission-issues rollup. A probe poll is not a refused user session. Failed boots still appear as crashed or degraded apps, and an upstream failure after readiness remains a proxy_upstream_error warning.

proxy_access records the actual downstream status. When ShinyHub serves a starting, deploying, stopped, or crashed page instead of app content, it also records fallback: true and fallback_reason. Browser starting pages can return 200; an unreachable previously ready upstream is identified by fallback_reason: "upstream-error" rather than appearing as an unqualified success.

render-deferred counts page loads deferred, not sessions refused, and it is inflated by design: one waiting browser re-polls roughly every 1.75 s until capacity frees, so a single user can contribute dozens of increments. Use it to see that users are waiting; use render-paced to count sessions actually turned away. Alerting on render-deferred as if it were a refusal rate will page you for one patient user.

Both session gauges are exported per control-plane instance, like every metric here. On a single-node deployment they are exact. In a clustered deployment, scrape every instance and aggregate in PromQL (sum by (slug) (...)) rather than reading one instance in isolation - the example alert below already does this.

Usage analytics durability

Metric Type Labels Description
shinyhub_usage_persistence_events_total counter result Exceptional durable-usage outcomes: start_overflow, start_retry, start_failed, start_dropped, end_retry, or policy_refresh_failed. Overflow and retries are recovered automatically; failed or dropped starts mean the Usage report may undercount connections. A policy-refresh failure leaves the cached connection fast path conservative while every insert remains clamped against durable policy.

This counter is process-local. In HA, sum it across every control-plane instance. An increase in start_failed or start_dropped is the completeness alert; the retry and overflow results are useful early warnings.

Fleet and lifecycle

Metric Type Labels Description
shinyhub_apps_running gauge - Apps currently in the running state (evaluated at scrape time).
shinyhub_replicas_running gauge - App replicas currently running (evaluated at scrape time).
shinyhub_deploys_total counter result Deployments by outcome (success / failure). Alert on a rising failure rate.
shinyhub_generation_handoffs_total counter outcome Generation handoff events from a closed vocabulary: success, deferred, explicit_downtime, candidate_failure, cutover_repair, forced_retirement, or cleanup_deferred.
shinyhub_generation_draining gauge - Old generations currently draining on this control-plane process.
shinyhub_generation_draining_sessions gauge - Open HTTP streams and WebSockets still attached to draining generations on this control-plane process.
shinyhub_generation_drain_duration_seconds histogram - Time from cutover until the old route retired or cleanup was deferred.
shinyhub_app_state_transitions_total counter event App lifecycle transitions (hibernate, wake).
shinyhub_replica_restarts_total counter - Replica crash-restarts performed by the watchdog. A flapping app shows up as a rising restart rate.
shinyhub_schedule_last_success_seconds gauge slug, schedule Unix timestamp of the most recent successful command run; absent until the first success.
shinyhub_schedule_activation_status gauge slug, schedule, status Current durable serving-data activation state; the labeled state has value 1. Alert on repairing, deferred_capacity, failed, cancelled, or blocked_unsupported according to age and policy.
shinyhub_schedule_activation_age_seconds gauge slug, schedule Age of a nonterminal activation. Use with the status metric so ordinary interval damping is not mistaken for a fault.
shinyhub_schedule_activation_target_generation gauge slug, schedule Target serving-data generation of the latest activation.

Application-log durability and delivery

These process-local signals cover the shared-log path used by clustered Postgres deployments. They deliberately carry no app, replica, or run labels: those identities are unbounded and belong in the log viewer, not in Prometheus series. Scrape every control-plane instance and aggregate with sum(...) in HA.

Metric Type Labels Description
shinyhub_app_log_flush_attempts_total counter result Shared-log database flush attempts (ok / error).
shinyhub_app_log_flush_duration_seconds histogram - Database-call duration for every shared-log flush attempt.
shinyhub_app_log_persistence_lag_seconds histogram - Time from entering the retry buffer to a successful database flush. Failed attempts do not produce a lag sample.
shinyhub_app_log_pending_bytes gauge - Bytes currently queued for persistence by this process, excluding the one chunk in flight. A brief non-zero value is normal because writes are batched.
shinyhub_app_log_buffer_dropped_bytes_total counter - Bytes evicted from the bounded retry buffer, including bytes abandoned when the final shutdown flush fails. Any increase means shared history may be incomplete.
shinyhub_app_log_runs_pruned_total counter - Immutable run records removed from database retention by the owner.
shinyhub_app_log_files_pruned_total counter - Orphaned immutable files removed from this node's private disk.
shinyhub_app_log_followers gauge - Active per-run shared database followers on this process. Compare with shinyhub_app_log_viewers to confirm concurrent viewers are sharing followers.
shinyhub_app_log_viewers gauge - Active retained-log viewer subscriptions on this process.
shinyhub_app_log_follow_errors_total counter - Failed database reads by shared retained-log followers. A rise means live delivery is degraded; durable history remains available for catch-up once reads recover.

The flush and buffer metrics remain present at zero on single-node SQLite deployments, where the viewer reads its local files directly. Retention cleanup counters can still rise there. Like all Prometheus counters, cleanup and loss totals reset when a ShinyHub process restarts.

After a follower read fails, its polling interval backs off exponentially with jitter to a five-second ceiling. A successful read restores the normal 200 ms cadence, while a locally committed log chunk wakes the follower immediately. This bounds database pressure during an outage without delaying healthy local delivery.

Example alerts

groups:
  - name: shinyhub
    rules:
      - alert: ShinyHubDeployFailures
        expr: increase(shinyhub_deploys_total{result="failure"}[15m]) > 0
        annotations:
          summary: "A ShinyHub deploy failed in the last 15m"

      - alert: ShinyHubGenerationForcedRetirement
        expr: sum(increase(shinyhub_generation_handoffs_total{outcome="forced_retirement"}[15m])) > 0
        annotations:
          summary: "An old app generation exceeded its drain deadline"

      - alert: ShinyHubGenerationCleanupDeferred
        expr: sum(increase(shinyhub_generation_handoffs_total{outcome="cleanup_deferred"}[15m])) > 0
        annotations:
          summary: "Old generation process cleanup was deferred to startup recovery"

      - alert: ShinyHubGenerationDrainStuck
        expr: sum(shinyhub_generation_draining) > 0
        for: 10m
        annotations:
          summary: "An app generation has been draining for more than 10m"

      - alert: ShinyHubExplicitDowntimeDeploy
        expr: sum(increase(shinyhub_generation_handoffs_total{outcome="explicit_downtime"}[15m])) > 0
        annotations:
          summary: "A deploy used the explicit stop-first fallback"

      - alert: ShinyHubReplicaFlapping
        expr: increase(shinyhub_replica_restarts_total[10m]) > 5
        annotations:
          summary: "A ShinyHub replica is crash-restarting repeatedly"

      - alert: ShinyHubSessionsNearCap
        # sum by (slug) aggregates across control-plane instances; on a single
        # node it is simply the one series.
        expr: sum by (slug) (shinyhub_app_sessions) / sum by (slug) (shinyhub_app_sessions_limit) > 0.9
        for: 5m
        annotations:
          summary: "{{ $labels.slug }} is above 90% of its admission ceiling"

      - alert: ShinyHubUsageHistoryIncomplete
        expr: sum(increase(shinyhub_usage_persistence_events_total{result=~"start_failed|start_dropped"}[10m])) > 0
        annotations:
          summary: "Durable app-usage history may be incomplete"

      - alert: ShinyHubAppLogPersistenceErrors
        expr: sum(increase(shinyhub_app_log_flush_attempts_total{result="error"}[10m])) > 0
        annotations:
          summary: "Shared app-log persistence has failed"

      - alert: ShinyHubAppLogBacklogStuck
        expr: sum(shinyhub_app_log_pending_bytes) > 0
        for: 5m
        annotations:
          summary: "Shared app-log bytes have remained queued for 5m"

      - alert: ShinyHubAppLogDataDropped
        expr: sum(increase(shinyhub_app_log_buffer_dropped_bytes_total[5m])) > 0
        annotations:
          summary: "Shared app-log history may be incomplete"

      - alert: ShinyHubAppLogFollowErrors
        expr: sum(increase(shinyhub_app_log_follow_errors_total[10m])) > 0
        annotations:
          summary: "Shared app-log live delivery has encountered database errors"

Access log

Every request emits one structured api_access record through log/slog (replacing chi's unstructured stock logger), so the control plane has a single structured log stream a log aggregator can ingest. Fields:

  • request_id - per-request correlation ID (see below)
  • method, path, route (matched pattern), status
  • bytes, duration_ms
  • client_ip - the trusted-proxy-aware client IP (honest even when ShinyHub sits behind an edge proxy; see server.trusted_proxies)
  • trace_id - present when tracing is enabled and a span is active

Request-ID correlation

Each request is assigned a correlation ID echoed on the response as the X-Request-Id header and threaded through the request context so downstream handlers tag their own logs with the same ID. A well-formed inbound X-Request-Id (e.g. minted by a trusted edge proxy) is honored so a request stays correlated across tiers; a malformed or oversized value is rejected and replaced, closing a log- and header-injection vector.

Log <-> trace correlation

When server tracing is enabled (see tracing.md), the access log and the trace are linked in both directions:

  • the api_access record carries the active span's trace_id, so you can pivot from a log line to the trace, and
  • the server span carries the request_id attribute, so you can pivot from a trace back to the log line.

Server (control-plane) tracing

Enabling tracing also instruments the control-plane API: ShinyHub emits one server span per request and spans for background lifecycle operations (lifecycle.wake, lifecycle.restart, lifecycle.hibernate, tagged with shinyhub.app.slug), exported to the same OTLP endpoint the managed apps use. Spans use OpenTelemetry HTTP semantic-convention attribute names (http.request.method, http.route, http.response.status_code) and carry a resource identifying the instance (service.name, service.version, service.instance.id). An inbound traceparent is adopted as the parent, so a client/edge trace links through ShinyHub to the app it proxies.

This reuses the existing tracing config block; there is no separate server-tracing switch. See tracing.md for the configuration fields and the per-app proxy trace buffer.

Database connection pool

Metric Type Meaning
shinyhub_db_open_connections gauge Open connections, including idle connections.
shinyhub_db_in_use_connections gauge Connections currently borrowed by database operations.
shinyhub_db_max_open_connections gauge Configured pool limit; zero means unlimited.
shinyhub_db_wait_count_total counter Connection acquisitions that had to wait for the pool.
shinyhub_db_wait_duration_seconds_total counter Cumulative time waiting to acquire a connection.

These metrics observe the server's actual pool and carry no DSN or query-text labels. Wait time excludes query execution and SQLite lock contention. Compare counter increases over the same interval as request latency and CPU usage; connection waits alone do not identify a slow query.