OpenTelemetry Tracing¶
ShinyHub propagates W3C trace context through its reverse proxy and injects the OTEL_* environment variables every app process needs to export its own spans to your OpenTelemetry collector. Apps export their spans directly to the collector (ShinyHub never sees the bytes) and the Traces tab in the UI shows a per-app ring buffer of recent slow or failed proxy spans, deep-linkable into your backend (Tempo, Jaeger, Honeycomb, etc.).
This keeps ShinyHub a single binary with no embedded OTLP receiver: the operator picks the backend, the apps export, and ShinyHub just propagates and surfaces.
How it works¶
client ──► ShinyHub proxy ──► Shiny app process
│ │
│ └──► OTLP collector (operator-owned)
│ │
└────────────────────────────┘
traceparent header full app spans
flows end-to-end delivered directly
For every proxied request, ShinyHub:
- Parses any incoming
traceparentheader. Missing or malformed headers start a new trace; valid ones continue it with a fresh span ID for the proxy hop. - Sets
traceparenton the upstream request so the app sees ShinyHub's span as its parent. Shiny for Python's built-in OpenTelemetry support then reports a single connected trace. - Records the proxy-level span (method, path, status, duration, replica, sampled flag) into a per-app ring buffer if the request was slow, returned 5xx, or errored.
- Drops everything else; the buffer never grows beyond
ring_buffer_sizespans per app.
The sampling decision uses W3C parent-based traceidratio: child spans honor
the parent's sampled flag, and roots fall under sample_ratio of all
traces.
Configuration¶
Enable tracing in shinyhub.yaml:
tracing:
enabled: true
otlp_endpoint: http://collector.observability.svc:4318
otlp_protocol: http/protobuf # or "grpc"
otlp_headers: "x-api-key=secret" # optional, for hosted backends
sample_ratio: 0.1 # 10% of new traces
slow_request_ms: 1000 # slow-threshold for buffer admission
ring_buffer_size: 200 # spans retained per app
trace_link_template: "https://tempo.example.com/explore?trace={trace_id}"
auto_instrument_apps: false # wrap supported Python apps, jobs and hooks
asgi_events: false # omit low-level ASGI send/receive spans
auto_instrument_extra_packages: # extras for the auto-instrument overlay
- opentelemetry-instrumentation-botocore
resource_attributes: # tags added to every span, here and in apps
deployment.environment.name: production
Every field has an env-var override (last-wins over YAML):
| YAML field | Environment variable |
|---|---|
enabled |
SHINYHUB_TRACING_ENABLED |
otlp_endpoint |
SHINYHUB_TRACING_OTLP_ENDPOINT |
otlp_protocol |
SHINYHUB_TRACING_OTLP_PROTOCOL |
otlp_headers |
SHINYHUB_TRACING_OTLP_HEADERS |
sample_ratio |
SHINYHUB_TRACING_SAMPLE_RATIO |
slow_request_ms |
SHINYHUB_TRACING_SLOW_REQUEST_MS |
ring_buffer_size |
SHINYHUB_TRACING_RING_BUFFER_SIZE |
trace_link_template |
SHINYHUB_TRACING_TRACE_LINK_TEMPLATE |
auto_instrument_apps |
SHINYHUB_TRACING_AUTO_INSTRUMENT_APPS |
asgi_events |
SHINYHUB_TRACING_ASGI_EVENTS |
auto_instrument_extra_packages |
SHINYHUB_TRACING_AUTO_INSTRUMENT_EXTRA_PACKAGES |
resource_attributes |
SHINYHUB_TRACING_RESOURCE_ATTRIBUTES |
auto_instrument_extra_packages requires auto_instrument_apps; its env form
is whitespace-separated PEP 508 requirements (name, optional [extras],
optional version specifiers; no spaces inside one requirement, no URLs):
SHINYHUB_TRACING_AUTO_INSTRUMENT_EXTRA_PACKAGES="opentelemetry-instrumentation-botocore opentelemetry-instrumentation-redis>=0.40".
resource_attributes keys must start with a letter, followed by any mix of
letters, digits, ., _, or - (up to 128 characters); the env form
is the same k=v,k2=v2 wire format as OTEL_RESOURCE_ATTRIBUTES, with values
percent-encoded (unreserved-character set per RFC 3986) so separators and
non-ASCII survive decoding: SHINYHUB_TRACING_RESOURCE_ATTRIBUTES="deployment.environment.name=production,team=data-eng".
Setting the env var replaces the whole YAML map rather than merging into it.
service.name, service.version, service.instance.id, and any key
prefixed shinyhub. are reserved (ShinyHub sets them itself) and rejected at
startup; see Multiple instances, one backend.
Defaults applied when enabled: true and the field is unset:
otlp_protocol:http/protobufsample_ratio:0.1slow_request_ms:1000ring_buffer_size:200
Environment variables injected into each app¶
When tracing is enabled, every app replica is launched with:
OTEL_SERVICE_NAME=<app-slug>
OTEL_RESOURCE_ATTRIBUTES=shinyhub.app=<slug>,shinyhub.app.slug=<slug>,shinyhub.replica=<index>[,<your resource_attributes, sorted by key>],shinyhub.deployment.id=<id>,service.version=<app version or bundle digest>
OTEL_EXPORTER_OTLP_ENDPOINT=<your collector>
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf | grpc
OTEL_TRACES_SAMPLER=parentbased_traceidratio
OTEL_TRACES_SAMPLER_ARG=<sample_ratio>
OTEL_PYTHON_STARLETTE_EXCLUDED_URLS=/websocket/?$
OTEL_EXPORTER_OTLP_HEADERS=<headers if configured>
SHINYHUB_TRACING_ASGI_EVENTS=false | true
tracing.resource_attributes pairs are appended after the built-in
app identity attributes, sorted by key, and percent-encoded the
same way as the config env override above. OTEL_PYTHON_STARLETTE_EXCLUDED_URLS
keeps Shiny's session WebSocket out of the auto-instrumented Starlette spans
(see What you get, and what you don't
below); it is set for every app whenever tracing is enabled, whether or not
auto-instrumentation is on, since it is harmless for an app that reads no
OTEL_* vars.
Deployment attributes come from the process's own launch parameters, including
rollback and overlapping deployment generations. service.version uses the app
version, falling back to its content digest; unknown deployment/version values
are omitted. service.instance.id remains the SDK's unique process identity.
The existing shinyhub.app resource attribute remains available alongside
shinyhub.app.slug for compatibility.
These are platform defaults. Per-app env vars (set via UI or
PUT /api/apps/<slug>/env/<KEY>) win on duplicate keys, so any app can
override the collector endpoint, service name, sampler, or headers
independently. The SHINYHUB_ prefix is the only reserved namespace;
OTEL_* is intentionally user-settable. OTEL_RESOURCE_ATTRIBUTES is merged
by attribute: app values override matching keys while unrelated platform
identity and fleet tags remain. Secret resource values remain in the runtime's
secret environment. An empty resource variable does not clear platform tags.
Multiple instances, one backend¶
Point every ShinyHub instance that shares a fleet (a zero-downtime upgrade
pair, staging vs. production, or two unrelated fleets sharing a collector) at
the same OTLP endpoint and tell them apart in the backend with
tracing.resource_attributes:
Give each instance a value that differs (environment, region, cluster) and
filter or group on it in Tempo/Grafana. The pairs land in three places: every
app replica's OTEL_RESOURCE_ATTRIBUTES (above), every scheduled job run's
(see Scheduled jobs below), and the resource of every span
the server process itself exports.
service.instance.id on the server's own spans is server.instance_id
(shinyhub.yaml only - there is no SHINYHUB_SERVER_INSTANCE_ID env
override), which defaults to <hostname>-<pid>. That default changes on
every restart, so if you want to correlate a server's spans across restarts,
or need a stable identity for a specific host rather than one process's
lifetime, set server.instance_id explicitly.
If apps or jobs run behind an outbound HTTP proxy (HTTP_PROXY/HTTPS_PROXY
set in their environment), add the collector's host to NO_PROXY on those
hosts so the OTLP exporter reaches it directly instead of through the proxy.
Auto-instrumentation (zero-config app spans)¶
With one fleet-level flag, every Python app gets transport-layer spans with
no change to its pyproject.toml, requirements.txt, or run command:
tracing:
enabled: true
otlp_endpoint: http://collector.observability.svc:4318
auto_instrument_apps: true # default: false
Env override: SHINYHUB_TRACING_AUTO_INSTRUMENT_APPS. Individual apps opt in
or out against the fleet default in their bundle's shinyhub.toml:
The override travels with the bundle: it is re-read at every boot (deploy, crash restart, hibernation wake) and applies per deployed version, including rollbacks.
ShinyHub already injects the OTEL_* env vars and propagates traceparent;
auto-instrumentation adds the remaining two pieces. The app is launched as
uv run [--with-requirements requirements.txt] \
--with opentelemetry-distro \
--with opentelemetry-exporter-otlp \
--with opentelemetry-instrumentation-starlette \
--with opentelemetry-instrumentation-requests \
--with opentelemetry-instrumentation-httpx \
[--with <tracing.auto_instrument_extra_packages>] \
opentelemetry-instrument python -c <ShinyHub bootstrap> app-module shiny run app.py --host ... --port ...
The bootstrap executes shiny as a module under the overlay's
interpreter; the app's own shiny console script would run under its own
environment's interpreter, which cannot see the overlay packages.
uv's --with overlay resolves these packages alongside the app's own
dependencies without modifying its venv or lockfile; turn the flag off and
the overlay is gone. tracing.auto_instrument_extra_packages (a fleet-wide
setting) is layered in after the built-in set for instrumentors the built-in
list doesn't cover, such as opentelemetry-instrumentation-botocore.
This applies identically to a pool replica's initial boot and to an elastic worker spawned on demand: both resolve the launch command through the same seam, honor the same fleet default and manifest override, and get the same uninstrumented retry (below) if the instrumented launch fails.
ASGI send/receive spans are suppressed by default. The shared bootstrap applies
exclude_spans=["receive", "send"] before the app constructs its middleware;
Starlette's instrumentor does not expose that middleware option itself. Set
tracing.asgi_events: true to restore the events for debugging. This keeps
request and library spans, and does not remove Shiny's reactive/output spans.
Explicit app middleware exclude_spans lists remain authoritative; an empty
list restores all events for that middleware.
What you get, and what you don't¶
- Transport-layer spans for free. Shiny for Python runs on Starlette
(ASGI), so each request gets a server span that nests under ShinyHub's
propagated trace context, and outbound
requests/httpxcalls become client spans. This closes the trace at the request boundary: "slow at the proxy" becomes "slow inside the app's HTTP hop". - The reactive graph, from Shiny itself. Shiny for Python 1.6+ ships its
own OpenTelemetry instrumentation: once the SDK is wired (by
auto-instrumentation, or manually per below), Shiny emits
session_start,reactive_update,output <id>, andreactive.calc <name>spans under the instrumentation scopeco.posit.python-package.shiny. With the default WebSocket exclusion (below) the session has no request span to nest under, so eachsession_startandreactive_updateis the root of its own trace; overriding the exclusion per app nests them under the session's WebSocket span, at the cost of one span that lasts the whole session. Verbosity is controlled per app withSHINY_OTEL_COLLECT; Posit suggestssessionin production andreactive_updatein staging. This covers renders and calc/effect invalidation without writing a single span by hand - reach for manual spans (next section) for your own library code inside a calc or output, which Shiny's instrumentation cannot see into. - WebSocket span excluded by default. Shiny holds one long-lived
WebSocket per session, so an uncontrolled auto-instrumented WS span would be
one long, low-signal span for the whole visit. ShinyHub sets
OTEL_PYTHON_STARLETTE_EXCLUDED_URLS=/websocket/?$in every app's environment whenever tracing is enabled, so the Starlette instrumentor skips it. Override the pattern per app the same way as any otherOTEL_*default (table below) - set it to an empty string to trace the WebSocket after all (which also nests Shiny's reactive spans under it, above), or to a different regex. - Logs and metrics too, unless you turn them off. The
opentelemetry-distrothat auto-instrumentation launches defaultsOTEL_LOGS_EXPORTERandOTEL_METRICS_EXPORTERtootlpas well as traces, and sends them to the sameOTEL_EXPORTER_OTLP_ENDPOINT. A collector with only a traces pipeline rejects those exports (overhttp/protobuf, a 404 on/v1/logsand/v1/metrics), and the exporter reports each rejection in the app's log. The spans are unaffected. To stop it for an app, setOTEL_LOGS_EXPORTER=noneand/orOTEL_METRICS_EXPORTER=nonein that app's env.
Failure semantics¶
Instrumentation can never take an app down. If the overlay cannot resolve,
or the wrapped process crashes at startup (for example the app pins an old
opentelemetry-api that breaks opentelemetry-instrument's imports), or it
fails its health check, ShinyHub retries the boot uninstrumented and logs
a warning (instrumented launch failed; retrying without
auto-instrumentation) in the server log; the uv resolution error or Python
traceback is visible in the app's own log. Persistent offenders should set
[tracing] auto = false in their manifest. Note the failed instrumented
attempt costs up to one health-check timeout before the fallback boots.
Scope and caveats:
- Python only. R apps (
app.R/Rscript) are never wrapped; there is noopentelemetry-instrumentequivalent for R. - Inferred commands only. Deploys that supply a custom command are never wrapped; wrap your own command if you need both.
- Docker runtime: the overlay resolves inside the container at start, so the first start (and starts after image replacement) download the OTEL packages; subsequent starts hit uv's cache only if you persist it. Budget a few extra seconds of cold start, including hibernation wakes.
Tracing your app¶
Two layers of per-app control sit on top of auto-instrumentation. Both
assume auto_instrument_apps (or the app's [tracing] auto = true).
Layer 1 - config knobs, no code. Per-app env vars win over the injected
platform defaults, so tuning is a few settings (UI → app → Configuration, or
PUT /api/apps/<slug>/env/<KEY>):
| Env var | Effect |
|---|---|
OTEL_TRACES_SAMPLER_ARG=1.0 |
Sample this app harder than the fleet sample_ratio |
OTEL_PYTHON_STARLETTE_EXCLUDED_URLS=<regex> |
Change which paths the Starlette instrumentor skips; the platform default excludes only the session WebSocket |
OTEL_PYTHON_DISABLED_INSTRUMENTATIONS=starlette |
Drop all Starlette/ASGI spans, including the per-request server spans, not just the WebSocket; most apps want the exclusion above instead |
OTEL_RESOURCE_ATTRIBUTES=team=analytics,owner=data-eng |
Ownership tags on every span. Merges with the platform resource attributes; explicit matching keys override defaults, while deployment/replica identity and fleet tags remain |
OTEL_SERVICE_NAME=my-name |
Override the default service name (the app slug) |
OTEL_EXPORTER_OTLP_ENDPOINT=... |
Send this app's spans to a different collector |
OTEL_LOGS_EXPORTER=none, OTEL_METRICS_EXPORTER=none |
Stop the distro exporting logs and metrics, for a collector that only accepts traces |
Layer 2 - custom spans in two lines. opentelemetry-instrument has
already wired the global TracerProvider, the OTLP exporter, and incoming
traceparent extraction, so an app adds its own spans with just:
from opentelemetry import trace
tracer = trace.get_tracer(__name__)
def load_cluster_data():
with tracer.start_as_current_span("load_cluster_data"):
return load_data() # nests under the ASGI request span, exports for free
This is where the reactive-graph gap closes: wrap your heavy data loads, renders, and calcs by hand and they appear inside the request trace.
The one footgun: rely on the auto-configured global provider - call
trace.get_tracer(...) and emit. Do not call
trace.set_tracer_provider(...) yourself; that double-initialises the SDK
and breaks export.
Manual instrumentation (without auto-instrumentation)¶
If the fleet flag is off and the app cannot opt in, the pre-existing route
still works: add opentelemetry-distro, opentelemetry-exporter-otlp, and
the instrumentors to the bundle's own dependencies and deploy with a custom
command that wraps shiny run in opentelemetry-instrument. The injected
OTEL_* env vars apply either way. Posit's guide:
https://shiny.posit.co/py/docs/opentelemetry.html
The Traces tab¶
UI → App detail → Traces polls GET /api/apps/<slug>/traces every 5
seconds and shows the most recent slow or failed proxy spans, newest first.
The buffer is in-memory and per-process, so it resets on ShinyHub restart and
holds at most ring_buffer_size spans per app.
Each row shows:
- When the request started
- Method / Path (the path after stripping the
/app/<slug>prefix) - Status (HTTP status from the backend)
- Duration (ms)
- Replica index that handled the request
- Trace: the short trace ID, with a link to your backend if
trace_link_templateis configured ({trace_id}is replaced with the full 32-hex trace ID).
API¶
GET /api/apps/<slug>/traces uses the same auth model as /metrics (any user who
can view the app):
{
"enabled": true,
"trace_link_template": "https://tempo.example.com/explore?trace={trace_id}",
"spans": [
{
"trace_id": "0af7651916cd43dd8448eb211c80319c",
"span_id": "b7ad6b7169203331",
"parent_id": "00f067aa0ba902b7",
"app_slug": "my-app",
"replica": 0,
"method": "GET",
"path": "/session/abc/dataobj",
"status": 502,
"duration_ms": 1843,
"started_at": "2026-05-13T10:34:01Z",
"sampled": true,
"error": "context canceled"
}
]
}
When tracing is disabled the endpoint still returns 200 with enabled:
false and an empty spans: [] so the UI can render an "off" state without
extra error handling.
Server-side spans (control plane)¶
The propagation and ring buffer above cover the proxy hot path and the
app processes. Separately, when tracing.enabled is set, ShinyHub's own
server process exports spans through the OpenTelemetry SDK to the same OTLP
endpoint, so a client/edge trace links through ShinyHub to the app it
proxies, and background work (deploys, wakes, restarts) is visible even when
no client is watching.
- One server span per control-plane API request, named by the matched
route pattern (not the raw path, so cardinality stays bounded), e.g.
POST /api/apps/{slug}/deploy. Spans use HTTP semantic-convention attributes (http.request.method,http.route,http.response.status_code) and adopt an inboundtraceparentas the parent. - One span per
/app/<slug>/...proxy request, namedGET /app/{slug}(method varies, route is always the pattern). Attributes:http.request.method,http.route,url.path(the path after stripping the/app/<slug>prefix),http.response.status_code,shinyhub.app.slug, and, when known,shinyhub.replica,shinyhub.deployment.id, andshinyhub.proxy.reject_reason. It adopts an inboundtraceparentas its parent and uses the same sampling decision (ParentBased(TraceIDRatioBased(sample_ratio))) as everything else, so the trace ID and sampled flag it hands to the app in the outboundtraceparentare the exported span's own: the app's spans always find their parent in the backend. For an accepted WebSocket upgrade the HTTP span ends once the 101 headers have been flushed. The trace buffer also records only this handshake duration. Rejected upgrades retain normal HTTP spans. - A separate WebSocket session span, named
WS /app/{slug}, is a child of the handshake span. It is an internal span with nohttp.route, so its duration does not enter HTTP route latency aggregates. It recordsshinyhub.ws.bytes_to_client,shinyhub.ws.bytes_to_upstream, close code when available, close initiator, transport-ending side, and abnormal status. Application-provided close reasons are not copied into traces. Existing WebSocket lifecycle metrics and access logs retain session duration. - Background lifecycle spans for the watchdog's wake, restart, and
hibernate operations (
lifecycle.wake,lifecycle.restart,lifecycle.hibernate), each tagged withshinyhub.app.slug, so cold-start latency and restart storms are visible in the backend.lifecycle.wakeadditionally carriesshinyhub.wake.trigger, one of: request- a proxied request found the app hibernated. The span is a child of that request's/appproxy span above, and inherits its sampled flag: an inboundtraceparentwith the sampled bit unset suppresses the wake span too, along with everything nested under it.reconcile- the watchdog's periodic reconciler found an app stuckwaking(e.g. the instance that started the wake died mid-flight) and resumed driving it. This is a trace root; there is no request to parent on.
There is no third trigger: warm restore at startup (re-freezing hibernated
apps back to their pre-restart state) does not go through this wake
machinery at all, so its replica boots below are roots, not
lifecycle.wake children.
- Deploy phase spans, opened as internal spans nested by call structure:
- deploy.receive, deploy.extract, deploy.validate, deploy.handoff -
API-layer phases of a deploy/rollback/restart request, each a direct
child of the control-plane request span (siblings of each other and of
deploy.run below, not nested inside one another).
- deploy.run - the root of one pool deploy (all replicas of one app
version), tagged shinyhub.deploy.replicas (count) and
shinyhub.deploy.prepare_only. A deploy/rollback/restart submitted
through the API is its child (via traceDeploy with the request's
context, so it inherits that request's trace and sampling). Two triggers
run with no live request and are trace roots instead: a replica-count,
placement, resource-limit, or worker-isolation change from
PATCH /api/apps/<slug> launches its redeploy in a detached goroutine
that carries no request context; and scheduled activation's disruptive
capacity-fallback restart (stopping and rebooting the whole pool when a
surge-based roll cannot find capacity) runs from the background
activation coordinator loop, which likewise carries no span. A
schedule's own deploy_trigger = "bundle_change" reconvergence never
reaches this code path at all: it reruns the producer command as an
ordinary job run (a schedule.run span, see
Scheduled jobs below), not a deploy. Scaling and
warm-pool operations boot replicas one at a time and never open a
deploy.run at all - see the standalone case below.
- deploy.build - building the bundle's environment (dependency
install), once per deploy regardless of replica count.
- deploy.post_deploy_hooks - the bundle's post-deploy hook run.
- deploy.replica (one per replica, tagged shinyhub.replica) -
booting one replica.
- deploy.replica.start - launching the process.
- deploy.replica.health - polling until healthy, tagged
shinyhub.deploy.health_timeout_ms.
- Outside a pool deploy, deploy.replica (and its .start/.health
children) also appears standalone, with no deploy.run wrapper:
- Under lifecycle.restart - the watchdog's crash-restart path boots
exactly the one replica that died.
- Under lifecycle.wake - waking a suspended app first tries an
abbreviated resume (deploy.replica tagged
shinyhub.deploy.resume=true); if that fails, it falls back to a full
cold boot. A failed-resume-then-cold-boot wake therefore produces
two deploy.replica spans nested under the same lifecycle.wake:
the resume attempt (recorded as an error) followed by the cold boot.
- At startup, warm restore re-adopts every replica that was warm
when the server last stopped, one deploy.replica per replica. This
does not go through the wake machinery at all (see above), so these
spans are trace roots, not children of a lifecycle.wake.
- Scale up/down and warm-pool operations (growing or shrinking a
pool, thawing a suspended replica to keep it in the warm floor) boot or
thaw replicas individually the same way, each producing its own
standalone deploy.replica span rather than a shared deploy.run.
- Scheduled activation's normal roll boots the same way: it surges in
one replica of the new generation, then cuts canonical replicas over to
it one at a time, each boot its own standalone deploy.replica. Only
the disruptive capacity-fallback restart mentioned above opens a
deploy.run.
- Regardless of which span parents a deploy or replica boot, the
background work itself never observes the triggering request's
cancellation or deadline - only its place in the trace. A client
disconnecting mid-deploy does not cut the deploy short; it only stops
that one client from seeing the outcome.
- Every exported span carries a resource identifying the instance
(service.name, service.version, service.instance.id, plus any
tracing.resource_attributes; see Multiple instances, one
backend).
This reuses the same tracing config block above; there is no separate
server-tracing switch. Server spans and the access log are correlated in both
directions (the span carries the request_id; the access-log line carries the
trace_id); see metrics.md for the access-log fields.
Scheduled jobs¶
When tracing.enabled is set, every scheduled job run gets a root
schedule.run span and the same OTEL_* defaults an app replica gets (see
Environment variables injected into each app
above), so the job's own spans, once it starts one, land in the same trace.
The span is opened the moment the run's row is inserted, so it covers the
whole admitted lifetime, including time spent waiting on the schedule's own
lock or an execution fence, through to the terminal status
(succeeded/failed/etc.). A run that never starts because another run is
still in flight (skipped_overlap) gets no span at all - there is nothing to
trace. Attributes: shinyhub.app.slug, shinyhub.schedule.name,
shinyhub.schedule.id, shinyhub.schedule.run_id, shinyhub.schedule.trigger
(see Run provenance),
shinyhub.deployment.id, and, once the run ends,
shinyhub.schedule.status, shinyhub.schedule.persisted, and
process.exit.code. Sampling follows the fleet sample_ratio like any other
root span (ParentBased(TraceIDRatioBased(sample_ratio))).
When auto_instrument_apps is enabled (or [tracing] auto = true in the
bundle), supported Python job commands use the app's instrumentation overlay,
including auto_instrument_extra_packages. Examples:
The shared bootstrap extracts TRACEPARENT and TRACESTATE, starts an active
process.run child span, and runs the script or module in that context. AWS
calls instrumented by botocore therefore appear beneath
schedule.run -> process.run. The process span covers Python execution,
whereas schedule.run also includes admission/lock waits and uv startup.
The SDK flushes completed spans at normal interpreter exit, including nonzero
Python exits. Forced termination can lose buffered child spans; the platform
still records the terminal schedule status.
The same bootstrap applies to native post-deploy Python hooks, with
deploy.post_deploy_hooks as parent. Hooks receive the shared exporter defaults
and their deployment identity, without a serving-replica attribute. Container
hooks remain skipped under the existing runtime policy. Jobs and hooks are
never retried as an instrumentation fallback: their side effects may already
have occurred. Process error spans record the exception type and a generic
status, without exception messages or stack traces.
Supported commands explicitly invoke python or python3 with a script,
-m module, or -c code, directly or through uv run. Common uv run options
and basic interpreter flags are retained. Shell wrappers, explicit/venv or
version-specific interpreter paths, uv script mode, and unsupported options
are passed through unchanged. [tracing] auto = false opts both apps and
one-shot commands out; tracing disabled leaves scheduled commands unchanged.
For unsupported commands or manual instrumentation, extract the environment
context in the job yourself. TRACEPARENT/TRACESTATE identify the current run
and override static per-app values. Exporter settings remain overridable:
import os
from opentelemetry import trace
from opentelemetry.propagate import extract
carrier = {"traceparent": os.environ.get("TRACEPARENT", "")}
if os.environ.get("TRACESTATE"):
carrier["tracestate"] = os.environ["TRACESTATE"]
ctx = extract(carrier)
tracer = trace.get_tracer(__name__)
with tracer.start_as_current_span("refresh", context=ctx):
refresh_data()
A job run with tracing disabled has no TRACEPARENT set; the snippet above
still works unmodified, since extract on an empty carrier returns the
current (background) context and start_as_current_span simply starts its
own root trace.
What ShinyHub does not do¶
- No embedded OTLP receiver. ShinyHub exports its own spans and propagates trace context, but it does not receive, collect, or visualise traces for other services. Run a collector (Tempo, Jaeger, Grafana Alloy, Honeycomb, etc.) and point ShinyHub and the apps at it.
- No app-span correlation in the ring buffer. The Traces-tab ring buffer is proxy-level metadata only; for full request data, follow the trace ID into your backend. (Server spans and the access log are correlated separately, as noted above.)
- No sidecar. The OTEL_* env approach uses the OpenTelemetry SDK that Shiny already loads, with no separate agent and no exporter binary on the host.
Admission rejection alerts¶
The existing Prometheus counter
shinyhub_admission_rejects_total{slug,reason} records proxy admission decisions
independently of trace sampling. For refused capacity admissions, for example:
sum by (slug, reason) (
rate(shinyhub_admission_rejects_total{reason=~"pool-saturated|render-paced|memory-pressure|cpu-saturation"}[5m])
)
Keep render-deferred separate: a wait page can re-poll several times for a
single visit, so that counter measures deferral responses rather than distinct
refused sessions.