Fixes the prod failed to upload metrics: Post ".../v1/metrics": EOF / connection reset by peer / http: server closed idle connection errors.
Cause. Services push metrics every 60s (PeriodicReader default). Alloy's otelcol.receiver.otlp closes HTTP connections that have been idle for 1m (idle_timeout default). The exporters' own transport keeps idle connections for 90s, so a push could reuse a connection the collector was already closing. net/http doesn't retry a POST, and the exporters don't retry transport errors, so that push was dropped.
Over 6h in prod: 595 failed metric pushes across 14 services and 33 pods (~3 per pod per hour), plus 26 failed trace batches.
Alloy itself is healthy (no restarts, low CPU) and the failures aren't clustered around its events.
A scaled local repro (server idle 100ms, pushes every 100ms) failed 138 of 300 pushes with client idle > server idle, and 0 of 300 with client idle < server idle.
Fix. Both exporters get WithHTTPClient(otlpHTTPClient()). That is a copy of the exporters' default transport (v1.46.0 ourTransport) with IdleConnTimeout 30s, below the collector's 1m. The client timeout is 10s, the exporters' default.
Trade-off: a custom client makes the exporters ignore OTEL_EXPORTER_OTLP_*TIMEOUT and the TLS certificate env vars. No service sets them (all only set OTEL_EXPORTER_OTLP_ENDPOINT). This is noted in the code comment and CLAUDE.md.
Impact of the bug: temporality is cumulative, so counters and histograms stayed correct. It cost a missing sample about 5% of the time and dropped some trace batches.
Review (Go Backend Expert): no Critical or High findings. It confirmed the transport matches upstream apart from the idle timeout, and that headers, compression and retry are unaffected. I applied its Mediums: the same fix for the trace exporter (confirmed failing in prod), and documenting the ignored env vars. Accepted gap: the unit test pins the invariant (client idle < 1m, timeout 10s) but not the WithHTTPClient wiring; an end-to-end test would need a fake clock. I skipped closing idle connections on shutdown, since the 30s timeout reaps them anyway.
Tests: go test -race ./... passes; prek is clean.
Rollout: after the release PR, bump all 14 services from v0.6.0.
Fixes the prod `failed to upload metrics: Post ".../v1/metrics": EOF` / `connection reset by peer` / `http: server closed idle connection` errors.
**Cause.** Services push metrics every 60s (PeriodicReader default). Alloy's `otelcol.receiver.otlp` closes HTTP connections that have been idle for 1m (`idle_timeout` default). The exporters' own transport keeps idle connections for 90s, so a push could reuse a connection the collector was already closing. net/http doesn't retry a POST, and the exporters don't retry transport errors, so that push was dropped.
- Over 6h in prod: 595 failed metric pushes across 14 services and 33 pods (~3 per pod per hour), plus 26 failed trace batches.
- Alloy itself is healthy (no restarts, low CPU) and the failures aren't clustered around its events.
- A scaled local repro (server idle 100ms, pushes every 100ms) failed 138 of 300 pushes with client idle > server idle, and 0 of 300 with client idle < server idle.
**Fix.** Both exporters get `WithHTTPClient(otlpHTTPClient())`. That is a copy of the exporters' default transport (v1.46.0 `ourTransport`) with `IdleConnTimeout` 30s, below the collector's 1m. The client timeout is 10s, the exporters' default.
**Trade-off:** a custom client makes the exporters ignore `OTEL_EXPORTER_OTLP_*TIMEOUT` and the TLS certificate env vars. No service sets them (all only set `OTEL_EXPORTER_OTLP_ENDPOINT`). This is noted in the code comment and CLAUDE.md.
**Impact of the bug:** temporality is cumulative, so counters and histograms stayed correct. It cost a missing sample about 5% of the time and dropped some trace batches.
**Review (Go Backend Expert):** no Critical or High findings. It confirmed the transport matches upstream apart from the idle timeout, and that headers, compression and retry are unaffected. I applied its Mediums: the same fix for the trace exporter (confirmed failing in prod), and documenting the ignored env vars. Accepted gap: the unit test pins the invariant (client idle < 1m, timeout 10s) but not the `WithHTTPClient` wiring; an end-to-end test would need a fake clock. I skipped closing idle connections on shutdown, since the 30s timeout reaps them anyway.
Tests: `go test -race ./...` passes; prek is clean.
Rollout: after the release PR, bump all 14 services from v0.6.0.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_014rfQ5HJ7zwfuYQcWXd3Mm3
Metrics are pushed every 60s and Alloy's OTLP receiver closes connections idle for 1m, while the exporters keep them for 90s. A push could reuse a connection the collector was closing and fail with EOF or connection reset; the data was dropped (no retry for a POST or for transport errors). In prod that was ~3 failed metric pushes per pod per hour, plus occasional trace batches.
Both exporters now use an HTTP client that closes idle connections after 30s. A custom client makes the exporters ignore the OTLP timeout and certificate env vars; none are set.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014rfQ5HJ7zwfuYQcWXd3Mm3
argoyle
scheduled this pull request to auto merge when all checks succeed 2026-09-19 13:44:50 +00:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Fixes the prod
failed to upload metrics: Post ".../v1/metrics": EOF/connection reset by peer/http: server closed idle connectionerrors.Cause. Services push metrics every 60s (PeriodicReader default). Alloy's
otelcol.receiver.otlpcloses HTTP connections that have been idle for 1m (idle_timeoutdefault). The exporters' own transport keeps idle connections for 90s, so a push could reuse a connection the collector was already closing. net/http doesn't retry a POST, and the exporters don't retry transport errors, so that push was dropped.Fix. Both exporters get
WithHTTPClient(otlpHTTPClient()). That is a copy of the exporters' default transport (v1.46.0ourTransport) withIdleConnTimeout30s, below the collector's 1m. The client timeout is 10s, the exporters' default.Trade-off: a custom client makes the exporters ignore
OTEL_EXPORTER_OTLP_*TIMEOUTand the TLS certificate env vars. No service sets them (all only setOTEL_EXPORTER_OTLP_ENDPOINT). This is noted in the code comment and CLAUDE.md.Impact of the bug: temporality is cumulative, so counters and histograms stayed correct. It cost a missing sample about 5% of the time and dropped some trace batches.
Review (Go Backend Expert): no Critical or High findings. It confirmed the transport matches upstream apart from the idle timeout, and that headers, compression and retry are unaffected. I applied its Mediums: the same fix for the trace exporter (confirmed failing in prod), and documenting the ignored env vars. Accepted gap: the unit test pins the invariant (client idle < 1m, timeout 10s) but not the
WithHTTPClientwiring; an end-to-end test would need a fake clock. I skipped closing idle connections on shutdown, since the 30s timeout reaps them anyway.Tests:
go test -race ./...passes; prek is clean.Rollout: after the release PR, bump all 14 services from v0.6.0.
🤖 Generated with Claude Code
https://claude.ai/code/session_014rfQ5HJ7zwfuYQcWXd3Mm3
Coverage Report
Total coverage: 23%