项目文件夹

文件
Asim Aslam 550033dcce Optimize image formats and fix lease re-registration issue (#2959)
* perf: convert generated PNGs to optimized JPEGs (12MB -> 1.5MB)

The landing and docs loaded 18 AI-generated PNGs at 0.5-1MB each. They're
1200x800 RGB illustrations with no transparency, so they recompress ~8x
as progressive JPEG (quality 82) with no visible loss. Convert all,
update every reference (.png -> .jpg), and drop the originals (including
the unused hero.png). Generated images: 12.3MB -> 1.5MB.

* fix(registry/etcd): re-register when a lease silently expires (#2956)

The keepalive rework (long-lived KeepAlive instead of KeepAliveOnce)
moved lease renewal entirely onto the keepalive goroutine; the 30s
periodic Register now skips on the 'unchanged' check. The goroutine only
reacted to the keepalive channel closing, so a lease that expired
server-side without a prompt channel close (e.g. a partition that
outlasted the 90s TTL) left the node de-registered from etcd while the
cache still believed it was registered — and nothing re-registered it.
That is the hidden-failure mode reported in #2956.

React to a non-positive TTL keepalive response the same as a channel
close: drop the cached lease/hash so the next Register performs a full
re-registration. Extract the loop into keepAliveLoop and unit-test the
TTL-expired, channel-closed, and healthy paths (no etcd required).

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-06-10 08:24:49 +01:00

2.7 KiB

layout
layout
default

Observability

Observability

Observability in Go Micro spans logs, metrics, and traces. The goal is rapid insight into service behavior with minimal configuration.

Core Principles

  1. Structured Logs – Machine-parsable, leveled output
  2. Metrics – Quantitative trends (counters, gauges, histograms)
  3. Traces – Request flows across service boundaries
  4. Correlation – IDs flowing through all three signals

Logging

The default logger can be replaced. Use env vars to adjust level:

MICRO_LOG_LEVEL=debug go run main.go

Recommended fields:

  • service – service name
  • version – release identifier
  • trace_id – propagated context id
  • span_id – current operation id

Metrics

Patterns:

  • Emit counters for request totals
  • Use histograms for latency
  • Track error rates per endpoint

Example (pseudo-code):

// Wrap handler to record metrics
func MetricsWrapper(fn micro.HandlerFunc) micro.HandlerFunc {
    return func(ctx context.Context, req micro.Request, rsp interface{}) error {
        start := time.Now()
        err := fn(ctx, req, rsp)
        latency := time.Since(start)
        metrics.Inc("requests_total", req.Endpoint(), errorLabel(err))
        metrics.Observe("request_latency_seconds", latency, req.Endpoint())
        return err
    }
}

Tracing

Distributed tracing links calls across services.

Propagation strategy:

  • Extract trace context from incoming headers
  • Inject into outgoing RPC calls/broker messages
  • Create spans per handler and client call

Local Development Strategy

Start with only structured logs. Add metrics when operating multiple services. Introduce tracing once debugging multi-hop latency or failures.

Roadmap (Planned Enhancements)

  • Native OpenTelemetry exporter helpers
  • Automatic handler/client wrapping for spans
  • Default correlation IDs across broker messages

Deployment Recommendations

Scale Suggested Stack
Dev Console logs only
Staging Logs + basic metrics (Prometheus)
Prod (basic) Logs + metrics + sampling traces
Prod (complex) Full tracing + profiling + anomaly detection

Troubleshooting

Symptom Cause Fix
Missing trace IDs in logs Context not propagated Ensure wrappers add IDs
Metrics server empty Endpoint not scraped Verify Prometheus config
High cardinality metrics Dynamic labels Reduce labeled dimensions