|
| 1 | +--- |
| 2 | +layout: page |
| 3 | +title: "Log Triage v2: Sub-Millisecond Incident Log Search" |
| 4 | +date: 2026-03-20 |
| 5 | +description: > |
| 6 | + Go-based log triage system that finds the closest log entry to an incident timestamp in sub-millisecond time across millions of log lines — reimagined from a production system at Amazon. |
| 7 | +--- |
| 8 | + |
| 9 | +**[View on GitHub](https://github.com/Mister-Raggs/log-triage-v2)** |
| 10 | + |
| 11 | +## Overview |
| 12 | + |
| 13 | +A reimagining of a production log triage system originally built at Amazon. The original system handled logs from a high-volume internal service generating ~42M rows/hour. During on-call incidents, engineers needed to quickly locate the relevant log entry nearest to a reported timestamp — the original solution reduced average triage time from 15 minutes to under 45 seconds. |
| 14 | + |
| 15 | +v2 rewrites the core in Go, strips out internal dependencies, and adds observability and Kubernetes-native deployment so anyone can run it. |
| 16 | + |
| 17 | +## Key Achievements |
| 18 | + |
| 19 | +- **Sub-Millisecond Lookups**: Binary search over a time-sorted in-memory index finds the nearest log entry across millions of lines in under 1ms |
| 20 | +- **Production-Scale Simulation**: Log generator replicates real volume (42M rows/hour / 11,667 lines/sec) without internal dependencies |
| 21 | +- **Built-In Observability**: Prometheus metrics for request counts, latency histograms, index size, ingestion totals, and parse error counts |
| 22 | +- **Kubernetes-Native**: Multi-stage Docker build (~15MB image), K8s manifests with health/readiness probes and resource limits |
| 23 | +- **Clean Separation of Concerns**: Ingestion, indexing, query, and analysis are distinct packages with clear interfaces |
| 24 | + |
| 25 | +## Architecture |
| 26 | + |
| 27 | +``` |
| 28 | +Generator (sidecar) --> Log File --> Ingestion Goroutine --> Sorted Index (binary search) --> HTTP Server (:8080) |
| 29 | +``` |
| 30 | + |
| 31 | +Deployed as a Kubernetes pod with the generator running as a sidecar that writes logs, and the server ingesting, indexing, and serving queries. |
| 32 | + |
| 33 | +## Why I Built It |
| 34 | + |
| 35 | +At Amazon, I built the original version of this system (Lambda + Java) to solve a real on-call pain point — digging through massive log volumes to find the entry closest to an incident timestamp. The production system used internal log aggregation tooling (Timber) with a 1-hour publish delay, binary search for fast lookups, and an MCP endpoint for access. It plugged into AWS Step Functions and LLM confidence scoring to automate on-call SOP workflows. |
| 36 | + |
| 37 | +I rebuilt it in Go to make the core idea portable: better concurrency primitives, smaller binaries, observability from day one, and no internal dependencies — so the architecture speaks for itself. |
| 38 | + |
| 39 | +## Technologies |
| 40 | + |
| 41 | +- Go |
| 42 | +- Docker (multi-stage, ~15MB image) |
| 43 | +- Kubernetes |
| 44 | +- Prometheus |
| 45 | +- Zap (structured logging) |
0 commit comments