Put context in reach.
Retrieve relevant runbooks and postmortems, filtered by service and incident context. Keep the source attached to each piece of evidence.
TraceHarbor is an AI agent for production incident investigation. It brings runbooks, deployment events, and logs together, chooses read-only checks, and builds hypotheses with sources engineers can verify.
A backend engineering portfolio project. Evidence before automation.
Alerts tell you something broke. An investigation needs context, sources, and a way to challenge the first explanation.
Retrieve relevant runbooks and postmortems, filtered by service and incident context. Keep the source attached to each piece of evidence.
Separate observations from possible explanations. Cite supporting evidence and say when there isn't enough to draw a conclusion.
Authorize read-only tools before execution. Record inputs and outputs so the investigation can be reviewed and reproduced.
Explore a fictional checkout incident. This interactive walkthrough uses fixed sample data; it does not call an AI model or connect to production.
[1] event-081[2] sample-metrics[3] pool-runbook §3Temporal correlation is a starting point, not proof of causation.
The deployment timing [1] and connection-wait signal [2] are consistent with the pool-exhaustion scenario described in the runbook [3].
Still unknown: whether the release changed pool configuration, whether traffic increased, or whether database queries slowed independently.
Evidence incomplete · no confirmed root causeNo rollback or infrastructure change is executed by this walkthrough.
The planned agent helps an on-call engineer decide what to inspect next. It gathers evidence, revises hypotheses, and keeps operational decisions with the engineer.
Use the affected service, severity, symptoms, and time window to retrieve relevant runbooks and past incidents through RAG.
Select an authorized read-only tool to inspect deployment history, metrics, or logs. Compare the result with the current hypothesis.
Update the hypothesis, cite its sources, and choose the next useful check. Stop when evidence is insufficient or the investigation budget is reached.
searchRunbooksgetDeploymentHistoryqueryMetricssearchLogs
Start with mock adapters. Connect real operational data after the core workflow is tested.
Check permissions before every tool call. Record tool inputs, outputs, and model versions in an audit trail.
The planned scope excludes automatic rollbacks, restarts, and database writes.
The MVP retrieves documents and returns cited answers. A bounded investigation loop with authorized tool calls follows in Production Core.
This landing page demonstrates the concept with sample data; the AI backend is not implemented yet.
The planned implementation starts with a modular monolith. Domain logic stays independent of the database, cloud services, and model provider.
The project is in development. These are planned milestones, not a list of shipped capabilities.
Incident creation, sample document ingestion, filtered retrieval, and cited answers that acknowledge missing evidence.
Proof: a reproducible local setup and an end-to-end example.
A bounded agent loop with authorized read-only tool calls, plus queue processing, idempotency, retry and replay, access control, audit, and observable failure paths.
Proof: evaluation results, load tests, threat model, and runbooks.
Real metric and deployment adapters, improved retrieval, and richer investigation workflows.
Expand only when the core is supported by evidence.
Join the waitlist for product updates and early-access announcements.
TraceHarbor · AI engineering through a backend reliability lens.