AI incident investigation

Every incident
leaves a trace.
Follow the evidence.

TraceHarbor is an AI agent for production incident investigation. It brings runbooks, deployment events, and logs together, chooses read-only checks, and builds hypotheses with sources engineers can verify.

A backend engineering portfolio project. Evidence before automation.

•••   investigation / INC-042ILLUSTRATION
SEV-2 · CHECKOUT SERVICE

Checkout latency rises.
Where do you look first?

Deployment v2.8.1 completeddeployment-history · event-081
Database connection wait increasessample-metrics · checkout-db
Runbook matches pool exhaustionrunbook · connection-pool / §3
WORKING HYPOTHESIS / NOT A ROOT CAUSEA connection-pool configuration change may explain the latency. Compare the deployment diff and pool metrics before acting.
[1] Deployment event   [2] Pool runbook
Tools: read-onlyHuman decision required ↗
PLANNED ENGINEERING STACKJava / Spring BootPostgreSQL + pgvectorAWS / SQSOpenTelemetry
01 / The approach

Less guesswork.
A clearer next step.

Alerts tell you something broke. An investigation needs context, sources, and a way to challenge the first explanation.

01 — GATHER

Put context in reach.

Retrieve relevant runbooks and postmortems, filtered by service and incident context. Keep the source attached to each piece of evidence.

02 — REASON

Make hypotheses inspectable.

Separate observations from possible explanations. Cite supporting evidence and say when there isn't enough to draw a conclusion.

03 — VERIFY

Keep engineers in control.

Authorize read-only tools before execution. Record inputs and outputs so the investigation can be reviewed and reproduced.

02 / A sample investigation

From an alert to
a testable hypothesis.

Explore a fictional checkout incident. This interactive walkthrough uses fixed sample data; it does not call an AI model or connect to production.

Three signals. One investigation.

A release completed five minutes before the signal.[1] event-081
Connection wait rose during the incident window.[2] sample-metrics
The runbook describes connection-pool exhaustion.[3] pool-runbook §3

Temporal correlation is a starting point, not proof of causation.

03 / Meet the investigation agent

An investigator that
shows its working.

The planned agent helps an on-call engineer decide what to inspect next. It gathers evidence, revises hypotheses, and keeps operational decisions with the engineer.

01 — UNDERSTAND THE INCIDENT

Start with the context.

Use the affected service, severity, symptoms, and time window to retrieve relevant runbooks and past incidents through RAG.

02 — CHOOSE A CHECK

Ask tools for evidence.

Select an authorized read-only tool to inspect deployment history, metrics, or logs. Compare the result with the current hypothesis.

03 — REASON AND REPEAT

Explain what holds up.

Update the hypothesis, cite its sources, and choose the next useful check. Stop when evidence is insufficient or the investigation budget is reached.

Observe → Select a tool → Read the result → Revise the hypothesis → Verify
Planned agent loop · bounded steps, time, and model cost
THE TOOLBOX

Read before acting.

searchRunbooks
getDeploymentHistory
queryMetrics
searchLogs

Start with mock adapters. Connect real operational data after the core workflow is tested.

ENGINEER IN CONTROL

Decisions stay human.

Check permissions before every tool call. Record tool inputs, outputs, and model versions in an audit trail.

The planned scope excludes automatic rollbacks, restarts, and database writes.

IMPLEMENTATION PLAN

RAG first. Agent next.

The MVP retrieves documents and returns cited answers. A bounded investigation loop with authorized tool calls follows in Production Core.

This landing page demonstrates the concept with sample data; the AI backend is not implemented yet.

04 / Under the hood

A small core.
Explicit boundaries.

The planned implementation starts with a modular monolith. Domain logic stays independent of the database, cloud services, and model provider.

  • Retrieval with metadata filters and source citations
  • Read-only tool authorization and an audit trail
  • Timeouts, retry budgets, and graceful failure
  • Evaluation, latency, token usage, and cost visibility
Incident → Investigation
↓
Spring Boot · Application + Domain
↓ adapters / ports ↓
PostgreSQL
+ pgvector
LLM provider
+ read-only tools
↑
Java ingestion worker · SQS
Cross-cutting: authorization · audit · observability
05 / Build in deliberate steps

Useful first.
Production-minded next.

The project is in development. These are planned milestones, not a list of shipped capabilities.

01 / PLANNED · MVP

The investigation loop

Incident creation, sample document ingestion, filtered retrieval, and cited answers that acknowledge missing evidence.

Proof: a reproducible local setup and an end-to-end example.

02 / PLANNED · PRODUCTION CORE

Reliable by design

A bounded agent loop with authorized read-only tool calls, plus queue processing, idempotency, retry and replay, access control, audit, and observable failure paths.

Proof: evaluation results, load tests, threat model, and runbooks.

03 / OPTIONAL · AFTER THE CORE

Deeper integrations

Real metric and deployment adapters, improved retrieval, and richer investigation workflows.

Expand only when the core is supported by evidence.

06 / Stay in the loop

Follow the next
TraceHarbor release.

Join the waitlist for product updates and early-access announcements.

Your email will be used for TraceHarbor updates.

Investigate with evidence.
Build with intent.

TraceHarbor · AI engineering through a backend reliability lens.

Walk through the case ↗