Skip to content

Incidents & defects ​

When something goes wrong, keep more than the fix. The cause, and what not to do again, are what save the next person time. The ledger has three topic types for this. This page explains which type to use, what to record, and how to search before you investigate a problem again.

Incident, defect or investigation? ​

TypeUse it forExample title
IncidentA failure in a live system or service, and its root causeCheckout down for 40 minutes after a certificate expired
DefectA flaw found before it caused an incident, such as a bug in code or an error in a template or process, and its fixInvoice totals round tax per line instead of per invoice
InvestigationA question you researched, where the finding is what matters, even if nothing changedThe CRM export includes deleted contacts, flagged but not removed

If a problem fits more than one type, choose the one a future reader would filter by. A decision that comes out of an incident, such as "Renew certificates automatically", is a separate decision topic.

What to record ​

A useful topic answers six questions. Each answer has a place:

QuestionWhere it goes
What did people see? (the symptom)Title and summary
Who or what was affected, and for how long? (the impact)Summary, then the current state once it is known
Why did it happen? (the root cause)Current state, once known
What fixed it?An entry in the history, and the current state
What should nobody do again?Current state
How could it be caught sooner?Current state, or a separate topic for the follow-up work

The title and summary are fixed when the topic is opened. So:

  • If you open the topic during the incident, title it with the symptom. The cause goes in the current state when you find it.
  • If you open it afterwards, put the symptom and the cause in the title.
  • Write the summary as the problem looked when the topic was opened: the symptom, when it started, the impact, and how you noticed.

As the work goes on, ask your agent to append an entry for each step that matters: what you tried, what you ruled out, and when the problem was mitigated. Each entry is timestamped and records who wrote it. When you know the root cause, set a new current state that leads with it.

Statuses ​

Incidents, defects and investigations start as open, unless your agent sets another status.

StatusUse it when
openWork is in progress, or the question is not answered yet
resolvedIt is fixed, or the question is answered
wontfix (shown as Won't fix)You decided not to fix it. Say why in the entry
supersededA newer topic replaced this one, for example a later and more accurate root-cause analysis

Workstate does not enforce an order between statuses. Agree how your team uses them, and write it into your agent instructions.

Search before you investigate again ​

Before anyone digs into a problem, ask your agent to search the ledger for it. The same failure may already have a known cause and a workaround.

  • Describe the symptom in plain words. The search matches by meaning, and it covers every entry in a topic's history, not only the title.
  • If there are too many results, narrow them by type (incident, defect or investigation), by status (for example open), or by repository.
  • Before acting on a match, have the agent read the whole topic with ledger_get. A search result is a pointer, not the full record.

People can do the same in the console. Open Ledger and choose Incidents, Defects or Investigations under Type. The page lists every topic of that type, most recently updated first. To search, describe the symptom in Search the ledger and select Search. The console has no status filter, but each row shows the topic's status. See Read decisions in the console.

Ask your agent ​

For example:

  • "Checkout is failing. Before you dig in, search the ledger for past incidents about checkout or payment errors."
  • "Open an incident under checkout: every checkout has failed since 09:12 UTC. Put the symptom and how we noticed in the summary."
  • "We issued a new certificate and checkout recovered. Append that to checkout-0031 as a state change."
  • "We found the root cause. Set the current state to the cause, the fix, and what not to repeat. Then mark it resolved."
  • "That rounding bug is a separate problem. Record it as its own defect under billing."

One problem per topic ​

If an investigation turns up a second, unrelated problem, give it its own topic, and mention its id in an entry on the first. A topic that tracks several problems is hard to resolve and hard to find.

Example ​

Illustrative

This example is not real data.

This topic was opened during the incident, so its title names the symptom. The cause is in the current state.

FieldExample
Idcheckout-0031
TypeIncident
StatusResolved
TitleEvery checkout failing with "Payment could not be processed"
SummarySince 09:12 UTC, every checkout fails with "Payment could not be processed", and no orders are going through. A customer reported it. No alert fired.
Current stateResolved. Checkout was down from 09:12 to 09:52 UTC, and about 1,200 orders failed. Root cause: the payments service's certificate expired. It was renewed by hand a year ago and never set to renew automatically. Fix: automatic renewal, and an alert 14 days before any certificate expires. Don't renew certificates by hand.

Its history, newest first:

EntryWritten byText
Status changedlee@example.com with Claude CodeAutomatic renewal and the expiry alert are live. Resolved.
State changedlee@example.com with Claude CodeRoot cause found: the certificate was renewed by hand last year and had no automatic renewal.
State changedari@example.com with CursorMitigated at 09:52: issued a new certificate, and checkout recovered.
Noteari@example.com with CursorEvery checkout failing since 09:12 UTC. Investigating.

A month later, someone's agent searches for "payments failing after a certificate change". It finds checkout-0031, reads the cause, and checks the certificate first.

Next steps ​

Workstate is built by Nerdstorm Pty Ltd, Sydney.