Incidents & defects
When something goes wrong, keep more than the fix. The cause, and what not to do again, are what save the next person time. The ledger has three topic types for this. This page explains which type to use, what to record, and how to search before you investigate a problem again.
Incident, defect or investigation?
| Type | Use it for | Example title |
|---|---|---|
| Incident | A failure in a live system or service, and its root cause | Checkout down for 40 minutes after a certificate expired |
| Defect | A flaw found before it caused an incident, such as a bug in code or an error in a template or process, and its fix | Invoice totals round tax per line instead of per invoice |
| Investigation | A question you researched, where the finding is what matters, even if nothing changed | The CRM export includes deleted contacts, flagged but not removed |
If a problem fits more than one type, choose the one a future reader would filter by. A decision that comes out of an incident, such as "Renew certificates automatically", is a separate decision topic.
What to record
A useful topic answers six questions. Each answer has a place:
| Question | Where it goes |
|---|---|
| What did people see? (the symptom) | Title and summary |
| Who or what was affected, and for how long? (the impact) | Summary, then the current state once it is known |
| Why did it happen? (the root cause) | Current state, once known |
| What fixed it? | An entry in the history, and the current state |
| What should nobody do again? | Current state |
| How could it be caught sooner? | Current state, or a separate topic for the follow-up work |
The title and summary are fixed when the topic is opened. So:
- If you open the topic during the incident, title it with the symptom. The cause goes in the current state when you find it.
- If you open it afterwards, put the symptom and the cause in the title.
- Write the summary as the problem looked when the topic was opened: the symptom, when it started, the impact, and how you noticed.
As the work goes on, ask your agent to append an entry for each step that matters: what you tried, what you ruled out, and when the problem was mitigated. Each entry is timestamped and records who wrote it. When you know the root cause, set a new current state that leads with it.
Statuses
Incidents, defects and investigations start as open, unless your agent sets another status.
| Status | Use it when |
|---|---|
open | Work is in progress, or the question is not answered yet |
resolved | It is fixed, or the question is answered |
wontfix (shown as Won't fix) | You decided not to fix it. Say why in the entry |
superseded | A newer topic replaced this one, for example a later and more accurate root-cause analysis |
Workstate does not enforce an order between statuses. Agree how your team uses them, and write it into your agent instructions.
Search before you investigate again
Before anyone digs into a problem, ask your agent to search the ledger for it. The same failure may already have a known cause and a workaround.
- Describe the symptom in plain words. The search matches by meaning, and it covers every entry in a topic's history, not only the title.
- If there are too many results, narrow them by type (
incident,defectorinvestigation), by status (for exampleopen), or by repository. - Before acting on a match, have the agent read the whole topic with
ledger_get. A search result is a pointer, not the full record.
People can do the same in the console. Open Ledger and choose Incidents, Defects or Investigations under Type. The page lists every topic of that type, most recently updated first. To search, describe the symptom in Search the ledger and select Search. The console has no status filter, but each row shows the topic's status. See Read decisions in the console.
Ask your agent
For example:
- "Checkout is failing. Before you dig in, search the ledger for past incidents about checkout or payment errors."
- "Open an incident under
checkout: every checkout has failed since 09:12 UTC. Put the symptom and how we noticed in the summary." - "We issued a new certificate and checkout recovered. Append that to
checkout-0031as a state change." - "We found the root cause. Set the current state to the cause, the fix, and what not to repeat. Then mark it resolved."
- "That rounding bug is a separate problem. Record it as its own defect under
billing."
One problem per topic
If an investigation turns up a second, unrelated problem, give it its own topic, and mention its id in an entry on the first. A topic that tracks several problems is hard to resolve and hard to find.
Example
Illustrative
This example is not real data.
This topic was opened during the incident, so its title names the symptom. The cause is in the current state.
| Field | Example |
|---|---|
| Id | checkout-0031 |
| Type | Incident |
| Status | Resolved |
| Title | Every checkout failing with "Payment could not be processed" |
| Summary | Since 09:12 UTC, every checkout fails with "Payment could not be processed", and no orders are going through. A customer reported it. No alert fired. |
| Current state | Resolved. Checkout was down from 09:12 to 09:52 UTC, and about 1,200 orders failed. Root cause: the payments service's certificate expired. It was renewed by hand a year ago and never set to renew automatically. Fix: automatic renewal, and an alert 14 days before any certificate expires. Don't renew certificates by hand. |
Its history, newest first:
| Entry | Written by | Text |
|---|---|---|
| Status changed | lee@example.com with Claude Code | Automatic renewal and the expiry alert are live. Resolved. |
| State changed | lee@example.com with Claude Code | Root cause found: the certificate was renewed by hand last year and had no automatic renewal. |
| State changed | ari@example.com with Cursor | Mitigated at 09:52: issued a new certificate, and checkout recovered. |
| Note | ari@example.com with Cursor | Every checkout failing since 09:12 UTC. Investigating. |
A month later, someone's agent searches for "payments failing after a certificate change". It finds checkout-0031, reads the cause, and checks the certificate first.