Search & citations
Search finds passages by meaning, not only by matching words. Every result comes with a citation: the repository, the file and the lines it came from, or the ledger topic. A citation lets a person, or another agent, check the source.
Where to search
- Agents use
search_corpus. For ledger work,ledger_searchadds filters for topic type and status. - People use the Search and Ledger pages in the console.
- Your own software can call
POST /v1/searchon the API. See Your own agent.
Every search runs inside one namespace. To search another, switch namespace in the console, or use that namespace's servers.
How results are ranked
- Match by meaning. Workstate turns the question into an embedding, a list of numbers that represents its meaning. It then finds the chunks whose meaning is closest. A chunk is a passage of a file.
- Rerank. A second model reads the question together with each candidate, and orders the candidates by how well they answer it.
- Return the best. Search returns the top results, 8 by default. A ledger topic appears only once, even when several of its entries match.
Search ranks by meaning, so describe what you want in plain words. You do not need the exact function name or phrase.
What each result tells you
| Field | What it tells you |
|---|---|
repo | The repository. For GitHub, it is the repository's name. For an upload, Confluence or Jira source, it is the source's name in lower case with spaces as hyphens, such as team-handbook. |
rel_path | The file's path in the repository. A Confluence page looks like ENG/Onboarding (1234).md: space key, title and page id. A Jira issue looks like ENG/ENG-42.md. |
start_line, end_line | The lines the passage covers. |
text | The passage itself. |
corpus | code, wiki or ledger. |
language | The detected language, such as python or markdown. |
ann_score, rerank_score | How close the passage was in the first pass, and how relevant it is after reranking. |
topic_id, type, status | Ledger results only: which topic matched, and where it stands. Ledger results have no file or lines, so you cite them by topic id. |
Each Confluence page and Jira issue is indexed starting with its title, where it lives, and a link back to the original, so its first passage includes that link.
With Each top-level folder is a repository turned on, each top-level folder of an upload source is a repository named after the folder. The Repositories page lists the repository names in the selected namespace.
Filters
These are the parameters of search_corpus. The API takes the same ones except rerank, with repositories as a list named repos.
| Parameter | Use it to |
|---|---|
corpus | Search only code, wiki (Confluence pages, Jira issues and uploaded files that are not code) or ledger. |
repo | Search one repository, or a list of repositories. |
language | Search one language, such as rust or markdown. |
path_prefix | Search only files under a path, such as docs/. |
top_k | Change the number of results. The default is 8, and the API returns at most 100. |
rerank | Turn reranking off, for debugging only. It is on by default. |
On the Search page, the corpus control offers All, Code, Wiki and Ledger, beside a repository filter.
Citing results
Ask your agents to cite each result they rely on as repo/path:start-end, for example payments-api/src/retry.ts:40-72. Ask them to cite ledger topics by id, for example payments-0012. Agent instructions has wording to start from.
A result is one passage, not the whole file. When the agent has the file itself, such as a local checkout, have it read the surrounding lines before it relies on the passage.
Limits
- Search sees what the last sync indexed. For GitHub, that is the latest commit on the default branch when the sync ran.
- Every GitHub file is in the
codecorpus, including Markdown. To find documents in repositories, filter bylanguage: "markdown". - Code in languages other than Rust, Python, TypeScript, JavaScript, C, C++, Bash and Ruby is split as plain text, so a passage can start or end partway through a function.